Sivanrayana retweeted
i co-lead a16z speedrun alpha. applications close sunday, oct 11. if you've been sitting on it, this is the week. bet on yourself.
Sivanrayana retweeted
There's a huge opportunity right now in being the deployment layer for AI into the economy. The amount of work it takes to change out workflows in enterprises tends to be far greater than anyone realizes or would prefer. Clearly this is what the applied layer of AI is going to look like in the form of software and agents, but also it opens up new services firms opportunities.
Legacy systems need to be moved to the cloud, data organization and access needs to be updated, software needs to be connected to agents in new ways, workflows need to be reengineered for agents, HITL needs to be figured out for the process, evals need to be generated and maintained, and the entire system needs to be continually updated as new models get released and new capabilities emerge. And the full list may even be longer.
AI is not the same as just deploying software. Software you generally did the implementation of an existing, well understood category of technology, then stepped back and the customer kept running. With AI agents, you're delivering actual work augmentation to the organization, which has a completely different set of complexities associated with it. You're no longer deploying tools that the company is merely enabled by, you're deploying work output in a process. Completely different implementation and enablement process.
As a result, this is going to open up lots of new kinds of firms and plays for existing firms to diffuse AI into organizations. We're going to see approaches by industry, by size of company, and by problem inside of companies. Traditional SIs will modernize and adapt (some will clearly not adapt as well), and new entrants will also be founded in this period that take advantage of this window. Great time to be an FDE or FDE firm.
Box CEO Aaron Levie (@levie) calls out the MASSIVE opportunity for AI deployment services:
"every single one of those companies, whether that's a 50-person firm or a multi-100,000 person firm, is gonna need an army of people to go in and help them with that transformation"
"when you go to that law firm and you go to that pharma company and you go to that bank, they need something that bridges the core technology to their workflow in their business process"
"somebody has to go into that organization and get it set up, and somebody has to go and provide domain expertise to this model so it really understands our particular business process"
Vendor FDEs are incentivized to get you hooked on their platform and to spend more money.
Indie FDEs are incentivized to use the best tool for the job and to save you money.
Hire indie FDEs.
I've tried 21 general-purpose AI Agents 😵💫
OpenAI Dots, Muse, Buzz, Manus, Grok Bot, Claude Cowork, Gemini Spark, Perplexity Computer, Instinct,Poke, OpenClaw, Microsoft Scout, Genspark,Lindy,Genii, Simular,Hermes, Toyo, Hello Haven, Zo,Vellum.
Quick take + demo for each ↓
Sivanrayana retweeted
Assumptions from even a month ago no longer hold.
The new bitter lesson for agent harnesses:
Sivanrayana retweeted
the fundraising market is very hot and very founder-friendly atm, but I'm still seeing a lot of founders make misguided errors that complicate or derail their raises entirely. a non-exhaustive list of own goals to avoid:
> don't use exceptional rounds to set expectations for your own raise. i.e. " raised $ X at $ Y valuation so I should do this too". don't do this, exceptions don't make the rule
> in same spirit, don't raise your valuation cap because of vibes / don't mechanically start asking for a higher cap because *one* investor converted, but interest in your round remains meager. the "cuck tranche" and rounds with exploding valuation caps are indeed very common, however valuation caps increase as a consequence of excess $$ interest in a company at a certain cap, not because there is a pre-determined mechanic in fundraising that dictates investor 2+ must pay more than investor 1 irrespective of external market conditions and round dynamics. only hot rounds (a spectrum) command increasing valuations, and if its not already hot you're not playing the same game. trying to extort new investors with a higher valuation while your round is not remotely hot is a good way to push interested folks away (and consequently slow your momentum even more)
> don't be needlessly greedy about dilution + valuation. to the last point, the way you get hot (thus enabling to raise more capital at more friendly/less dilutive terms) is to by closing capital quickly. thus the beginning of your raise should optimize for speed (i.e. make it easy for an investor to want to give you money), and as your momentum builds, you can increasingly optimize for price and dilution. the mistake here is to be a stickler about dilution from the beginning (too many founders seem to believe they are entitled to 10% dilution no matter what)
> don't think you're entitled to a certain investor base. i've met a lot of founders that say some flavor of "I only want X,Y,Z stratetic investors in my round and everyone else is beneath me". this once again misses the reality of how things work: the "hotter" you are, the choosier you get to be. as covered, you get hot by attracting a lot of capital to your round. it's orders of magnitude harder to attract a lot of capital when you're only pitching a small group of investors you consider to be "good enough"
Sivanrayana retweeted
Teaching LLMs to Plan — Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
Research paper: arxiv.org/abs/2509.13351 [30-page PDF]
Sivanrayana retweeted
An Empirical Study of Harness Design for Coding Agents
Researchers from UMass Amherst, Zoom, Emory University, and UNC Charlotte kept the AI model the same, changed only the software around it, and ran 176 test setups on AI coding agents.
They used four models and two benchmarks, one with 500 real problems from GitHub projects and one with 89 command-line tasks.
The researchers call the part they changed the harness.
It's the layer that decides how the model plans, which tools it can use, and what it keeps in memory as a task gets long.
Memory produced the largest swing in the study.
A model can hold a limited amount of text at once, and that limit is its context window.
With a 32k-token window and no memory management, 78.7% of the GitHub tasks failed because the model ran out of room, averaged across the four models.
NVIDIA's Nemotron-3 550B solved 6.4% of those tasks in that setup.
When the harness trimmed old tool outputs, summarized earlier steps, or did both, the same model solved between 51% and 58%.
Planning's effect depended on how strong the model was.
Without a plan, the smallest model, Nemotron-3 30B, stopped after a median of 5 turns, and its success rate on the GitHub tasks fell from 25.2% to 13.6%.
The two stronger models kept accuracy within 2 points with planning and cost about 30% less on the GitHub tasks, because they stopped re-checking work they had already finished.
Tools showed the same split.
When the 550B model got a plain terminal instead of ready-made file tools, it solved more GitHub tasks and cost 53% less.
Mistral Medium 3.5 went the other way and dropped from 68.6% to 45.4% on the same set.
The authors conclude there's no single best setup, and each part should be picked for the model, the task, and the budget.
A benchmark score measures the model plus its harness, and this paper shows how much of that number the harness controls.
Sivanrayana retweeted
The presumption that the Fed raising short-term rates reduces inflation is predicated on the belief that higher rates reduce demand and investment.
But what if higher rates don’t reduce demand and investment because the demand for intelligence and energy is unaffected by higher rates because winning the race for super intelligence has a near infinite ROI and the demand for compute will remain incalculable.
Why won’t higher rates at this unique moment in history therefore lead to more inflation as interest costs are embedded in everything?
And the problem is compounded as the more the Fed raises rates, the more inflation we will have and the more the Fed will need to raise rates further and so on.
But what if the old models don’t apply to the current paradigm and the Fed is wrong?
I think the Fed might have just made a mistake. Am I right or am I wrong?
Sivanrayana retweeted
The Consupocalypse
Consumer commerce is having its own SaaSpocalypse moment thanks to Muse, Instinct and personal agents. (eg: Month to date Booking share price is -21%, Expedia is -18%, and Airbnb is -18%)
IMO, when the dust settles, companies that will be left standing will have some or all of the following:
1. Differentiated physical (not just digital) infrastructure and fulfillment
2. Hard to access supply
3. Powerful local network effects
4. Proprietary data loops that allows better personalization than a general purpose agent can
5. Differentiated Membership / reward programs that drive true loyalty
6. Consumer trust
We will go beyond theoretical moats discussions and truly see which moats are durable.
Banger paper introducing Jev-as-a-Judge.
The overall finding is that you want to use a cheap judge for most of your evals and send only the uncertain calls to a frontier model.
This paper measures how well that works with JEV, TypeSafe AI's decision-only judge.
On 510 held-out preference pairs, a cascade that accepted JEV's confident verdicts and escalated the rest to GPT-6 Astra kept 99% of GPT-6's accuracy at about 57% of its fee.
JEV returns a verdict and label probabilities with no reasoning text.
It costs $0.044 per 1,000 judgments at a median latency of 0.152 seconds, against $12.182 and 1.885 seconds for GPT-6, about 277 times cheaper.
On ordinary preference and evidence-grounded factuality it stays within 3 points of GPT-6 (92.2% against 93.5% on RewardBench, 87.5% against 86.7% on HaluEval).
The gap grows to 9 to 20 points on tasks that require checking a derivation or rejecting an elaborately written wrong answer, such as JudgeBench (78.6% against 93.1%).
On several benchmarks, JEV's gap to GPT-6 is concentrated in its low-confidence decisions, which is why the cascade works.
The escalation threshold did not transfer for every fallback model, so the authors recommend setting it on your own data.
Paper: academy.dair.ai/papers/jev-a…
Nice paper showing how to re-evaluate a production agent at a fraction of the cost.
200 questions, 38.5% of the full benchmark, reproduce the full score to within 1.03 points.
The authors studied an analytics agent that serves tens of thousands of monthly users, using 574 historical benchmark runs split by date into calibration and held-out periods.
They compared random sampling, cached results, fixed representative subsets and adaptive testing based on item response theory.
Multidimensional 2PL adaptive testing gave the best fidelity.
The team deployed difficulty-stratified fixed subsets instead because they are simpler to run. Those subsets transferred to five other agent families without recalibration and stayed stable with calibration windows as short as one day.
Paper: arxiv.org/abs/2609.21267
Chat with Paper: academy.dair.ai/papers/effic…
Should you raise money for your startup? Ask these three questions:
1. Good valuation?
2. Good Investor?
3. Clean Terms?
JUST DO IT!
MAKE HEY WHEN THE SUN SHINES
DON'T OVERTHINK IT FOUNDERS
STARTUPS WITH CASH IN THE BANK DON'T GO OUT OF BUSINESS
Sivanrayana retweeted
In the obsession with AI, too many founders, leaders, managers seem to have forgotten that an engineer with motivation + drive + focus + AI >> hundreds of AI agents without any of those
The best products and companies are built by people and teams like that, and will continue to be so
Ask yourself, again, why the AI labs hire the best of the best (some of the most driven people in the industry; founders, CTOs etc) if people aren't making the differrence when everyone has access to similar tools already
DigitalOcean Managed Agents are here!
They work with major harnesses like Claude Code and Codex.
16,000+ tools. Idle agents stop using CPU.
Optimized to help builders scale agents in production. Worth checking out.
🤝 Paid partnership
DigitalOcean Managed Agents is now in public preview.
Run Claude Code, Codex, or your own LangGraph agent in a runtime environment that pauses when idle. Put its tools behind one governed endpoint, and pick from 75+ open and proprietary models. One cloud, one bill.
Prompts to get started available in the blog: do.co/4ysh3it
Sivanrayana retweeted
For top talent, four years is a lot of time to spend on a model designed for a different era.
Building things, working with exceptional people, & learning by doing is a better path.
Excited to play a small part, teaching “Building Financial Freedom” at the Academy!
Today, we’re launching an ambitious new school called The Horowitz Andreessen Academy.
Based in San Francisco, The Academy serves the most promising young high school graduates.
We think this can be an elite institution that attracts top tier talent. One that prepares students for the future rather than remaining stuck in the past.
The #1 goal is to help students learn to build, which is the most important skill in the AI era. They'll learn primarily by pursuing their own projects, either individually or in groups. There are classes and guest lectures, too, from some truly amazing people who have built modern-day Silicon Valley.
The Academy is designed as a network, since that’s the reason students go to school in the first place. Core to that network are our 10 Founding Partners: Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit, and Stripe. The network includes over 50 hiring partners and over 200 speakers and mentors. To join as a hiring partner or faculty member, you can apply on our website.
We raised $42M in funding led by @a16z. I'll be CEO and @pmarca and @eriktorenberg will join me on the board.
Applications are open for our Founding Class Fellowship, which will be one year and tuition-free. Eventually, pending regulatory approval, we plan to offer a two-year program that charges tuition, similar in cost to an elite private university.
We're looking for the most unusually ambitious young builders on the planet. Come join us in San Francisco:
theacademysf.com/
Sivanrayana retweeted
Conceptually you can access APIs for control panels in the enterprise use case. The API and MCP is less of an issue - the access, audit, ownership of actions and liability is. In the enterprise use case, one who accesses is responsible.
Replying to @nikesharora
Couldn't agree more.
@nikesharora do you take the same approach for Palo products? E.g will an "enterprise muse" be able to configure NGFW?
every major company is arriving at roughly the same offering:
persistent memory, email/calendar/messages, browser + computer use, background tasks, proactive notifications, voice, app/tool execution, ambient context, & some notion of a personal agent sitting above everything.
remarkable levels of convergence with very little differentiation whatsoever.
Banger paper from MIT and Sakana AI.
They show that self-improving coding agents work.
The best part is that their approach, Self-Improvement via Fast Tree-search (SIFT), runs at a tenth of the CPU hours of DGM.
They reach 35.1 percent on Polyglot with o3-mini after 30 expansions. DGM reaches 30.7 percent after 80 nodes of tree search.
SIFT does it in under 50 CPU hours and under 5 hours of wall clock. The Qwen3-30B configuration runs its full search at 224 CPU hours and $34 of API spend, a tenth of the DGM baseline.
The saving comes from where the money goes.
Benchmark evaluation is the runtime bottleneck, so an LLM judge ranks candidate self-modifications first and only promising candidates get evaluated.
Judge quality decides the run.
On TerminalBench, gpt-5.4-high as the pairwise judge finds a 36.7 percent agent against a 29.2 percent starting point. gpt-5 finds 34.5 percent, and its top-ranked candidate is not the best agent its search produced.
Paper: academy.dair.ai/papers/self-…
Build your own harness, folks.
This is absolute banger paper from NVIDIA on self-evolving agent harnesses.
(bookmark it)
They introduce SoL-Pi which cuts token traffic by nearly half.
And it matches its baseline harness on GPT-5.6 Sol and Opus 5.
More details below:
Instead of tuning a harness by hand, they run auto-research loops at the harness layer across many repository-derived and verifier-driven environments, keeping only the mechanisms that survive selection.
Four mechanisms survived:
> Action Fusion changes how actions execute
> Online Context Compact handles compaction during a run
> ObservationPack reshapes observation handling
> Evidence-Preserving Reducer covers delegated reading
On the 51-task EdgeBench evaluation, the savings translate to about a third off API cost. In dollars that is an estimated $8.75 to $13.50 per hour against native Codex and Claude Code harnesses, and $4.36 to $5.71 against the baseline harness.
Because the search runs across many environments rather than one, the retained mechanisms keep working outside the setting that produced them. Code is on GitHub under NVlabs.
Paper: arxiv.org/abs/2609.20519
Chat with Paper: academy.dair.ai/papers/sol-p…
Sivanrayana retweeted
talked to a devtool founder doing $2M ARR who gets zero traffic from AI search
meanwhile a competitor with 1/10th their users is getting cited by chatgpt on every query in their category
the difference came down to one thing:
the competitor wrote pages titled exactly like the questions ppl ask before they know the product exists
> "how do i do X without Y"
> "best tool for [use case] under $50/mo"
> "X alternative that works with [integration]"
the founder had better docs, more github stars, bigger brand
but when someone asked chatgpt for a recommendation, there was nothing to link to
we ran the audit in 30 minutes. found 14 queries where competitors were getting cited and they weren't.
they wrote 6 pages in a week. already showing up on 3 of them.
dm me if u want me to run this on ur company
takes 30 min and its free