@af3i
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Builder. Ex-McK → PE → YC Founder → Enterprise Product Leader. Dad to three, plus a dog, wedded to my high school debate opponent. ☧
Atlanta
Joined April 2008
- Tweets483
- Following2.1K
- Followers415
- Likes4K
We're all talking about what AI means for software. The real question is what it means for everyone - and everything - else. nitter.cf/af3/status/20177116574…
A question I keep wrestling with this year is: what will be valued when intelligence is too cheap to meter?
I used Jev to classify 1,018 AI research papers.
The result: $0.08 total cost and 256ms median end-to-end latency per paper.
The pipeline was:
1. Summarize each paper with DeepSeek V4 Flash
2. Send the title + summary + 24 possible topics to Jev
3. Use Jev to classify each paper
4. Visualize everything on 1kpapers.com
The summaries cost $3.99 on @togethercompute. The classifications cost $0.08 on @typesafeai.
So for just over $4 of inference, I ended up with a pretty useful way to explore the top AI research papers from the past year.
I think this is where things are heading: different models for different parts of the workflow, instead of using one model for everything.
I’m running evals on the Jev classifications before replacing the current ones, but the site is already live: 1kpapers.com
ashby retweeted
i lost the lid to my coffee grinder and codex tracked down the part, found the dimensions, and made me a 3d print file that fit perfectly in one shot
ashby retweeted
I Haven’t retired from the @NFL … But if you're a WR and you want to be great, come to @GeorgiaTechFB
ashby retweeted
TIME TO RETIRE THE CLASSIC PRD
The PRD is a relic. Time to retire it.
The word "PRD" brings to mind a 10-page Word doc or spreadsheet with a Background section, a Goals section, a Customers section, and 7 pages of feature descriptions. Half of it is throat-clearing. The other half is so vague that engineering and design end up guessing what to build anyway. Nobody reads it twice. Most of the time nobody reads it once.
Replace it with the Product Spec.
A Product Spec is a different kind of artifact, designed for the people who actually consume it (engineers, designers, AI agents)
The Product Spec has four mandatory pieces:
• The problem: who is hurting, what they are doing today, why now
• The bet: a falsifiable hypothesis (if we ship X, then [specific user] will [observable change] within [time], measured by [metric])
• The success criteria: what concrete behaviors we will see when this is working
• The evaluation: how we will measure it, what the kill / scale / graduate thresholds are
The shift from PRD to Product Spec is structural. It is what Coach is designed to enable.
In the PRD era, the bottleneck was getting alignment from a room of humans.
Long docs and exhaustive sections were the price of that alignment.
In the agent era, the bottleneck is giving an agent or an engineer a tight enough specification that it can ship without 12 follow-up questions, ideally using the goal loop in your agent of choice (h/t Peter Yang)
A 10-page PRD fails that test. A 1-page Product Spec with clear acceptance criteria and evals passes it.
Founders and product leaders: stop calling your docs PRDs. Stop writing them like PRDs. The artifact you need is a Product Spec with a falsifiable bet, a specific problem statement, concrete acceptance criteria, and a measurement plan.
Sorry, it’s too late. I’ve already pictured you as the feckless, price-taking Haiku market participant, and myself as the shrewd, disciplined Opus market participant.
Replying to @AnthropicAI
But the quality of the model mattered a lot. In the simulated runs where Opus and Haiku models negotiated with one-another, the Opus models got substantially better deals.
Interestingly, though, participants in our survey didn’t pick up on this disparity.
I’ve seen a drastic rise in ex-McKinsey founders for this exact reason.
In SF it’s easy to get trapped in the AI groupthink. The rest of America is different. Consultants were basically FDEs for data/business decades before palantir made it cool
If you read this and don’t understand why it’s happening it’s an opportunity to reset your understanding of how the real world works.
The real world will need a ton of help actually getting agents going in the enterprise. Companies have legacy tech stacks they need to modernize, data in tons of fragmented tools, knowledge that isn’t captured or digitized, and change management needed to actually utilize agents effectively. And they have to do all this while still running their business day-to-day, unlike startups.
This is why there is so much opportunity for companies (software or services) to actually deploy agents in specific domains and workflows. This remains a big opportunity for both existing services providers but also tons of new startups as well. Every new technology wave produces a new era of consulting firms that can deliver on that technology.
It’s also why the FDE model is going to be alive and well for a long time because companies will want to have their vendor actually help drive the change management and implementation for their new workflows.
The people aren’t going away. Far from it.
ashby retweeted
how did Allbirds pivot to AI compute hardware before the shoe company literally called ASICS
What if the Pimiento hats and Map and Flag exclusives are a shibboleth to separate proud corporate partners from true ball knowers
Everyone knows about the merch at Augusta National but just down the road there are some pretty sweet lids ⛳️ #thepatch
maybe this is not yet clear, so let me state it plainly: as of right now Anthropic, and really a small number of individuals at Anthropic, has the capacity to directly attack and cause major damage to the United States Government, China, and generally global superpowers. government agencies like the NSA do not have internal models or defense capabilities that outclass frontier models. if they chose to do so, they could likely exfiltrate top secret information from government systems, gain control over critical infrastructure including military infrastructure, sabotage or modify communications between members of government at the highest level, and potentially carry on activities for some time without detection. the thing about having access to a huge number of zerodays your adversaries don't know about is it gives you a massive asymmetric advantage.
they did not exploit this to gain power or destabilize the world order. they publicly released the information that they had these capabilities and worked to mitigate these flaws. you should be grateful american frontier labs have proven themselves remarkably trustworthy and concerned with the public good. but it's critical you understand we are in a new regime. private entities now have power that directly rivals and impacts the government's monopoly on influence and violence. and anthropic is certainly not the only one, there's little chance OpenAI's internal models are far behind.
this trend will accelerate on virtually every dimension, not slow down. my prediction for how it plays out is the relatively imminent seizure and nationalization of labs by the US government, sometime over the next two years. it's very tough for me to see how they accept the existence of this kind of threat. but this adds a whole new class of governance issues, as then we've handed these extremely wide-reaching capabilities from private entities to public ones.
ashby retweeted
My 3yo wanted to use the computer like me so I made him his own terminal. He types whatever he wants, it responds with fun messages. No external deps, no ads, just keyboard practice and cause-and-effect thinking. He thinks he's hacking. github.com/meimakes/tiny-ter…
Maybe this is intuitive to those of us who build software, so let me crib a line from @johnrichards here:
Tools used to build software can be probabilistic, but the software itself — at least for now — needs to be deterministic.
In John’s world at @officiallyiru, AI-native SDLC doesn’t just speed up velocity. It can be used in the product to improve threat detection, surface risks faster, and improve policy generation. But the application of all those still has to be applied deterministically.
Enterprise software buyers are seeking burden transfer. They are buying dependable outcomes.
Could OpenAI and Anthropic build their own ERP? Sure, but I think they know that 1/ there’s far more value to be created building for greenfield and enhancing their own core product (models / inference), and 2/ if they succeed at #1, ASI will let them migrate off of their ERP pretty easily, having inflated away the data portability costs.