CEO @TandemAIx. AI brains, agents & risks for FDEs & implementation teams.
San Francisco
Joined July 2013
- Tweets329
- Following381
- Followers115
- Likes42
Amazon blocking AI agents opens a door, but not the one people think.
3rd-party sellers make 61% of units sold on Amazon. They'll follow the agents wherever they shop.
But Amazon delivered ~30% of US parcels last year.
A startup can win the storefront. Amazon still ships the box.
Everyone assumes AI will eventually replace annotators. I'm betting the opposite.
Models are bad at extrapolating, and they learn from yesterday's problems. Every new thing AI touches (code, science, everything) creates new blind spots.
More AI → more surface → more blind spots → more annotators.
We’re already starting to see what kind of new jobs AI is creating.
AI requires significant technical work and surrounding services to deploy into the economy. This means jobs for AI engineers that build applied AI products sold to (or within) enterprises, FDEs to deploy agents into companies, new services firms for deploying AI, and more.
And even the published stats of new jobs undercount all the existing jobs that are transitioning to new areas of AI work in an enterprise. Many prior data, research, and software jobs large enterprises are also being repositioned for working with AI.
Every bank, life sciences company, manufacturer, and even law firm is bringing on more technical talent -or repositioning existing roles- to help with agent deployment in their companies. It’s a lot easier to picture what AI can replace vs. what it creates until it starts happening. Now we’re seeing what this looks like.
I love the story of Podium (a YC companies a few years back). They were selling to tire shops a way to send SMS to collect google maps reviews. Now it's a billion-dollar business.
I've seen several comments saying “no one uses AI.” That's just not true.
The trick is distinguishing *paying* from *using*.
People absolutely use AI for many search/“Googling” use cases: plan my trip, find something, explain something, etc.
And no one was paying Google directly for most of those searches either.
In fact, I'm not paying for AI either. I just use it through my existing Pro account.
A more interesting question is how much more usage makes a power user VS a free trier one. 100x? 1000x?
98% of US households aren't paying for AI yet
More charts in State of Markets II: a16z.news/p/state-of-markets…
Beyond explanation, also execution.
I give Claude my outreach list and ask it to build a temporary SaaS with accounts, persons, contacts, links etc.
Missing emails? @FullEnrich MCP. HubSpot URLs? Just ask. Need more info? Claude or ChatGPT can now browse for you.
All in batch.
You can execute super fast, keep it a few days and throw it away.
(But I'm not letting it send the messages directly without review, model still isn't good enough - or lacking context - to write without giving that cringe)
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
First 1000X developer
here's how i shipped 2,500 PRs last month to production
this was originally supposed to be for Cursor Compile in London. i couldn't make it since i was livestreaming for Grok @Bot Galaxy so i'm making it available for free here on X! watch it on 2x speed, i talk slowly
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Astra is just stupidly good.
Now the same question as always (same for SaaS), what about a full, finished product? how does it work when you need all the details together to work? when complexity explodes?
Is it there yet?
“harness capabilities are increasingly shifting into the model itself.”
This is significant news.
The bar to build AI-native apps just dropped.
The bar to build valuable AI-native apps just went up.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Read our post on Astra and what these results mean: arcprize.org/blog/astra
Christophe Barre retweeted
The hardest part of building an adjacent product inside a big company isn’t building.
It’s getting credible customer signal fast enough to know what deserves to be built.
A short thread on what I learned in my first month as an FDE at OpenAI:
Christophe Barre retweeted
Engineers who can find a company’s real problems are rare.
Engineers who can then scope and build the AI to solution are rarer. They barely exist.
Announcing Gauntlet FDE:
A hyper-intensive training ground for the next generation of forward-deployed engineers.
FDEs even in crypto now 🫨
Russ runs PS at Omni ($1.5B, revenue up 4x).
His hottest AI take: "The most overrated use case is asking for the latest status."
You burn tokens trolling Slack and Drive to learn what a good operator already knew.
His whole playbook is inverted like this. 👇
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Christophe Barre retweeted
Everyone's suddenly building company brains.
Nobody agrees on what's inside one. 😵💫
So we opened up 9+ company brains to see how they're actually built.
Every single one does the same four things: getting signals, remembering, dreaming & pruning, speaking & searching.
𝟭. 𝗚𝗕𝗿𝗮𝗶𝗻
Garry Tan's open-source personal brain. Your email and calendar flow into a git repo, and a nightly job re-links everything and flags what's gone stale.
𝟮. 𝗺𝗲𝗺𝟬
A memory library you call from your own code. It only stores what you explicitly tell it to, and ranks fresh facts above idle ones at search time.
𝟯. 𝗟𝗲𝘁𝘁𝗮
For building agents that remember across sessions. The agent decides what's worth keeping, and a second agent tidies up its memory in the background.
𝟰. 𝗭𝗲𝗽 / 𝗚𝗿𝗮𝗽𝗵𝗶𝘁𝗶
A knowledge graph with a clock in it. When a fact changes, the old one gets an end date instead of being overwritten, so you can still ask what was true last March.
𝟱. 𝗦𝘆𝗹𝗽𝗵
A content brain that lives entirely in a git repo. Agents write drafts, humans publish, and afterwards the agent reads your edits to learn what it got wrong.
𝟲. 𝗗𝗜𝗬 (𝗖𝗹𝗮𝘂𝗱𝗲 𝗖𝗼𝗱𝗲 + 𝗴𝗶𝘁)
What most engineering teams actually do. Markdown in the repo, grep instead of search, and pull requests as the only thing keeping it honest.
𝟳. 𝗣𝗹𝗲𝘁𝗼𝗿
A brand brain for marketing teams. Campaigns, assets and performance data in one tree, and the brand rules only move when a human signs off.
𝟴. 𝗚𝗼𝗿𝗴𝗶𝗮𝘀 𝗖𝗼𝗿𝘁𝗲𝘅
Built in-house by an eight-person AI team. 12,000 markdown nodes in GitHub, and every night the questions it got wrong become PRs that fix it.
𝟵. 𝗦𝗹𝗶𝘁𝗲 𝗔𝗴𝗲𝗻𝘁
For teams whose knowledge lives in docs and across sources. It watches ~20 connected tools (Slack, Drive, GitHub, Jira, etc.) for what's gone stale and sends the diff to whoever owns the page. Nothing changes without human approval.
We just launched an interactive ebook with architecture notes from real 149 teams of builders and users, interview insights, and the complete research.
The ebook is free, get it here: slite.com/ebooks/company-bra…
Which brain would you pick? 🧠
Context point is very true. In fact, you don't even need wait for FDEs to rotate every quarters to face it.
Context already doesn't circulate well:
- FDE need to discuss a feature/change with engineering team
- Hiring or junior FDE on a project
- Just your lead or head doesn't have all context about a project
We are working on solving that at Tandem. A brain per account, shared context to the team.
FDE interviews should test for:
- claude code
- product skills
- handle customer relationships
Christophe Barre retweeted
Forward-deployed engineer job postings grew 4x in 6 months while the overall AI Engineering market doubled.
I analyzed 113 FDE job descriptions to understand what companies actually want.
The results:
90% require direct client engagement
87% expect you to build production systems
62% need system integration work
51% want you to scope projects from scratch
The skill requirements are also broader than typical AI engineering positions.
I broke down:
- Which programming languages appear most (Python: 89%)
- The specific AI technologies companies ask for
- How FDE differs from solutions engineer, consultant, and AI engineer
- A self-assessment checklist to see if this role fits you
Full analysis with data and charts: aishippingblog.com/p/what-ai…