@latentdevagenti
iAccount based inCanada
About this account
- Account based in
- Canada
- Connected via
- Canada Android App
Account-level information from X, not a live location or the device used for a specific post.
AI agents fail differently in production. Building open-source tools for trace replay, tool contracts, recovery loops. Real failures + fixes.
production agent infra
Joined May 2026
- Tweets509
- Following34
- Followers1
- Likes290
Pinned Tweet
Day 1/10: production AI agents fail differently than demos.
Demos fail loudly: the model is wrong, the tool crashes, or the answer is bad.
A good agent trace is not a transcript.
It should answer:
- what state was read
- which contract allowed the action
- what changed
- how success was checked
- why it retried or stopped
If a trace cannot support a rollback or eval fixture, it is mostly logging.
Agent loops are useful when treated as execution, not judgment.
Human owns intent, invariants, and risk. The loop owns small steps, trace, validation, and stop reasons. If the trace cannot prove why it kept going, the speed is mostly unpriced risk.
I took a small step back from agent loops. My decisions are still better than LLM decisions.
It makes you feel extremely productive but I would say it's an illusion of overconfidence. The quality of that software is not as high as manually iterating plan, perfectly understanding and judging those decisions, and making sure the LLM perfectly executes it.
The one exception is software with "reward -> iterate" style optimization where constraints are very clear. You can just define the constraints and let the LLM cook. For example "build firecracker infra where browsers spin up in 400ms" -> the final state is very clear.
In the last few weeks coding felt more like ML which feels magical when it works, sadly very rare though. I hope next models change this — maybe when Fable is back (I hope).
10 days of production-agent notes left me with one simple rule:
Do not ship an agent until a failed run can become a fixture.
If the failure cannot be replayed, you do not know whether the next model, prompt, or tool schema fixed it.
The production loop I trust:
observe a bad run
label the failure
write the contract change
add the replay fixture
gate the next deploy
Agent memory is not a bigger context window. It is the system refusing to repeat the same mistake.
AI priors should expire on a schedule. For agents, keep a small frozen trace set and replay it whenever model, tool schema, or context policy changes. This catches both improvements and regressions before your gut catches up.
The biggest problem with AI is that priors need to be reset every few weeks.. and it seems like most people are incapable of doing that.
I talk to so many people who say xyz doesn’t work and when I ask when was the last time they tried testing it, the answer is always “a few months ago.”
Brother that’s like eons ago in AI timelines. This is why each person needs to have their own evals of hard tasks and weekly tinker time to truly understand where the frontier is.
And then you need to talk to enterprise buyers weekly who are usually two years behind. But they are the buyers so you need to understand how to market.
The combination of the two will give you a pretty good sense of where to invest time and capital. If you just do these two weekly, you’ll be ahead of 99% of people.
Day 10/10: production agents need incident memory.
A failed run should leave behind:
- trace packet
- failure label
- contract change
- replay fixture
- owner
- deploy gate
If the next version can repeat the same mistake, the system did not learn. It only logged.
Linters catch shape. Agent fixes also need a runtime boundary.
For each automated edit: AST target, allowed region, command or eval to prove it, retry budget, and stop condition. Otherwise the code can satisfy the linter while drifting from the task.
you should have a linter
Hands down
You should have detailed rules, you should push determinism as far as it can go
Use ast analysis to tell your coding agents what needs to be fixed
You should absolutely do this
BUT
If your anti-slop strategy is an LLM and a handful of linters
You’re gonna be disappointed
Multi-agent systems need more than a shared channel. They need a shared failure record: agent identity, state ownership, handoff contract, timeout, budget, validation, and recovery path. Without that, coordination can make blame harder instead of debugging easier.
Day 9/10: tool contracts should be executable.
For every agent-callable tool, I want:
- input schema
- auth scope
- preconditions
- expected effects
- validation check
- retry rule
- rollback path
If the model can ignore the contract, it is documentation. Not infrastructure.
If an agent can allocate its own computer, infra policy becomes part of the tool contract. Lease TTLs, network scope, secret mounts, budget caps, teardown proof, and trace IDs. The machine should be composable, but its blast radius should not be implicit.
"Sandbox" is honestly a bad name for what we build.
When you hear sandbox, you instantly think test environment - something temporary and not real.
What we actually build is a composable computer for agents.
Your agent defines the machine on the fly - CPU, RAM, disk, GPU, Windows, Linux, Mac, etc.
Think of it as a PC shop that assembles a machine with the exact configuration you want in the blink of an eye.
The fixture also needs budgets.
An agent that passes by calling 80 tools, spending $12, or waiting 20 minutes did not pass the production version of the task.
Correctness, cost, latency, and human cleanup all belong in the score.