@zencoderaii
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
The Most Intuitive AI Coding Agent - Code faster, smarter, and stay in the flow.
Joined May 2024
- Tweets590
- Following296
- Followers1.3K
- Likes171
Writing code got cheap. Reading it did not. The engineer who can look at a large diff and say no is worth more than they were two years ago, not less. Every workflow worth building still assumes a human reads the thing.
An agent writes eighty percent of the change in about a minute. The last twenty percent is where prod breaks, and that part does not compress. Teams that optimized only the writing half moved their bottleneck. They did not remove it.
When something ships broken the ticket says the AI got it wrong. Which model. On what task. With what context. If nobody on the team can answer that, it is not a model problem, it is a logging problem. Put the model name in the commit.
Every AI coding demo builds the same thing. Empty folder, todo app, four minutes, applause. Your repo has a decade of decisions in it and one service nobody wants to touch. The gap between those two is the entire job. Judge tools on the second one.
Most agent failures are not reasoning failures. The model did not have the file. It did not know the convention. It never saw the ticket. Before you reach for a bigger model, go look at what you actually handed the one you have.
Loud failures are cheap. A test goes red, the agent retries, nobody is hurt. The expensive ones are quiet: correct syntax, plausible logic, wrong assumption, and it clears review because it reads well. A second model from a different family catches those. A tired human at five does not.
Eighteen frontier models in one workspace. Opus 5, GPT 5.6, Gemini 3 Pro, and the rest of the list. The length is not the point. The point is that the right model for a migration is not the right model for a formatting pass, and which one that is keeps moving.
A rename does not need your most expensive model. Neither does a doc fix or a formatting pass. Your team already routes compute by task size everywhere else in the stack. Model selection is the one place that habit stopped.
The eval that beats every leaderboard: take ten tickets your team already closed, run them through the model you are considering, and read the diffs yourself. That number will not match the leaderboard. Yours is the one that predicts your repo. It takes an afternoon.
The model that wrote the code is the worst available reviewer of it. Same weights, same context, same blind spot. It says yes because it genuinely cannot see what it already could not see. Route the diff to a different family than the one that wrote it. Who reviews yours?
One model plans well. Another refactors better.
Force one to do everything and you leave quality on the table.
Which one do you actually trust with a refactor?
CI/CD Migrations: What Agents Can and Cannot Own nitter.cf/i/broadcasts/1YxNrbbak…
Thursday, 1pm Central. Praveen G on the parts of a CI/CD migration that cannot be agentic: cutover, secrets, permission models, and proving nothing got dropped.
youtube.com/watch?v=sjYqqsvk…
One person can run 20 agents today. Email, content, code, operations, all executing at once.
The thing they still cannot do is decide what's worth doing.
James runs about 18 across his own businesses. The agents do the work. He picks the work, and he reads anything that goes out.
Where's your line on what an agent gets to send on its own?
There is no best model. There's a best model for the step you're on.
How we route inside Zenflow:
Planning goes to the model with the deepest reasoning.
Implementation goes to whatever is fast enough to run often.
Review goes to a different model than the one that wrote the code.
That last rule is the whole point. A model grading its own work agrees with itself.
How are you splitting yours?
AI made everyone a 10x coder. It also made everyone a 10x debt machine.
Code you didn't read is a loan. It ships Friday and the interest shows up in the sprint where nobody can explain what it does.
What's the oldest piece of AI code in your repo that nobody understands?
Four things we check before an agent's pull request gets merged:
Does it touch auth, money, or anything a user sees? A human reads every line.
Did the tests already exist, or did the agent write its own?
Is the diff bigger than the ticket asked for?
Can one person explain the change without opening the file?
The last one catches the most.
What's on your list?