@BuiltinMindi
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
AI Coding Harnesses & Agentic Workflows
Mumbai
Joined June 2009
- Tweets3.1K
- Following864
- Followers171
- Likes1.6K
Pinned Tweet
Ten Grok Bot templates for the actual founder loop: survive, decide, ship, review, sell, post. Plus a shared memory repo so the whole org gets smarter at once, not ten private diaries.
A pile of bots is not a startup. A pile has no graph, no memory.
People saying Opus 5.5 >> Sol 6.1 dont know that there is a thing called token efficiency
nitter.cf/BuiltinMind/status/210…
Pareto Frontier is now shifted after GPT 6.1 sol launch
Its Mimo 2.6 pro -> GPT 6.1 sol med -> GPT 6.1 sol high -> sol max -> opus 5.5 xhigh -> opus 5.5 max
GPT 6.1 sol now becomes the work horse and undercuts opus 5.5 low, med and high efforts.
Sonnet 5.5 never made the cut as i mentioned yesterday.
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task
Pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, except that the cache read discount rises from 90% to 95%. GPT-6.1 Sol’s overall blended price for agentic workloads is therefore slightly lower than GPT-6 Sol. This represents an additional price cut, following GPT-6 Sol’s original 50% discount from GPT-5.6 Sol.
Key takeaways:
➤ Achieves near-Astra Intelligence: GPT-6.1 Sol gains 4 points in the Intelligence Index vs GPT-6 Sol, and 5 points vs GPT-5.6 Sol - landing 1 point below GPT-6 Astra. It makes significant gains in agentic knowledge work, improving 4 points and 5 points in AA-Briefcase v1.1 and GDPval-AA v2.1 respectively. Other notable gains include a 12 point jump in Terminal-Bench 4.0, a 5 point jump in Humanity’s Last Exam, a 6 point jump in GDP.pdf, and an 8 point jump in AA-Omniscience Accuracy coupled with hallucination rate falling from 60% to 54%.
➤ Pushes cost efficiency frontier: At max effort, GPT-6.1 Sol costs less than a quarter of GPT-6 Astra per Intelligence Index task ($0.72 vs $3.26). It also costs 31% less per task than GPT-6 Sol ($1.05) and 64% less than GPT-5.6 Sol ($1.99). All effort levels of GPT-6.1 Sol push out the cost efficiency Pareto frontier: for a given level of intelligence, there is no cheaper model.
➤ Pushes token efficiency frontier, but uses slightly more output tokens than GPT-6 Sol: GPT-6.1 Sol uses ~10-30% more output tokens than GPT-6 Sol across effort levels. However, due to the increase in Intelligence Index score, its low and medium effort levels are Pareto optimal for token efficiency.
➤ Gains in Coding Agent Index: GPT-6.1 Sol gains 3 points on GPT-6 Sol at max effort in the Artificial Analysis Coding Agent Index, and sits 2 points below GPT-6 Astra.
Congratulations @OpenAI and @sama on the launch!
things get even more clear on agentic coding cost index
Look at that gap between sol 6.1 Xhigh and opus 5.5
Insane how expensive opus 5.5 is. sonnet 5.5 doesnt make any sense at all
OpenAi Plugins and plugin extensions will kill a lot of apps, both web and native.
ChaptGPT reach makes it hard to beat.
Plus the convenience of all the tools being available to integrate.
Control the harness, control the cost
Accenture Responsible AI's September 24 paper, Control the Harness, Control the Cost, argued that coding agent spend lives in harness defaults, its cache-safe router recovered about 14% to 21% of model spend across a 10,000-seat enterprise emulation.
Its rule was to move work only where no running conversation had to rebuild its prompt cache: at session start, in side lanes, and at subagent launch.
Got access to Muse yesterday.
Here is Grok Bot and Muse working side by side now.
Same same but different.
Love how Muse is less verbose and thinks visually, it creates these visually rich Artefacts without me asking.
Grok bots is complete ecosystem with hundreds of connectors. Its much more robust and capable due to connection with cursor.
AutoTrust AI has released JEV-27B, an Apache-2.0 open-weights model that answers yes/no, multiple-choice and 0–5 rating questions in a single forward pass, returning calibrated probabilities (announcement).
A decision takes a median 137 ms on one B200, because the recipe trains only a small decision block on a frozen Qwen3.8-27B backbone. The reasoning path is untouched: HumanEval matches the base model, with every completion byte-identical.
For agentic coding Opus 5.5 high and med is always better and cheaper than sonnet 5.5 high, xhigh and max
For terminal coding too its similar.
and this is cost per task.
Now why should anyone use sonnet ?
tried it loved it, checking client mails, replying linkedin messages has and prospecting has never been easier
I led engineering at Google DeepMind.
Today, I'm proud to introduce Fo to give personal AI something no lab ever has... Humans.
Other personal AI's pretend AI can do everything. Fo employs humans to do tasks that AI cannot.
- 2x better at real-world task completion (beats other agents by 69%)
- 94% trust rate (4x less likely to leak private info vs Muse, Instinct)
Sign up for free: wajo.ai/join-wajo
I dont get it why frontier labs are competing on price ?
Ultimately lower end model prices will approach zero.
Its better to release an opensource model optimised for their own compute and a top end frontier model.
This is the reason grok team has not released an opensource model in a while.
Elon is genius when it comes to economics.
If cursor decides to onboard all opensource model hosted on their own compute, they can crush opencode, Pi then spaceXAi can sell massive compute as an api and get 10 times more data in contributor mode.
Grok 4.8 can compete the top end.