@wittybricksi
iAccount based inUnited States!
About this account
- Account based in
- United States
- Connected via
- Web
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Building blocks & puns in equal measure.
Oxford
Joined November 2021
- Tweets935
- Following69
- Followers22
- Likes3.2K
Max Well retweeted
Businesses, it’s time to get ready to start selling via @muse:
1/ Add your API to Muse Connector Platform => muse.ai/platform
2/ Accept agentic payments via @stripe => docs.stripe.com/agentic-comm…
3/ Profit! (with high-value consumers with @link wallets with high-value intent)
we're opening muse connectors to developers!
we have seen so much excitement in the developer community, integrating muse into everything from robots to mood lights and more.
plug your API into muse, then people can use your service just by asking for it. every request runs in a secure VM, and muse will ask before anything consequential.
come build with us! muse.ai/platform
Max Well retweeted
"All the warning lights are flashing red."
Stuart Russell, AI researcher and professor emeritus at UC Berkeley, says we are building systems we still do not know how to control.
The race is scaling intelligence faster than our ability to understand or govern it.
Open intelligence means little without reliable control.
Max Well retweeted
Jensen Huang:
“𝐑𝐞𝐧𝐭𝐚𝐥𝐬 𝐨𝐟 𝐭𝐡𝐚𝐭 𝐨𝐧𝐞 𝐠𝐢𝐠𝐚𝐰𝐚𝐭𝐭 𝐀𝐈 𝐟𝐚𝐜𝐭𝐨𝐫𝐲 𝐢𝐬 𝐚𝐛𝐨𝐮𝐭 $𝟓𝟎 𝐁𝐈𝐋𝐋𝐈𝐎𝐍 𝐩𝐞𝐫 𝐲𝐞𝐚𝐫”
$NVDA + $IREN are targeting deployment of up to 5GW of AI Factories.
5GW = $250B/year in potential AI-factory rental economics.
$IREN current market cap: $17.2B.
People ask "𝐖𝐡𝐞𝐧 𝐃𝐞𝐚𝐥??"
The deal is already done.
The market just hasn’t caught up yet.
Max Well retweeted
It’s here, seeing rave reviews and the “magic” moment for a lot of folks once they get over the intimidation of setting it up (pretty low on consumer)
The obvious gap right now is credentials (2fa, passkeys, etc.) and financials (credit cards not marking as fraud, controls that limit spend but don’t block usefulness of it) to enable seamless agentic experience for commerce, reservations, etc.
This would be another takeoff moment. Average person is too scared to nor wants to understand vibe coding, different model types & providers, work on laptop, etc. we tinkerers do but most just want to talk to their magic AI black box and just have it work.
Max Well retweeted
Odyssey announced Odyssey-3 today: One foundation world model that drives robots, humanoids, cars, drones & game agents.
Not one model per domain. One model, many bodies.
An LLM predicts the next word. A world model predicts the next state of reality like momentum, contact & what happens when this thing hits that thing. 🤷♂️
The demo they led with was driving on Indian roads. Unmarked lanes, mixed traffic, informal right-of-way.
If you want to prove a policy generalizes, that's the hardest test there is.
No independent benchmarks yet. Every performance claim here is still the lab's own.
Max Well retweeted
OpenAI's Astra just flipped Anthropic's Fable for enterprise AI spend.
OpenAI's share of OpenRouter spend went from 20% to 50% since June, while Anthropic fell from 80% to 50% over the same stretch.
That's OpenAI's best showing there since Feb 2024.
The catalyst map:
→ Anthropic shipped Claude Fable 5.1 on Sept 1
→ OpenAI answered two days later with GPT-6 Astra
→ Astra now holds 19% of wallet share, Fable 5.1 sits at 6%
@tryramp's enterprise data shows the same break. Astra grabbed roughly 13% of frontier-model spend within two weeks of launch, while Fable 5.1 is holding near 8-9% after peaking closer to 10%.
Note: Ramp's broader company index still puts Anthropic ahead overall, 43.8% to 39.8%, so this fight covers the newest models, not the whole business (at least, for now).
Follow for more @MilkRoadStocks
Max Well retweeted
OpenAI co-founder and president Greg Brockman says the company pulled a quarter of its production engineers off their roadmaps and told them their new job was defending the company:
"And in the case of cybersecurity, how I think about it, we at OpenAI took our models and applied them to finding vulnerabilities."
"We took 25% of our production engineers and said, sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the models to find all the holes."
"And we found a number of serious issues and we fixed them."
"And I've talked to a number of CISOs over the past couple of weeks and months and there are many companies who are also telling me that they've applied these models, they found some very significant issues, but that they're able to fix them."
"And one positive sort of part of the story is that when we took Astra, pointed out our systems, we found some new problems, but eventually it saturated. We basically have found, to our knowledge, all of the P zeros, all of the critical problems that Astra is smart enough to find."
Sit with that saturation: the scan stopped at the edge of what one model could see, not at a clean codebase, and a bug Astra can't reason about looks exactly like an empty queue.
The lesson isn't the tooling, it's the staffing. Brockman's own evidence is that OpenAI put every project on hold for 25% of its production engineers and reassigned them to harden the company's systems. The model surfaced the issues, but people still had to investigate and fix them. A scanner you buy is a purchase; a quarter of your engineers is a decision someone has to make and defend.
(I run TrustModel, which does independent AI evaluation.)
Max Well retweeted
An orchestration lesson from the Grok Bot Galaxy Livestream with Lauren (@poteto):
Tell Tater to use Cloud agents going forward.
Grok Bot is powerful for orchestrating bots.
Cursor's harness is really good for coding.
What to do:
1. Have Grok Bot spin up Cursor Cloud agents for you.
2. Those agents run on Cursor's harness with their own VM.
3. They can test the app, click around, and burn CPU while you stay in the Grok Bot chat.
4. Do engineering work from Grok Bot by telling your bot to spawn a cloud agent.
Next action: open your coding Grok Bot and say: "spawn a cloud agent for this repo and come back with a screenshot."
Three SpaceXAI employees are building a company in 3 days with Grok Bot. This is Day 1.
Matt Palmer (@mattyp), Lauren Tan (@poteto), and Roshan Sadanani (@roshan_s) start with research, a plan, and a product.
Live now, plus sessions for engineering, product, and founders.
nitter.cf/i/broadcasts/1AxRnZbVp…
the $0 output cost completely changes the unit economics of ai automation.
typesafe just shipped jev. it replaces sequential token generation with parallel computation. you get smart if-statements with calibrated probabilities directly in your code.
this is the infrastructure required to scale workflows without bleeding margin on api calls. clean architecture.
🤝 Paid partnership
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Max Well retweeted
Yale professor Robert Shiller explains how easily luck gets mistaken for skill.
An AI trading model can fool us the same way.
One strong backtest is not proof of an edge.
Before scaling:
→ Test on unseen data
→ Compare with simple baselines
→ Repeat across market regimes
→ Separate process from outcome
A weak model can still get lucky.
The edge is not winning once.
It is proving the result can repeat.
I share more breakdowns like this here.
Want the details?
Drop “VALUE” in my DMs.
Max Well retweeted
In 2008 the UK paid the same price for natural gas as the US.
Today we pay eight times more.
Yet another example of how badly this country is and has been led.
If govt had just stayed out of it, this would not be the case. But no. They had to get involved.
Max Well retweeted
‼️ OpenAI just turned a year of engineering into one API call
The Agents API ships the exact harness that runs Codex as a managed service. Compaction, tool search, subagents, parallel calls, all for the price of tokens with zero platform fee: the-ai-corner.com/p/you-spen…
Launch customers report:
▫️ 4x lower latency
▫️ 60% lower cost per task
▫️ 86% fewer failed responses
▫️ Ciridae moved its eval score from 0.71 to 0.85
Most of the internet has the conclusion backwards, and the interesting part sits in three lines of fine print nobody is posting, including the one that rules out regulated and European work today.
Full playbook with the code, the migration framework and 7 agent ideas 👇
the-ai-corner.com/p/you-spen…
Did you build a harness this year?
Max Well retweeted
Most people will use a 1M token context window like it's a 32K one.
That's the real gap. Not the model.
Three habits that change how useful V4.1 Flash feels:
✔ Stop pasting single files → drop the whole folder in, then ask your question. It holds the project structure, so it stops inventing function names that don't exist in your code.
✔ Use the vision side → hand it a screenshot of a layout you like, or a photo of something you sketched on paper, and ask it to build that. No separate tool needed.
✔ Try the harness modes → Standard gives the agent the full toolset. Code mode has the model write code to orchestrate multiple rounds of tool calls.
Minimal strips it back to a shell and a file editor. Creator lets you inspect the runtime and build your own. → Run the same task in Standard, then again in Code mode, and keep whichever fits the way you actually work.
Most people leave it on Standard forever and never find out what the other three do.
Save this video, you'll get more out of a free model than most people get out of a paid one.
Want the SOP? DM me.
Max Well retweeted
"We don't believe in this everything app thesis."
"The everything apps of the world won't be such a thing in a couple years because all the crazy new experiences that get built will need their own dedicated front end."
"We think if you can solve the movement between apps, you allow this crazy app ecosystem to emerge and thrive."
"So people will have 20 financial apps on their home screen instead of one or two. That's the world that we're building towards."
@jjjjacobx, CEO of @BlinkCashX on the live show.
Max Well retweeted
Mistral raised 3B euros at a 21B euro valuation, Europe's largest tech equity round on record. CEO Arthur Mensch says it closes the compute gap with Chinese AI labs. Europe's frontier AI bet just got real money behind it.
#AI #ArtificialIntelligence
Max Well retweeted
ONE PROMPT TRIGGERS FIVE SEPARATE SYSTEMS BEFORE A SINGLE ANSWER COMES BACK
context intake alone pulls from six different sources into one 32k window, a system prompt, twelve tool schemas, eighteen turns of history, four retrieved docs, the actual query, and scratch space, every one of them competing for the same room
memory isn't one bucket either, working memory expires in 6.7 seconds, episodic holds eight logged events, semantic carries twenty four facts, procedural holds five skills, and every fact stays linked back to the source it came from in both directions
the loop only gets three tries out of five before it has to stop, each pass narrows scope down to read-only, frees seventy two percent of the context back, and promotes two facts actually worth keeping
the output isn't generated, it's routed, recall pulls from memory, route reads the graph, gate checks the harness, act is the only step allowed to touch anything real, and all four trace back to a decision you can point to
one prompt in, five systems decide, one answer out
save this before someone calls a single model call "the agent" ↓
Max Well retweeted
Excluding China, Revolut is the most-downloaded financial app worldwide this summer.
Europe, Asia, Australia, the US and Latin America are all in the mix.
Max Well retweeted
Next on the block is Snap.
2027 Prediction: Bending Spoons signs agreement to acquire Snap Inc.
Watch this space.
BREAKING: we've officially entered a deal to acquire Miro for $1.355B! 😍