Pinned Tweet
How good are VLMs at geological reasoning? Can we teach to reason where no labels can be established / training examples are hard to create?
We create a stratigraphic diagram exam environment and train VLMs on it. - then show they seem to transfer to seismic domain.
Gonna leave this here: oneusefulthing.org/p/centaur… jagged can also mean “very jagged”
OpenAI confirmed to the New York Times that they have made "substantial progress" on another Millennium Prize problem in the last five days, and are preparing to announce.
The rumors for the last 48 hours have been OpenAI solved the Hodge Conjecture, and that Anthropic has solved the Birch and Swinnerton-Dyer Conjecture. Since Navier-Stokes rumors abound, so I was reluctant to post about either. However, OpenAI's statement to the NYT now gives the Hodge rumors some very serious support.
In general people have not updated yet that the new unnamed OpenAI model, the one that finished training about two weeks ago, which I believe will be named Aeon, is massively better at math than Astra, which two weeks ago was the best in the world. Aeon solved Navier-Stokes in 88 hours, start to finish. Follow the trend line. That means everything is on the table. Literally everything. And this does not end with math. Please update. We are taking off.
What surprises me the most is that people dont read their work contracts. I am willing to bet that anthropic claims all IP while employed. Doesnt matter if you do it in your spare time. Such a thing often doesnt exist especially when its somehow related to your day job.
drama summary for those confused:
- Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems
- they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help lead the way there
- Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models. Tristan is not related to Anthropic.
- Early Sep: Rumor spreads to OpenAI that Anthropic has solved a major problem. Tristan emails OpenAI to clarify. without revealing the problem they solved or how they did it.
- After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
- Sep 6th: OpenAI's Sebastien Bubeck tells Tristan that they solved the $1,000,000 Millenium Prize Navier Stokes problem. The approach is very similar to Tristan & Levent's approach to the non-Millenium problem.
- Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
- OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but only if they remove Levent as an author, as he works for Anthropic.
- Sep 8th: Tristan refuses to remove Levent, and rushes to publish their results independently.
Currently unclear is whether Anthropic had a separate solution for the $1,000,000 problem, or whether the rumor was about Tristan & Levent's independent research.
Copyright lawyers rn
it may soon be cheaper to rebuild/clone games than to buy them at the store...
it cost about $10 of compute each to knock off Super Smash Bros, as well as mario kart and a few others
built with Astra ofc
this isn't full game quality and no online vs. mode or all the maps, but we're also just getting started.
wonder how this impacts gaming companies since indie devs will be able to compete much easier, and it all comes down to execution
Maybe the lobsters won’t be first to go
For the first time, scientists have mapped the complete brain and central nervous system of an adult male fruit fly — a key model organism in science. 🪰
Working alongside HHMI Janelia Research Campus and the scientific community, @GoogleResearch scientists and researchers used AI to combine millions of 2D images into 3D neural shapes, reconstructing a record-breaking 166,000+ neurons. This foundational map of the adult male fruit fly brain can help accelerate our understanding of the brain, and is a major milestone in neuroscience.
Lukas Mosser retweeted
Google have, dare I say... RELEASED A GOOD MODEL 🥹
Gemini 3.8 Flash provides ~Opus 5 performance at much lower cost, while being super fast
Cycle 1: we need to fine tune our models!
Cycle 2: fine tuning is worthless, the frontier is too good!
We are here:
I have a feeling we should be looking at more HR and social sciences type literature on how to train employees and build good teams rather than how to build good software.
Two things needed in the MCP debate: observability to the outside e.g. we need tracing coming from what’s inside the mcp especially if it’s an agent behind a tool. Monetising skillsets, oh wait, that’s what I pay for docs anyway. So actually it’s just one thing.
I think this might end up being the worst integration pattern and might end up having similar poor social effects as social media.
I don’t want to talk to an LLM when I send someone a message.
Everything complicated gets built by software engineers at the frontier first. Then, over time, all that complexity disappears and becomes a one-click feature anyone can use, with zero technical background.
This is a perfect example. Yesterday, people were building complex agents that connected to email, SMS, calendars and other apps. Today, every consumer LLM will simply do all of this for you. Tomorrow it will do more including making calls on your behalf (this will be a killer feature)
What looked like a complex agent system yesterday will soon just feel like a normal feature... My friends often feel left behind like "omg I don't know how to run Openclaw on a MacMini." But friends, you don't need that! Everything will be put as a service wrapped in a few dollars a month subscription.