Post-training @OpenAI | PhD @StanfordAILab

San Francisco
Joined August 2017
💫 Astra is out Highlights: - Much smarter (evals are really saturating...) - Much more aligned & trustworthy - Better CUA (see this house Astra made in Blender) So proud of the team for training it! Still some issues: - Too much code slop - Astra asks for confirmation too often, which can feel lazy. Given how smart it is, we wanted it to be cautious, but likely overcorrected. We’ll fix those next! What else should we fix? Note: Astra follows instructions even more carefully than 5.6, including unwanted ones buried in old skills. I’d clean those up before using it. openai.com/index/gpt-6-astra…
82
58
19
1,045
155,112
Yann Dubois retweeted
We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because: 1. First time ever that OpenAI is #1 on Vending-Bench 2. The best model is no longer the unethical one.
115
371
146
5,279
724,187
Yann Dubois retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,731
20,164
14,767
120,582
74,534,486
Yann Dubois retweeted
since you guys loved the exploding tesla.. I used GPT-6 Astra to create a 3D website that pulls apart the male anatomy into 2,234 modeled pieces! we are in a renaissance of learning
1,153
5,429
1,013
53,415
7,874,246
Yann Dubois retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
970
2,527
1,088
15,259
7,583,867
We made a big push on decreasing hallucination & making Astra much more trustworthy! (Nit: to interpret those benchmarks you really need to control for the amount of claims the model make, which neither of those plots show.)
GPT 6 Astra finally fixed the hallucination problem. GPT 5.6 Sol was at 92%. It cancelled every one of my Stripe subscriptions during a migration. GPT 6 Astra is at 51%. That is the biggest hallucination improvement I have ever seen from any lab in a single release. And Fable 5.1 went the other way. Anthropic shipped a smarter model that lies more. The biggest weakness OpenAI models ever had is gone.
9
3
2
204
11,942
Yann Dubois retweeted
hey Astra, how fast are you really at using a computer?
Readers added context they thought people might want to know
The video shows Astra executing a pre-planned batched sequence of mouse clicks, keyboard presses and timed waits, not live real-time high-speed control. The poster clarified hybrid inputs and advance planning in replies. x.com/victornunez/st… x.com/victornunez/st… x.com/victornunez/st… openai.com/index/gpt-6-as… openai.com/index/computer…
318
685
292
10,002
1,945,601
Yann Dubois retweeted
Real-world results are in. There is a new #1 on Code Arena - GPT-6 Astra (Max)! It also reshapes the Pareto frontier as the best-performing model at $40/Mtoken, which matches the latest Claude model pricing. GPT-6 Astra by @OpenAI takes the top spot in Code Arena: WebDev with a score of 1797 pts. This opens up a solid +35pt lead over #2 Claude Fable 5.1 (Max) at 1762 pts and #3 Claude Opus 5 (Max) at 1688 pts. This is a significant improvement from GPT-5.6 Sol (xHigh) at +180 pts, ranked at #13. Category level votes still incoming, but already we see it at #1 in: Data & Analytics, Consumer Product, Content Creation Tools and #2 in Gaming and Simulations. Stay tuned for other categories like Brand & Marketing, Reference-Based Design and Full Stack rankings. Congrats to the @OpenAI team on this release!
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
176
398
148
4,195
803,270
Yann Dubois retweeted
GPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
88
107
68
1,293
853,421
Yann Dubois retweeted
We evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost.
146
491
125
5,166
900,814
Yann Dubois retweeted
GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.
65
323
83
2,211
638,350
Yann Dubois retweeted
We've updated FrontierCode 1.1 to reflect new discounts for GPT-5.6 Terra and GPT-5.6 Luna. With these new costs, the GPT-5.6 series sits on the pareto curve of price/performance efficiency.
47
70
37
1,375
419,893
Yann Dubois retweeted
We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
1,397
1,982
1,856
19,521
20,988,237
☀️Very proud of the team for training 5.6! A few of my highlights: - Great front-end aesthetics. E.g., I asked Sol to make a new blog for 5.6 with a celestial theme. - Much better CUA - Can work for much longer and is less lazy But it's not all great: - The writing quality has improved, but it's still bad in artifacts - Too many options with models/reasoning/fast mode/Cerebras/multi-agent We'll fix the above. What else should we improve? PS: belated post as there are so many other good models to train.
48
18
1
470
37,161
Tips when moving to 5.6 -Clean up old instructions. 5.6 follows them carefully, including undesirable ones buried in old skills. We had to clean many of those -Avoid Ultra if cost-constrained. Multi-agent improves evals+latency but costs much more. Or ask @thsottiaux for resets
14
10
2
215
20,963
Yann Dubois retweeted
BREAKING - OFFICIAL RESULTS: GPT-5.6 Sol by @OpenAI is 1st overall on Design Arena with an Elo of 1353. This puts GPT-5.6 Sol above Claude Fable 5 by @AnthropicAI and in the same performance band as GLM 5.2 by @Zai_org on frontend design. This is an 18-position and 60-point Elo leap from GPT-5.5. GPT-5.6 Sol also establishes a new Pareto frontier for preference vs. speed, faster than any model at this performance. Congratulations to the @OpenAI team on the launch!
109
183
102
1,919
521,827
Yann Dubois retweeted
The new ChatGPT Work comes with a new Computer Use experience. It's faster and introduces picture-in-picture, so you can keep an eye on Computer Use while it works!
54
44
35
849
197,643
Yann Dubois retweeted
GPT-5.6-Sol mogging Fable 5 on two of the most important coding evals
12
17
3
508
28,278
Yann Dubois retweeted
GPT-5.6 tops the DeepSWE leaderboard at 73%. Sol, Terra, and Luna results are now available.
71
168
88
2,037
346,837