@npew

Devices @OpenAI

San Francisco, CA
Joined April 2010
Peter Welinder retweeted
GPT-6 Astra turned my room into Studio Ghibli, on Apple Vision Pro 🤯 It constructs virtual objects located and sized exactly as real objects. Then they can be styled however we want, in REAL TIME. I'll share later how this all works. For now, I just want to say, wow.
42
80
19
945
73,711
Peter Welinder retweeted
We’ve fixed a bug that was degrading image understanding in GPT-6 Sol and GPT-6 Luna. You should now see better results on visual tasks in the API and Codex, including computer use.
259
305
186
5,999
679,912
Peter Welinder retweeted
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵
52
27
12
335
144,520
Wow, I remember having my car broken multiple times pre-2020, and it just seemed to get worse. If this trend continues, maybe we can start removing all the signs warning tourists to not leave anything in their car.
Remember car break-ins? San Francisco is on pace to end 2026 with a small fraction of what *used* to be our City’s most pervasive crime. Automated License Plate Readers (ALPRs) have been a game-changer in ending that. Yet, incredibly, some activists want ALPRs gone… (1/2)
2
1
11
3,639
You can now do almost anything with voice. Our voice team really cooked!
We heard you loud and clear. ChatGPT Voice can now: - Use plugins like your email, calendar, and Slack. - Be powered by GPT-6 Astra, Sol, and Luna. - Be used in ChatGPT Work on web and mobile, so you can create docs, decks, sites, and spreadsheets or tackle complex tasks in the browser, just by talking. Rolling out globally today in the latest version of the app.
5
6
2
62
6,177
Peter Welinder retweeted
The race for AGI Script: Sherpa by Pocket FM Video: Seedance 2.5
🤖 Made with AI
838
3,437
1,025
21,954
4,076,643
Astra crushing it again. Luna is a little beast.
With the latest results from the @rails agent evals, it's clear to see that @openai still has a solid lead, despite Opus 5.5 making a good jump. The dark horse here remains Luna Max. 18% completion at just $11! Not that far off GPT-6 Sol!
2
83
7,689
I guess that explains why I'm always confused about this 🇸🇪🫠
Whatever you got used to while growing up I guess...
2
10
3,282
Peter Welinder retweeted
AI now appears to be visible in increased productivity growth rates.
U.S. labor productivity, 2013–2026: > Unremarkable growth for most of the 2010s > Breaks sharply higher starting 2020 > Now running 2.2% above where the old trend said we'd be The last time this happened was 1995–2004, when the internet added close to 3% a year for a full decade. Looks like we're a few years into the sequel.
14
23
1
284
31,964
Peter Welinder retweeted
Astra is an incredible bargin compared to Fable! And look at Luna on max too!! @openai's return to the top is something else. Maybe this is why Anthropic finally agreed to do AGENTS.md? 😄
Agents on Rails: You asked, so we turned every model in Agents on Rails up to its max effort level. The result: more effort/reasoning doesn’t always mean better results. @OpenAI's models made the biggest gains, costs nearly doubled overall...and the newest agent in the benchmark, DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score. What an entry. Here’s what we learned and what max effort gets you with each model: rubyonrails.org/2026/9/21/ag…
76
60
11
1,516
171,240
RT @rapha_gl: in the future, all historical research breakthroughs will be announced via ChatGPT sites
This quoted post is unavailable.
1
296
Peter Welinder retweeted
Two days ago, GPT-6 Astra broke a yet unsolved German Army Enigma message from 1941. Amazingly Astra was able to autonomously: - Search historical archives - Compare uncertain letters - Find contextual clues - Build an Enigma simulator - Write cryptanalysis code - Run parallel experiments - Test competing keys - Recover the plaintext - Cross-check the results 1/n
131
460
139
3,956
1,091,061
Funny if the GPT moment for robotics is just another GPT.
We put GPT-6 Astra in the RoboDojo. 🥋🤖 The RoboDojo Team conducted a comprehensive evaluation of GPT-6 Astra as an embodied agent, including: • RoboDojo Sim & Real, compared with GPT-5.5 and DeepSeek-Flash • Humanoid high-level control • Dexterous piano playing with RoboPianist 🎹 • A systematic study of in-context learning (ICL) Our key takeaway: GPT-6 Astra demonstrates remarkably strong semantic and spatial understanding, together with impressive in-context adaptation. At the same time, physical commonsense remains a clear bottleneck — revealing an important gap between understanding the world and truly reasoning about its physics. Full report & demos: robodojo-benchmark.com/repor… @_wenbozhang (project lead), @wenhaocha1, @frankzydou, @JinWeiyang18434, @YutaoOuyang, @minifullcapsule, @x_h_ucb, @YutaoOuyang, @YueChen614
13
9
1
312
21,211
Peter Welinder retweeted
What’s the wildest thing you’ve built in 3D with GPT-6 Astra? I’m teaming up with @OpenAIDevs to see what you got. RT and drop a link, render or video in the replies. You’ve got 24 hours.
I fed GPT-6 Astra 9 crappy photos of my studio and it was able to understand the spatial arrangement, stitch them together and create a fully interactive 3D model of the space. Absurd. I'll link it below so you can check it out for yourself:
152
47
12
398
333,570
Peter Welinder retweeted
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
118
142
51
2,306
998,589
Astra is a fun model!
Extraordinary share gains for OpenAI vs. Anthropic over the last two months. Per Openrouter, OpenAI has gone from 20% share to 50% share vs. Anthropic (meaning Anthropic has gone from 80% to 50%) since June.
1
58
5,876
Look at that token efficiency!
it's not looking good Anthropic bros
4
1
112
6,627
Future with robots is looking to be quite joyful!
During Japan Mobility Show 2025, Toyota has unveiled ‘walk me,’ a concept autonomous wheelchair with foldable tentacle legs that can climb stairs and sit on the floor.
5
1
29
7,999
Peter Welinder retweeted
This will happen to Inference
One data point in technology making everyone richer: the cost of lighting has decreased by over 1000x.
29
18
5
371
80,034
Peter Welinder retweeted
⚡️ 2.4x more ChatGPT Voice in Desktop We've dropped prices by ~60% for voice in Codex and Work in the desktop app, giving you more time to orchestrate tasks and even more tokens for real work.
97
72
20
1,551
124,195