applied AI & intelligence. notes on how we think, and how our machines learn.

San Francisco
Joined August 2025
You might have noticed that the multi-agent feature was also released in beta for the Responses API with GPT-5.6 models: developers.openai.com/api/do… It's now supported in our Agents SDKs as well! github.com/openai/openai-age…
7
3
31
15,466
GPT Live is a full duplex audio model launching today in ChatGPT -- it's coming to API soon, sign up here to be notified openai.com/form/gpt-live-1-i…
2
3
25
1,913
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes: 1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship 2. Builder: quickly turns a prototype/idea into production-grade product/infra 3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance 4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit 5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS. A healthy team needs a mix of these, depending on the product: - A product that is new and pre-PMF needs people that are strong at 1+2+3 - A product that is growing and has found PMF needs 2+3+4 and some 5 - A product that has strong PMF needs 3+4+5 and some 2 Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
884
2,445
685
20,068
3,151,794
A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run. The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky. lilianweng.github.io/posts/2…
70
595
74
4,685
474,160
PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, arxiv.org/abs/2509.26114)
PPO: rejected from NIPS 2017
14
109
7
1,280
214,174
PPO: rejected from NIPS 2017
38
198
43
2,601
532,976
Today we're launching Exa Agent: Opus/GPT 5.5 quality web research at 2-10x lower cost. It's our most powerful endpoint, particularly good at deep research and list-building - as cheaply as possible. Basically you can now use a deep research API for close to the cost of a search API. As the world wrestles with the cost of big models, we'll keep pushing the price/performance frontier 🫡
Introducing Exa Agent: frontier web research at less than half the cost of GPT 5.5 and Opus. /agent orchestrates a mixture of cost-effective models to complete any web research task, from simple data enrichments to building gigantic lists.
11
23
2
241
45,071
Hosting a private night (<30 people) with @steipete to hear about loops I'd like to get a few more builders to come - who's in? June 25, SF
8
4
27
3,541
Memories and Chronicle are a great way for Codex to learn how you work and move you from prompt engineering to vague prompting nitter.cf/OpenAIDevs/status/2046…
Last week, we released a preview of memories in Codex. Today, we’re expanding the experiment with Chronicle, which improves memories using recent screen context. Now, Codex can help with what you’ve been working on without you restating context.
2
1
18
3,162
If you’re getting started with Codex Mobile, @Dimillian’s guide is worth a read! Your phone isn’t a tiny terminal. It’s a control center for directing and reviewing work running on your laptop or in the cloud. Thomas helped build the experience, and it shows.
17
24
291
61,161
Awesome to see this innovation in text diffusion. DiffusionGemma is lightning fast, 4x faster than other Gemma 4 models! Congrats to @bodonoghue85 and the team who worked so hard on this - excited to see what people build with it!
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
103
117
17
1,640
213,555
We are releasing River API, our first product, in early access. The API gives you access to the same battle-tested tools that we’re using internally at River for post-training, reinforcement learning and continual learning. Check it out and let us know what you think!
Introducing River API. Fine-tune and RL train leading open-source models at scale, ranging from 35B to 1T params. We’ve been using it internally to power our research and we love it. Today, we are opening up our public waitlist. Own your intelligence!
24
51
2
530
137,396
Training a generator with PROPEL roughly doubles the yield of goldilocks training tasks (not too easy, not too hard) across three domains: 2.4x on code induction, 1.7x on math, and 2.0x on a 27B SWE agent.
1
2
13
1,033
*​When you think of AI, think of humanity too* ​This was the main message of the GStar Summit on AI & Humanity that I organized two days ago (as a breather after Google I/O). For a long time, I was deeply focused on advancing model capabilities toward superhuman reasoning, and I found immense joy when we accomplished our goal, e.g., our IMO-gold milestone in July 2025. I always assumed that someone, somewhere, would just take care of the social impact side for me. ​It really hit me in Feb 2026 when our math research agent, Aletheia, solved FirstProof problem #7 by flawlessly utilizing heavy mathematical machinery, as noted by our surprised mathematicians. I started to wonder: what's left for humanity? ​Discussing human values with my wife, Wendy Nguyen, and our dear friend Jean DeSombre, I realized I hadn't been thinking much about humanity while actively charting the course of frontier AI. But as our conversations continued over the past few months, I noticed myself becoming more "humanity-aware." ​For example, during a team lunch at Google DeepMind, we were chatting about what we would do after we "solved" AI for math. I heard a number of suggestions, e.g., robotics or world models, but nothing explicitly centered on humanity. I told the team that we need to think about empowering humans: one way is to build a model capable of asking questions, from basic paraphrasing all the way to forming conjectures, which would be incredibly useful for both education and model training! ​At GStar Summit 2026, aside from the deep dives into agentic AI with @edchi, @YiTayML, @preslav_nakov, Phong Nguyen, Myungsub Choi, Noriyuki Kojima, and Hung Bui, we dedicated half of the summit to talking about humanity! ​@PoShenLoh, awesome as always, proposed the "Thought+Full" philosophy and reminded us how people find joy in helping others. Jay Kim shared with us new perspectives about human health in space, showing that space is not that scary to be a part of. Together with other speakers and panelist, Wendy Nguyen, Jean DeSombre, Marc Woo, Laurent El Ghaoui, Tuong Nguyen, Tuoc Huynh, Tuan Cao, and @CurtisSChin, we discussed all aspects of humanity in the age of AI. ​Our hope is that everyone who walked out of the summit can now find joy in talking and working with each other on human values (either before or after AI)! ​Thanks to everyone for coming from all over Asia Pacific and Silicon Valley! Many people told me that it's very rare in Vietnam for people to stay from 8am - 6pm with a fully-packed hall of 1,000+ people! See you again soon!
My second time as a speaker at gstar summit in Vietnam. This time I am honoured to be chilling (I mean sitting) on a panel with colleagues @lmthang and @edchi 😀. I really enjoyed the vibes and company and meeting new people. 🫡 Events by @lmthang and @newturing have always not only been super impressive in terms of speaker calibre density but also sota in organization and hospitality. 😆 These two days have been a much needed retreat for me and it's been really timely! 😀 Also enjoyed hanging out with @edchi, I wanted to try to get a photo with @lmthang but it was impossible to catch this Vietnamese national hero for a selfie. 😀
5
5
1
53
12,237
After AlphaGo, the skill of human Go players noticeably improved. I suspect we will see a similar pattern in math.
Another major problem, this time in additive combinatorics, has fallen, this time to humans rather than AI, but using methods related to the AI solution to the unit distance conjecture.
183
950
204
8,953
810,120
This is a general-purpose LLM. It wasn’t targeted at this problem or even at mathematics. Also, it’s not a scaffold. We have not pushed this model to the limit on open problems. Our focus is to get it out quickly so that everyone can use it for themselves.
40
70
41
996
250,324
Viraj retweeted
Codex is feeling codexy
320
41
28
2,204
154,983
Checking out Gemini 3.5 Flash, available today, which helped power Antigravity 2.0, Gemini Spark, and many more! blog.google/innovation-and-a…
1/ Today at #GoogleIO, we’re releasing Gemini 3.5, our latest family of models combining frontier intelligence with action. We’re starting by releasing 3.5 Flash, which is built to help you execute complex, long-horizon agentic workflows. Gemini 3.5 Flash is our strongest model for coding and agent yet.It outscores 3.1 Pro on agentic and coding benchmarks like Terminal-Bench and MCP Atlas, while running 4x faster than other frontier models. Used in Google Antigravity, 3.5 Flash is even further optimized to be up to 12x faster. It’s a powerful engine to deploy sub-agents that collaborate, run high-frequency iterative loops, and solve real-world problems at scale. Some highlights we’re excited about 🔽
2
2
1
27
3,148
Viraj retweeted
1/ Today at #GoogleIO, we’re releasing Gemini 3.5, our latest family of models combining frontier intelligence with action. We’re starting by releasing 3.5 Flash, which is built to help you execute complex, long-horizon agentic workflows. Gemini 3.5 Flash is our strongest model for coding and agent yet.It outscores 3.1 Pro on agentic and coding benchmarks like Terminal-Bench and MCP Atlas, while running 4x faster than other frontier models. Used in Google Antigravity, 3.5 Flash is even further optimized to be up to 12x faster. It’s a powerful engine to deploy sub-agents that collaborate, run high-frequency iterative loops, and solve real-world problems at scale. Some highlights we’re excited about 🔽
84
190
38
1,484
142,422
Confirmed that Andrej is joining Nick Joseph’s pretraining team at Anthropic and is starting a team that will use Claude itself for research
Personal update: I've joined Anthropic. I think the next few years at the frontier of LLMs will be especially formative. I am very excited to join the team here and get back to R&D. I remain deeply passionate about education and plan to resume my work on it in time.
12
13
8
709
95,764