@olcan

Engineer @GoogleDeepMind. Prev. Product @ GDM, Founder/CEO @ Scaled Inference, Engineer @Google (Search, Research, X, Brain).

Palo Alto, CA
Joined June 2008
ai will roofline all software
The asm version of ttfx continues to improve. The Rust version had been through many improvements, but even hours of frontier intelligence was starting to yield very modest improvements. The leaps going to asm has been huge. Now 8x mean on an Intel 135U! 🤩
1
300
open jev
Introducing Bespoke Nimble: an open data, open model, open recipe for an open Jev. Code and info: github.com/bespokelabsai/nim… Model: huggingface.co/bespokelabs/B… Data: * A new data curation recipe called contrastive data curation. * Slightly change facts to generate negative data. This pushes the model to discriminate better and become a better decision maker. The calibration is implicit. * Didn't do ablations but I think this is a critical piece! * This also means training data doesn't need probabilities. * Data covered 10 categories, and is fully synthetic. * This data is split into train and eval. Training * LoRA finetune of Qwen3.5-9B. * Distillation-free: we use Jev to only evaluate. * No RL yet! Serving * Parallel constrained decoding as suggested by @NielsRogge and @harshagundal. Results: * The post-trained Qwen (Nimble) became substantially better on our curated eval: 66% for Qwen to 90% for Nimble. Jev is at 93%. * 100ms on H100 and free to use on your macbook! Feel the AGI for free. * 2 days of building in public. :) Big caveat is that there is no standard benchmark to measure performance, and it's possible Nimble is much worse on other benchmarks compared to Jev. But it should be better than Qwen! We thank @typesafeai for making Jev and the inspiring discussions in the community. Hope this release lifts all the boats and encourages more research and activity in this space.
4
72
10,153
Olcan retweeted
A lot of otherwise smart people on Twitter seem 100% convinced AI risks are all fake and stupid and part of some marketing ploy. What is surprising is that some of these people seemingly also believe that AI’s positive uses are on an incredible trajectory of increasing capability with no end in sight. VC Twitter is particularly infected by this pattern. It’s not really coherent. Most positive use cases for AI have a corresponding “dark version”. If you are super human at coding, you are also super human at hacking. If you are superhuman at structural engineering you are likely superhuman at finding structural flaws to knock buildings down. If you are superhuman at designing drugs, you are superhuman at designing novel undetectable poisons. If you can cure viruses, you can create them. Some of these “dark versions” are not so bad, and some are actually pretty scary. Either way these are real societal and technical problems that need to be solved to get the good stuff and avoid the bad stuff. We’re experiencing the first of these with coding and computer security which is the most advanced, but that won’t be the last. I think we’ll be able to solve these problems, but they aren’t solved yet and if you believe in continued AI progress they are surely coming. But “bad people using AI” is not the only problem. Uncontrolled AI autonomously doing bad things, despite sounding kind of nutty, is also something we should be concerned about. AI “killing us all” is not the most likely outcome, but the chance of a major civilization-wide catastrophe doesn’t have to be very high for it to be a concern. Again, this is only a problem if capabilities advance to a point that AI can do really crazy things on their own, which hasn’t happened yet, but I think the Hugging Face incident is a good example of the outline of how things can go wrong when capabilities outpace alignment. We should be glad that the only available bad thing right now is hacking, which isn’t all that bad. It’s clear that as you get to superhuman capabilities you need a level of alignment and control that is correspondingly superhuman. Humans have plenty of misalignment problems themselves (serial killers, mass shooters, tyrants, etc), but it’s a manageable problem because most humans can’t do that much damage and we’ve developed systems to prevent dangerous humans from getting too much power. Talking about these issues is just common sense. It’s not a sign of some kind of neuroticism or pessimism. These are just hard problems that it’s very important to solve for AI to have a positive impact. I’m pretty confident we will solve them. But we haven’t solved them yet, and to my mind we are clearly on a trajectory of rapidly increasing capabilities which means this is important. Mocking people who are worried about this or talking about it, without anything substantive to say about how we can be sure these problems won’t arise, is not really a very helpful contribution.
166
128
32
737
104,006
qualitatively, gemini 3.8 flash is a big improvement over 3.7 imo
4
3
1
156
8,578
I once said to my former team: "I care less about what is the smartest thing our model can do than about whats the most stupid thing that it cannot do" robustness is still a limitation to increased automation
too many people working on making models smarter when they should be working on making models less stupid
45
74
17
1,466
248,838
the year of (agents on) linux on the desktop
Omarchy is blowing up. I've never been involved with anything in my career that has grown this fast. With Ruby on Rails, we had years to build solid institutions, teams, and relationships. With Omarchy, we've been forced to figure it all out in twelve days. It's exhausting, but also incredibly exciting. It was never sustainable with just Ryan and me running everything. That's how it was, more or less, up until Quattro. Lots of other contributors, but all the responsibility was on us to make sure the ship stayed afloat, the servers didn't crash, and fixes got pushed out quickly. Now it's time to build a proper institution. Durable, resilient, and competent. That's what we're doing here on Basecamp now. I'm spinning up teams for every facet of responsible distro management, and I'm getting an absolute outpouring of interest for all of them. Everyone wants to be part of this. We're winning hearts, minds, and volunteers at an astounding rate. Great! We need all of it to succeed. Because make no mistake: There are many people who'd love to see this rocket blow up before it reaches the moon. Aggrieved Linux users who don't like the sudden attention their exclusive hobby has received. Competing Linux distributions that are seeing our numbers explode. Mac stans who've sunk their identity into an apple. And, of course, any of the haters I've picked up in my quarter-century career speaking bluntly on the internet. They're not going to succeed. Because we've already become unstoppable. There's too much momentum, too much money, and too much support now backing this effort. We're living the Mandate From Heaven meme at the moment. And we are here to fulfill the prophecy: The Year of Linux on the Desktop! That has been a joke for two decades. But by the end of the year, nobody at Apple or Microsoft is going to be laughing. They're going to be scrambling. Because neither of these proud organizations currently has any method to counter the speed, vision, or ambition with which we're going to accelerate into the future of personal computing. This is the moment. This is the opening. This is our chance. For thirty years, we've been subject to one OS overlord or another. Dictating how we compute. Choking off competitors through platform malfeasance. Tollboothing the distribution. That ends now. Because Linux is going to win. And Linux is free. As in beer, speech, and source. But just because it's inevitable doesn't mean it's going to be easy. We have a lot of work in front of us if we actually want to make our mark. But there's never been a better time for this kind of delusional ambition. The age of agents is the unlocking factor. It sounds like a LinkedIn slogan, but it's true. Where the application of tokens goes, the innovation follows. We can fix everything. Let's do it together. Let's go. --- This is what I sent to the dozens of new volunteers who've signed up for teams within the new Omarchy organization yesterday. But we might as well broadcast our mission and intentions to the world too.
1
3
357
fairness aside, the "Complexity" bit here is hilarious 🤣
I tested Perplexity Computer. I've never laughed this hard in my entire life.
1
2
442
Today’s AI needs to be able to deal with uncertainty! This has been a central theme of my work, so I was delighted to talk about it at length with @FryRsquared for the latest @GoogleDeepMind Podcast. Intelligence means knowing the limits of one's own knowledge, and when to say you are unsure about something. If we want useful AI systems that we can trust to make important decisions, then they need to be able to do the same. They should recognise blind spots, avoid overconfidence, and reason carefully about their decisions. I think you’ll enjoy the episode, but I can’t say for sure. 🙂 You can watch it here: goo.gle/4y6WNlR
7
17
5
131
26,968
Olcan retweeted
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
236
890
276
6,334
1,611,146
Agree, agents LOVE (= are most trained on) Linux and other OSS and this is very bullish for Omarchy and any other open-source products that fully embrace AI.
I don't have to agree with DHH's political views to appreciate the energy he puts into this. I hope it also inspires others! (Agents love Linux, it's a real game changer) nitter.cf/dhh/status/20901686868…
1
368
3.7 flash is not a minor upgrade
Gemini 3.7 Flash just took #1 on @ArtificialAnlys new AA-AnalystAgent. AA-AnalystAgent evaluates against 80 real-world quantitative analysis tasks across 14 business and scientific domains (finance, healthcare, hydrology, government appropriations). The Agent run inside an isolated Python 3.12 sandbox using AA's open-source Stirrup harness. They receive reference spreadsheets (.xlsx) and documents (.docx) alongside standard data libraries (pandas, polars, openpyxl, scipy, PyMuPDF) to inspect schemas, write scripts, handle edge cases, and calculate final figures. Gemini 3.7 Flash: • Accuracy: #1 with 60.0% pass^5 (70.5% pass@1, 77.5% pass@5) • Speed: 1.32s per task (fastest) • Cost: $0.54 avg per task (middle)
3
3
173
10,483
Olcan retweeted
In this (extremely long) video, I'll build a distributed training framework from first principles using PyTorch. We'll derive everything from first principles, explaining the mathematics behind distributed training as well as every auxiliary component we build along the way, including Multi-head Latent Attention (MLA), Rotary Position Embeddings (RoPE) and YaRN. I'll explain process groups, ranks, collective communication operations, distributed autograd, device meshes and distributed tensors, and then use these foundations to implement the major forms of parallelism used to train modern language models. We'll derive and code pipeline parallelism, data parallelism, FSDP, HSDP, tensor parallelism and context parallelism, including pipeline schedules, block matrix multiplication, model sharding and ring attention. We'll then combine these techniques and examine how tensor layouts and gradients must be tracked as data moves between devices. Most importantly, we won't study each parallelism technique in isolation. We'll combine pipeline, data, tensor, context and expert parallelism into a single working framework and follow the tensors, communication operations and gradients as they move across GPUs. By the end, you'll understand not only how these techniques work individually, but how modern distributed training systems compose them to train large dense and mixture-of-experts models. We'll also build the important components of a modern transformer, including rotary embeddings, long-context extensions, multi-head latent attention and a mixture-of-experts architecture. This leads into expert parallelism, expert tensor parallelism, token dispatching, all-to-all communication and combining different parallel dimensions through dense and sparse device meshes. We'll also build the supporting training infrastructure, including the training loop, optimizer, learning-rate schedule, checkpointing and metrics. The goal is to give you the foundations required to understand how distributed training frameworks work internally and how to design and implement new parallelization strategies yourself. No prior knowledge of distributed training is required. Link to the video: youtube.com/watch?v=XoGvCBRn…
87
268
50
2,385
315,561
hear me out
3
1
1
50
2,267
I asked Gemini 3.7 Flash to write prompts for modern web UI video backgrounds (great challenge to see the creativity of any model) and I really love the resulting aeshetics, rendered with Gemini Omni Flash. Some of the prompts I liked best below - which one is your favourite?
31
49
9
767
71,496
let that sink in
120
250
97
2,773
747,373
Gemini 3.7 Flash just saturated one of our physical tool-use benchmarks at 92%. Gemini 3.6 Flash, released three weeks earlier, scored only 32%. This marks a step change in the robotics capabilities of LLMs. 🧵
18
59
15
640
112,070
I have decided to leave math academia. Not because of burnout, the academic job market, or a loss of love for mathematics. Surprisingly, it is because of how much mathematics I could do with LLMs over the last few months. For most of my life, pursuing mathematics felt like the obvious choice. It gave me a sense of purpose, rooted in the search for hidden truths in God’s Book. Over the past few months, that conviction has been shaken. With LLMs, I could get several breakthroughs on problems I care deeply about and had worked on for years during my PhD. This progress might have taken years of my academic life if not for LLMs. This could have strengthened my conviction that mathematics was my calling. Instead, it undermined it. With each passing week, my role in the discovery loop seemed to grow smaller. What made mathematics meaningful to me was never just the final answer. It was the struggle: months or years of exploration, failed approaches, and deep thought before a hidden structure finally became visible. That struggle gave me a sense that the result was truly mine. Prompting my way to answers, without the struggle that once defined the process, no longer felt like the same vocation. The answer might be just as beautiful, but it did not feel earned in the same way. I now believe that we are rapidly approaching a world in which most answers from God’s Book will be only a prompt away. Once I truly internalized that possibility, dedicating much of my life to finding those answers slightly earlier no longer felt as meaningful as it once had. But one problem kept bothering me: How do we know the oracle is right? Intelligence (and hence the amount of math papers) is becoming abundant and cheaper by the day. Trust is not. The proofs produced by these systems can be highly sophisticated, and their errors can be extremely subtle. Determining whether an apparent breakthrough was actually correct sometimes took me days. Until it is verified, a beautiful proof is not different from slop. That gap drew me toward formal verification and eventually toward formalizing math proofs in Lean, including proofs of several longstanding conjectures and Erdős problems discovered with the help of LLMs. I began to see verification as a central intellectual bottleneck. Intelligence is useful only when its outputs can be trusted. So I am leaving math academia. I am shifting my attention from discovering answers to building systems that can certify them. I am very excited to be joining @PramaanaLabs to work on this challenge. Over the next few years, I hope to help build a future in which increasingly powerful systems are not merely intelligent, but provably correct and reliable across problems far beyond mathematics.
216
628
173
4,681
563,042
Played around with Qwen 3.8 27B and DeepSeek 4 Flash on my mac over the weekend. I'd say both models represent a significant jump in quality and are very usable at 20-30 tps even for longer context (i tried up to 100k or so).
4
1
9
1,150
I find the "inception" pattern to be very useful in many agentic use cases. You can force the model to take an action when it thinks for too long by injecting a thought after a specified reasoning budget. Helps dealing with underspecified tasks which make the model reason for way too long.
Replying to @ggerganov
to limit the max reasoning length, add: ... \ --reasoning-budget 4096 \ --reasoning-budget-message "... I am thinking for too -- let me gather more info about the task." adjust to your needs
28
37
8
486
111,473