@bleysg

Helping people engineer the future. The core metric is task/GJ (gigajoule) and GJ/humanity.

SF / LA
Joined March 2009
I’ve been building an interactive model of the AI transition for a while, and wasn’t quite planning to make it public yet. AI 2040 is making the rounds today. There’s enough overlap, and enough difference, that this feels like the right moment. enia.cc
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
3
1
25
10,270
When I posted this I wasn't aware of Fujitsu-Monaka, but that chip is exactly following the principles I laid out here.
How do you make a DGX Spark that can run a 3T class model like Kimi K3 at good speed on one box? I break down how this is achievable before 2030.
3
605
Prediction: This diagram I made today is going to be the most important thing to understand in AI PCs and AI SMB servers for the next few years. No hardware yet follows this pattern. It's a natural evolution of what hybrid bonding gives to the PC. 🧵👇 nitter.cf/NVIDIAAIInfra/status/2…
Announcing the expansion of NVIDIA NVLink Fusion with NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs. Amazon's @AnnapurnaLabs will be the first to work with us on NVHBM, combining @awscloud custom silicon with our memory technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads. Learn how we're helping hyperscalers and AI innovators build the next generation of AI infrastructure: nvda.ws/4xpBtYN
3
1
9
4,026
Hybrid bonding makes the architecture practical: an N2/A16 compute tile above a larger mature-node platform die with shared cache, LPDDR6 PHYs, I/O, etc. The pieces are in @AIatAMD @intel datacenter parts already. What's novel is the unified-memory PC/SMB version.
3
226
Similar in concept, but that’s slower.
122
How do you make a DGX Spark that can run a 3T class model like Kimi K3 at good speed on one box? I break down how this is achievable before 2030.
Prediction: This diagram I made today is going to be the most important thing to understand in AI PCs and AI SMB servers for the next few years. No hardware yet follows this pattern. It's a natural evolution of what hybrid bonding gives to the PC. 🧵👇 nitter.cf/NVIDIAAIInfra/status/2…
1
18
3,183
Arithmetic on a ~900 mm² N6 platform die: 96 LPDDR6 channels (2,304-bit) ≈ 4.15 TB/s, 1.5–2 TB of DRAM, compute-die tax ≈ 0. Next act: add bonded DRAM... 20–40 GB per full-footprint layer, 64–160 GB at 10 TB/s. @Xiaomi fast/slow AI Cube architecture, scaled up. nitter.cf/ItsmeAjayKV/status/209…
Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀 - 3 custom chips: Xring O3, O100, D100 - 200 TOPS NPU - 1.22 TB/s AI memory bandwidth - Up to 160GB unified memory - 150W sustained power - 120B models running locally Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥 Once it hit's the marked, going to sell like hot cakes.
108
Prediction: M7 Ultra gets native FP4 matmul acceleration. M5 Ultra has the bandwidth-cutting primitives which help decode for FP4 and FP8, but still computes in FP16. Getting Blackwell style FP4 compute gains that multiply prefill perf will allow Apple to finally catch up on the whole picture of AI perf/W.
7
8
1
110
10,016
M5 Ultra has me excited for next year. 1.2 TB/s bodes well for next year's Rubin DGX Spark release and for M7 Ultra. Both will ship between late 2027 and early 2028, both will move to next-gen LPDDR6 memory. Both should move to 2 TB/s territory. That's faster than RTX PRO 6000! M7 Ultra will likely go up to 1 TB per box too. That would allow you to run full sized Kimi-K3 class frontier models fast on just two boxes. At this rate, frontier class models should run on just one box by the time of the Feynman DGX Spark generation late 2029.
2
2
1
33
6,020
M5 Ultra just boxed in my projection of "When do we get frontier models at home?" Why? Because they're copying the DGX Spark model: Dare you to have just one. 512GB M5 Ultra is $17.7k in Oct. It will take 4x to run Kimi K3. Apple is supporting this arrangement. This will be the most economical way to run Kimi K3 "at home" ... $71K. That's not really an "at home" price for most of us, but before that the most economical option was looking like it would be 8x B300 in a DGX B300 for around $550k. That's a rack-mount beast needing power and cooling consideration, not a cool and quiet desktop box that can plug into any average circuit. The costs and practicalities made it a real consideration only for SMBs with serious on-prem requirements. A 4x M5 Ultra cluster can sit anywhere and give practically fast frontier level model performance, before the end of this year!
The 512GB M5 Ultra will ship in October with a massive 1.2TB/s memory bandwidth. 50% more memory bandwidth than M3 Ultra, 4.4x more than DGX Spark and M4 Pro. Stacking 4 x M5 Ultra, I expect you'll be able to run Kimi K3 / GLM 5.3 faster than the API (>100 tok/sec).
40
17
3
281
58,947
Interesting charts I've come across recently on AI buildout, progress, and token usage.
1
3
551
CEO of Cerebras wants to land fabs without red tape in the US as a sovereign strategic priority. This is becoming a hot button issue. Local communities don’t want the noise, etc. He’s not pointing at that though, he’s pointing at the slowdowns that happen as a result of having to conform to other practical local norms like fire codes intended for strip malls. My current thesis is that we see a major compute crunch from demand outpacing global compute supply before 2028. Unlocking more power and compute, and getting obstacles out of the way now to scale buildouts at Terafab level ambitions are probably the only way we get a world where most compute happens in the US and generally a world where our collective ambitions aren’t heavily bottlenecked by a compute crunch for years.
There’s a simple solution to the current memory and compute shortages. It comes down to one policy decision: building fabs. I suggest giving TSMC, Samsung, SK Hynix, Micron, Global Foundries, Intel a 20-year window to build fabs in specific areas. Free from every local ordinance. That sounds extreme. It isn't. It’s a national, strategic imperative. Fabs are modern pyramids. They are the greatest things humans make, and the chips they produce are a foundation of our economy. Yet we have to build them around the same local building codes written for strip malls and parking lots. Fabs are the size of 50 football fields - it’s a red tape nightmare. One multibillion dollar fab in Texas and had to be completely redesigned over a single local fire ordinance. It set them back the better part of a year. In the US, we've lost the ability to quickly build the infrastructure we obviously need. Cutting-edge fabs included. And we didn't just lose the fabs. We lost the entire surrounding ecosystem. The packaging expertise. The strategic jobs. The deep manufacturing know-how. It is extraordinarily important that we get this full ecosystem back. Lets get to the root of the problem: permitting, local ordinances, zoning, the whole regulatory patchwork. Fabs are a strategic asset. We have to get them built in America.
1
2
928
I haven't slowed down in the two weeks since my last post. Entrpi/ds4 just hit v0.6.2 and is now at a great robustness and scaling milestone. 3M active tokens across many parallel agents is running no sweat under stress on just one DGX Spark today. vLLM and SGLang were designed around dedicated VRAM. I designed around the Spark in particular, where the engine shares the box with other apps and models. I want you to be able to run an STT model alongside without worrying about OOM and without compromising capacity more than absolutely necessary. To achieve that, the engine measures what each request actually uses, charges exactly that, sizes its budgets from what the box really has, and hands idle memory back. Context is demand-mapped: nearly free until a conversation fills it. When something does not fit, it reclaims idle state, then refuses with a reason instead of OOMing the machine. While idle it audits its own ledger and logs any drift. The one decision left to you is the memory floor: how much of the box stays free for everything else. Below that line the engine manages itself. To trust it I tried to break it. I ran every arrangement I could think of: dozens of small conversations at once, a handful of enormous ones side by side, deep ingestions landing while others were mid-flight. I planted needles at the bottom of the deepest contexts to prove nothing quietly goes missing, watched decode as the box filled to its floor, and left it serving heavy load around the clock. It held: 3 million tokens active, every needle found, nothing degraded after a day of continuous serving. Full 1M context in one bank is no problem either, staying fast and flawless. I loaded a single conversation of 975 thousand tokens: 25 minutes of ingest averaging 633 tokens a second. Needle still found. Next turn started answering in two seconds, everything still warm. I gave extra care to update and overhaul the main docs: both READMEs explain the memory model and every knob in plain language, and each capacity claim published with setup details and measurements. Full changelog: github.com/Entrpi/ds4/blob/b… Now is a great time to upgrade, from any version: curl -sSL enia.cc/ds4 | bash -s -- --start nitter.cf/bleysg/status/20834486…
⚡ If you have or want a DGX Spark, this post might be the most important one you come across this month. Serve the BEST model available (DeepSeek V4 Flash 07-31, 284B parameters) on ONE Spark @ 1,000 tok/s prefill, 59 tok/s multi-agent serving. One command to install. 🧵
13
13
7
168
30,505
Overwhelmed by the outpouring of support from the big v0.5 launch a week ago, the day after DeepSeek V4 Flash 07-31 shipped. I've spent the whole week shipping the top priority updates from everyone's feedback, and the engine is really in a better place thanks to your detailed reports and suggestions. Six releases in seven days, latest release is 0.5.6, and every single one includes work on your suggestions and reports. What that looked like in practice: The crash a few of you hit under heavy agent load? Root-caused to a timing window a few billionths of a second wide and fixed, with a reproducer that shows 5 crashes in 16 runs before, 0 after, and identical output token for token. One of you bisected it to the right kernel family before I found the race. That's the community this project has now. The one I'm most excited about: now you can point Claude Code or Codex straight at your Spark and they just work. All four APIs now get the full batching engine, streaming, thinking, tool calls and all. I watched Claude Code run a real tool loop against the box before writing this. Its second turn reused nineteen thousand tokens of already ingested context and paid for 42 new ones. Agent loops are nearly free after the first turn. And a dozen or more quality of life fixes you asked for, if you want to read the full changelog: github.com/Entrpi/ds4/blob/b… Now is a great time to upgrade, from any version: curl -sSL enia.cc/ds4 | bash -s -- --start nitter.cf/bleysg/status/20834486…
⚡ If you have or want a DGX Spark, this post might be the most important one you come across this month. Serve the BEST model available (DeepSeek V4 Flash 07-31, 284B parameters) on ONE Spark @ 1,000 tok/s prefill, 59 tok/s multi-agent serving. One command to install. 🧵
11
12
157
17,088
Some other notables: antirez took note of my work and has upstreamed a good chunk of my kernels, so the average Spark user on upstream will also benefit from faster prefill. My engine still has a lot of advantages and a performance edge, but it's good to know that many people who only known upstream will have a much better experience on Sparks now. There's also a lot of work to improve memory accounting and management, so hopefully cases where the engine slows down or you can't run deep banks become much rarer. I've personally tested filling multiple 200k banks and gotten to 1.15M tokens active between them without memory pressure this week.
8
583