@bleysgi
iAccount based inAustralia
About this account
- Account based in
- Australia
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Helping people engineer the future. The core metric is task/GJ (gigajoule) and GJ/humanity.
SF / LA
Joined March 2009
- Tweets696
- Following2.1K
- Followers3.4K
- Likes6K
Pinned Tweet
I’ve been building an interactive model of the AI transition for a while, and wasn’t quite planning to make it public yet.
AI 2040 is making the rounds today. There’s enough overlap, and enough difference, that this feels like the right moment.
enia.cc
Prediction: This diagram I made today is going to be the most important thing to understand in AI PCs and AI SMB servers for the next few years.
No hardware yet follows this pattern. It's a natural evolution of what hybrid bonding gives to the PC.
🧵👇
nitter.cf/NVIDIAAIInfra/status/2…
Announcing the expansion of NVIDIA NVLink Fusion with NVHBM, a next-generation high-bandwidth memory technology that brings higher memory performance and efficiency to XPUs.
Amazon's @AnnapurnaLabs will be the first to work with us on NVHBM, combining @awscloud custom silicon with our memory technology and the NVLink scale-up architecture to enhance performance and efficiency for AI workloads.
Learn how we're helping hyperscalers and AI innovators build the next generation of AI infrastructure: nvda.ws/4xpBtYN
How do you make a DGX Spark that can run a 3T class model like Kimi K3 at good speed on one box?
I break down how this is achievable before 2030.
Prediction: This diagram I made today is going to be the most important thing to understand in AI PCs and AI SMB servers for the next few years.
No hardware yet follows this pattern. It's a natural evolution of what hybrid bonding gives to the PC.
🧵👇
nitter.cf/NVIDIAAIInfra/status/2…
The trick is the same as Raptor's 3D DRAM, just here applied to moving all the traditional northbridge/southbridge elements into the lower die.
nitter.cf/bubbleboi/status/20915…
Arithmetic on a ~900 mm² N6 platform die: 96 LPDDR6 channels (2,304-bit) ≈ 4.15 TB/s, 1.5–2 TB of DRAM, compute-die tax ≈ 0.
Next act: add bonded DRAM... 20–40 GB per full-footprint layer, 64–160 GB at 10 TB/s. @Xiaomi fast/slow AI Cube architecture, scaled up.
nitter.cf/ItsmeAjayKV/status/209…
Xiaomi just showed its AI Cube Prototype and this could become a serious GB10 competitor from China 👀
- 3 custom chips: Xring O3, O100, D100
- 200 TOPS NPU
- 1.22 TB/s AI memory bandwidth
- Up to 160GB unified memory
- 150W sustained power
- 120B models running locally
Xring O100: 1.22TB/s + 330 t/s on a 150w AI box is 🔥
Once it hit's the marked, going to sell like hot cakes.
Prediction: M7 Ultra gets native FP4 matmul acceleration.
M5 Ultra has the bandwidth-cutting primitives which help decode for FP4 and FP8, but still computes in FP16.
Getting Blackwell style FP4 compute gains that multiply prefill perf will allow Apple to finally catch up on the whole picture of AI perf/W.
M5 Ultra has me excited for next year.
1.2 TB/s bodes well for next year's Rubin DGX Spark release and for M7 Ultra. Both will ship between late 2027 and early 2028, both will move to next-gen LPDDR6 memory. Both should move to 2 TB/s territory.
That's faster than RTX PRO 6000! M7 Ultra will likely go up to 1 TB per box too. That would allow you to run full sized Kimi-K3 class frontier models fast on just two boxes.
At this rate, frontier class models should run on just one box by the time of the Feynman DGX Spark generation late 2029.
M5 Ultra just boxed in my projection of "When do we get frontier models at home?"
Why? Because they're copying the DGX Spark model: Dare you to have just one.
512GB M5 Ultra is $17.7k in Oct. It will take 4x to run Kimi K3. Apple is supporting this arrangement. This will be the most economical way to run Kimi K3 "at home" ... $71K.
That's not really an "at home" price for most of us, but before that the most economical option was looking like it would be 8x B300 in a DGX B300 for around $550k. That's a rack-mount beast needing power and cooling consideration, not a cool and quiet desktop box that can plug into any average circuit. The costs and practicalities made it a real consideration only for SMBs with serious on-prem requirements.
A 4x M5 Ultra cluster can sit anywhere and give practically fast frontier level model performance, before the end of this year!
CEO of Cerebras wants to land fabs without red tape in the US as a sovereign strategic priority.
This is becoming a hot button issue. Local communities don’t want the noise, etc.
He’s not pointing at that though, he’s pointing at the slowdowns that happen as a result of having to conform to other practical local norms like fire codes intended for strip malls.
My current thesis is that we see a major compute crunch from demand outpacing global compute supply before 2028. Unlocking more power and compute, and getting obstacles out of the way now to scale buildouts at Terafab level ambitions are probably the only way we get a world where most compute happens in the US and generally a world where our collective ambitions aren’t heavily bottlenecked by a compute crunch for years.
There’s a simple solution to the current memory and compute shortages.
It comes down to one policy decision: building fabs.
I suggest giving TSMC, Samsung, SK Hynix, Micron, Global Foundries, Intel a 20-year window to build fabs in specific areas.
Free from every local ordinance.
That sounds extreme. It isn't.
It’s a national, strategic imperative.
Fabs are modern pyramids.
They are the greatest things humans make, and the chips they produce are a foundation of our economy.
Yet we have to build them around the same local building codes written for strip malls and parking lots.
Fabs are the size of 50 football fields - it’s a red tape nightmare.
One multibillion dollar fab in Texas and had to be completely redesigned over a single local fire ordinance.
It set them back the better part of a year.
In the US, we've lost the ability to quickly build the infrastructure we obviously need.
Cutting-edge fabs included.
And we didn't just lose the fabs.
We lost the entire surrounding ecosystem.
The packaging expertise.
The strategic jobs.
The deep manufacturing know-how.
It is extraordinarily important that we get this full ecosystem back.
Lets get to the root of the problem: permitting, local ordinances, zoning, the whole regulatory patchwork.
Fabs are a strategic asset.
We have to get them built in America.
I haven't slowed down in the two weeks since my last post. Entrpi/ds4 just hit v0.6.2 and is now at a great robustness and scaling milestone. 3M active tokens across many parallel agents is running no sweat under stress on just one DGX Spark today.
vLLM and SGLang were designed around dedicated VRAM. I designed around the Spark in particular, where the engine shares the box with other apps and models. I want you to be able to run an STT model alongside without worrying about OOM and without compromising capacity more than absolutely necessary. To achieve that, the engine measures what each request actually uses, charges exactly that, sizes its budgets from what the box really has, and hands idle memory back. Context is demand-mapped: nearly free until a conversation fills it. When something does not fit, it reclaims idle state, then refuses with a reason instead of OOMing the machine. While idle it audits its own ledger and logs any drift. The one decision left to you is the memory floor: how much of the box stays free for everything else. Below that line the engine manages itself.
To trust it I tried to break it. I ran every arrangement I could think of: dozens of small conversations at once, a handful of enormous ones side by side, deep ingestions landing while others were mid-flight. I planted needles at the bottom of the deepest contexts to prove nothing quietly goes missing, watched decode as the box filled to its floor, and left it serving heavy load around the clock. It held: 3 million tokens active, every needle found, nothing degraded after a day of continuous serving.
Full 1M context in one bank is no problem either, staying fast and flawless. I loaded a single conversation of 975 thousand tokens: 25 minutes of ingest averaging 633 tokens a second. Needle still found. Next turn started answering in two seconds, everything still warm.
I gave extra care to update and overhaul the main docs: both READMEs explain the memory model and every knob in plain language, and each capacity claim published with setup details and measurements.
Full changelog: github.com/Entrpi/ds4/blob/b…
Now is a great time to upgrade, from any version:
curl -sSL enia.cc/ds4 | bash -s -- --start
nitter.cf/bleysg/status/20834486…
Full setup repo: github.com/Entrpi/ds4-on-spa…
Overwhelmed by the outpouring of support from the big v0.5 launch a week ago, the day after DeepSeek V4 Flash 07-31 shipped. I've spent the whole week shipping the top priority updates from everyone's feedback, and the engine is really in a better place thanks to your detailed reports and suggestions. Six releases in seven days, latest release is 0.5.6, and every single one includes work on your suggestions and reports.
What that looked like in practice:
The crash a few of you hit under heavy agent load? Root-caused to a timing window a few billionths of a second wide and fixed, with a reproducer that shows 5 crashes in 16 runs before, 0 after, and identical output token for token. One of you bisected it to the right kernel family before I found the race. That's the community this project has now.
The one I'm most excited about: now you can point Claude Code or Codex straight at your Spark and they just work. All four APIs now get the full batching engine, streaming, thinking, tool calls and all. I watched Claude Code run a real tool loop against the box before writing this. Its second turn reused nineteen thousand tokens of already ingested context and paid for 42 new ones. Agent loops are nearly free after the first turn.
And a dozen or more quality of life fixes you asked for, if you want to read the full changelog: github.com/Entrpi/ds4/blob/b…
Now is a great time to upgrade, from any version:
curl -sSL enia.cc/ds4 | bash -s -- --start
nitter.cf/bleysg/status/20834486…
Some other notables:
antirez took note of my work and has upstreamed a good chunk of my kernels, so the average Spark user on upstream will also benefit from faster prefill. My engine still has a lot of advantages and a performance edge, but it's good to know that many people who only known upstream will have a much better experience on Sparks now.
There's also a lot of work to improve memory accounting and management, so hopefully cases where the engine slows down or you can't run deep banks become much rarer. I've personally tested filling multiple 200k banks and gotten to 1.15M tokens active between them without memory pressure this week.