@ablearn23

Data Scientist

San Francisco, CA
Joined May 2011
Abhishek Sreesaila retweeted
I completely stopped using Codex app. Pi + Herdr + Collie + Executor is unbeaten. Is not even close. The only thing I need to figure out is Scheduling tasks. Any advice ?
120
58
6
1,291
159,321
Abhishek Sreesaila retweeted
a really cool perspective on “compilers 2.0” from a key jalapeno engineer
"openai engineer explains /why/ he didn't need to understand the kernel line by line" 🙃 we're doing compilers 2.0
1
27
323
57,473
Abhishek Sreesaila retweeted
AI Kernel gen is easy with the rapid advancement of coding agents - if you have a compiler and dsl like triton/gluon. The real metric is end to end enablement of the entire stack, plus unlocking/integrating a developer community that can share code from NVIDIA ecosystem👇
Support for new AI hardware usually takes a big team, many repos, and a year. Two and a half engineers got a frontier open model serving on @Qualcomm Cloud AI 100 in under six months using the Modular stack, followed by GPT-2 on Qualcomm Dragonfly™ AI 200 in a week. Inside the bringup at ModCon 2026: youtube.com/watch?v=orJP0cxf…
10
17
1
216
39,064
Abhishek Sreesaila retweeted
I have never been more bullish on the future of AI compilers, including XLA. Frontier AI is starting to find compiler optimizations no human could write. Soon, frontier models will point hundreds of autonomous agents at the entire compiler stack — cycle count, latency, and throughput as the reward signal. It's the ideal RSI task. Fully verifiable, and every speedup compounds into cheaper training for the next model. Assuming AI keeps progressing at its current rate, AI compilers may very well eclipse handwritten kernels. Chris is right in a pre-AI world. Not in a post-AI future.
Modular is the only team that can do what it is doing, because we intentionally took another path Modular is more of “anti ai compiler”. These make great demos but have never scaled to the full generality of accelerator hardware at full performance.
18
14
2
207
32,016
Abhishek Sreesaila retweeted
The Modular ecosystem is now Open Source! 🚀 From Max to the Mojo compiler, built on MLIR/LLVM. Breaking news, huge thanks to @clattner_llvm and everyone who has worked on this over the past 4.5 years!
3
4
28
1,731
Abhishek Sreesaila retweeted
Weekend reading! A well documented ISA is a big advantage of AMD over NVIDIA, nice to see it continue for CDNA5.
8
29
1
453
42,466
Abhishek Sreesaila retweeted
BREAKTHROUGH: A full, unmodified 2.78-trillion-parameter Kimi K3 on a consumer laptop by streaming only the activated experts from NVMe. YOU CAN’T RUN KIMI K3 “ON THAT” THEY DECLARED. There are many paths to do it. This is one: Marco Bambini Just Gave Us the Full Kimi K3 on a Laptop Meet Marco Bambini he did something that felt impossible only a day ago. He built WASTE Weight-Aware Streaming Tensor Engine a clean, dependency-free C inference engine that runs the complete, unmodified 2.78-trillion-parameter Kimi K3 model by streaming only the activated experts straight from NVMe. No distillation. No pruning. No cloud. The full open-weight model. We have it running in the lab right now. What Marco Actually Built Kimi K3 is a sparse Mixture-of-Experts system. Only about 4 % of its weights fire on any given token. Marco’s insight was simple and ruthless: the idle experts do not need to live in RAM. They only need to be reachable in time. WASTE keeps the model’s “trunk” (attention, shared components, embeddings) resident in memory — roughly 27 GB on the converted container. The 82,000+ routed experts stay on disk as tightly packed residual vector-quantized records. When the router selects its 16 experts per layer, the engine issues direct, cache-bypassing reads from the internal NVMe and feeds them into a bounded expert cache. The rest of the machine’s RAM becomes working space for that cache. On a 64 GB MacBook Pro with the container on the internal SSD, we are measuring 0.32–0.34 tokens per second at a comfortable memory budget. Prefill sits a little higher. The vision tower works. Logits match the reference implementation to within a few parts in a million. It is the real model. The container itself is 982 GiB after conversion from the original 1.42 TB MXFP4 weights. Minimum RAM floor is just over 29 GB for short context. Push the budget higher and the expert cache hit rate climbs; push too high and you start paging and the speed collapses. The sweet spot on current consumer hardware is clear and measurable. How We Are Testing It We converted the official weights, verified the container, and began systematic runs the same day the engine stabilized. First we confirmed numerical fidelity against the PyTorch reference on short prompts. Then we moved to longer generation, vision inputs, and multi-turn chat using Kimi’s native XTML format. We are measuring wall-clock decode, expert I/O versus compute split, cache hit rates at different RAM budgets, and thermal behavior under sustained load. We are also exercising the OpenAI-compatible server that sits on top of the same C library so we can drop the model into existing agent loops without rewriting anything. Early observations: •Expert I/O dominates the timeline, as expected. On a fast internal NVMe the engine is already near the practical ceiling of the storage subsystem. •The architecture’s sparsity is the entire enabler. A dense model of this size would be dead on arrival for local use. •Context length is currently limited by RAM more than by the model itself. Practical working contexts sit comfortably in the tens of thousands of tokens on 64 GB hardware; the full million-token window will need more memory or smarter KV management. •Thinking tokens are expensive at this speed. Long internal monologues turn into multi-hour runs. For agent work we are already experimenting with tighter control over when full reasoning is requested. We are treating this as a research instrument, not a finished product. Every run teaches us something about expert locality, prefetch opportunities, and how far pure software streaming can push trillion-scale inference on ordinary machines. 1 of 2
86
184
37
1,253
149,208
Abhishek Sreesaila retweeted
🚨SHOCKING: An Indian developer just hit #1 on GitHub with a prompting framework that outperforms every major benchmark. No VC money. No research lab. Just a laptop and 14 months of testing. Here are the 11 prompt patterns from his repo that I've been using for 3 weeks:
7
56
2
444
73,417
Abhishek Sreesaila retweeted
Who wants to come on my podcast this week? 2,000,000+ listens per month. 1. You teach one AI tool, skill, or framework that helps people build a business 2. You can have 1 follower or 1M, doesn't matter 3. You come prepared Tag someone or tag yourself. Lets go.
807
51
21
1,390
148,152
Abhishek Sreesaila retweeted
If you're new or looking to get more connected to SF tech: bookmark these 35+ irl events 🗓️ A list of what's happening this week (July 6 - July 13) ⬇️
5
9
133
39,414
Abhishek Sreesaila retweeted
Costco is quietly hoping you never read page 47 of your membership terms. I did. There's $1,370/year sitting inside the Executive Membership most people already pay for — untouched. Here's everything they don't put on the checkout screen 🧵
95
195
6
2,747
2,999,683
Abhishek Sreesaila retweeted
For anyone attending AI Engineer SF this week I made a list of some of the interesting startups hiring there. I’ve grouped by location, and noted a few of the founders, and which people from their co are at the conference:
5
10
170
34,770
Abhishek Sreesaila retweeted
We’ve been working with NVIDIA in their HQ for the past month. We’re going to make Local AI The Default. BIG news to share at Local AI summit, SF, July 2nd.
MASSIVE NEWS Teamed up with NVIDIA to make Local AI The Default
123
113
33
2,028
304,250
Abhishek Sreesaila retweeted
Jane Street, one of the richest and most secretive firms in the world, paid him between $330,000 and $600,000 a year, and in just a couple of months he built an AI system that runs TRILLIONS of operations per second "we just hired a kid... and he turned out to be a supercomputer in a human body" is what they're whispering now at Jane Street a math genius who pushed supercomputers to their limit. now Wall Street's quant traders are in shock in this hour-long lecture he breaks down how to use his machine to process trillions of data points bookmark it right now and watch it instead of reels to learn how to do the same ↓
Jane Street showed the code that's made them billions, written in a language from 1996 that almost everyone else has abandoned. Ron Minsky: "type systems turn out to be a really good tool for catching a surprising number of mistakes" the whole bet is on the type system. the compiler catches errors before the code even runs. the Option type forces you to handle the case where a value is missing. a skipped null check that crashes Python in prod simply won't compile in OCaml. "the compiler essentially forces you to do case analysis" in trading a single bug costs millions, so the priority is reliability, not speed of development. 33 minutes, and you'll see the code that holds up the billions of one of the most secretive funds on Wall Street. bookmark it. this is worth more than any $500 vibe-coding course. ↓
22
192
8
1,617
344,959
Abhishek Sreesaila retweeted
China released an AI employee that works 24x7 on its own and runs 100% locally It researches, codes, builds websites, creates slide decks, and generates videos. All by itself. All on your computer. 100% Open Source.
39
258
22
1,677
150,729
Abhishek Sreesaila retweeted
AI storytelling is getting crazy ChatGPT Image 2 can turn one story idea into a full storyboard, then Seedance 2.0 turns it into a cinematic animated video in minutes. step by step tutorial with prompts:
🤖 Made with AI
19
31
4
176
38,755
Abhishek Sreesaila retweeted
As always, Thank you for reading this. If you enjoyed this post: 1. Follow me @hasantoxr for more of these 2. RT the tweet below to share this thread with your audience
This is genuinely impressive. Gauth just dropped Atlas and it might be the end of textbooks. Type any topic like "Silk Road," "how a camera works," "fall of Constantinople" and it builds you a hand-drawn, interactive visual world you can walk through. No more reading walls of text. You explore knowledge like a map. Here's how to use it (step by step): ↓
2
4
18
8,423