@cloudnative_ai

Cloud ⛅️ + Data 📚 + AI 🤖

Joined September 2019
KAI retweeted
We're coming out of stealth. We've built our first racks after a successful A0 tapeout, $1B+ in customer contracts, and $800m raised. Early customer tests show us achieving SOTA throughput, latency, and power efficiency on inference workloads. Our first racks ship this summer.
647
955
610
9,911
6,774,089
KAI retweeted
One of the best recent podcasts on AI. Neil has a gift for explaining all of the jargon and insights simply. (I’m not involved, just found it unusually educational).
Neil Movva (@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power. We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are. What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow. Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible. We discuss: - Latency versus throughput - Why there are no bad chips, only bad pricing - The end of kernel engineering - Buying chips and power no one else wants - New chip architectures - Nvidia lore + his contrarian view of the company - Open source and the frontier labs I learned a ton. Enjoy! TIMESTAMPS 0:00 Intro 0:38 Building a “Token Factory” 4:21 The Future of Background Agents 13:09 Nvidia and the GPU Stack 23:27 Chips, Memory, and Transformers 36:14 The Future of AI Training Data 44:32 Chip Scarcity and Compute Arbitrage 52:44 Reinventing the AI Data Center 59:01 Power and the “Scavenger Strategy” 1:10:10 Open vs. Closed AI
164
1,509
26
14,888
1,726,428
SpaceXAI engineer (ex-Cursor): "I'm running 10-20 GrokBot agents right now, and they handle 90% of my routine. I don't even manage them, I have a Chief of Staff agent that knows all the others and manages everything." 50 minutes from a SpaceXAI engineer, showing exactly how to build a team of agents that works for you 24/7. Worth more than any $700 course on agentic engineering. Watch it, then read the step-by-step guide on building your own Grok agent team for free, with ready-to-copy prompts.
32
99
3
623
136,027
Just set up 2 DGX Sparks. Here's what you should know before doing it yourself:
4
8
2
103
97,336
Had a lot of fun chatting again with my twin brother @dylan522p We went through lab economics over the next few years - the shift from inference to training as RSI draws near; and how Anthropic and OpenAI are on track to control most of the world’s usable FLOPs within the next few years (because they can monetize compute better and thus outbid everyone). And then we discuss whether the >$10T of total AI capex we’ll see by the end of the decade will cause a sovereign debt crisis, where hyperscaler debt raises interest rates, drives non-AI exposed countries into bankruptcy, and crashes non-AI equities. One question we weren’t able to resolve is whether there’s anything that can counter all the forces barrelling towards centralization in this industry - the economies of scale in training, the scarcity of compute, and eventually continual learning and RSI. 0:00:00 – Two labs will soon control most of the world’s compute 0:07:01 – $6 billion in fab capex enables $1t+ of end revenue 0:13:08 – Compute prices will rise if the labs outbid everyone 0:18:22 – Which layer will capture most of the surplus? 0:25:40 – Will datacenter regulation slow down AI? 0:29:43 – Labs are shifting compute from inference to R&D 0:33:27 – China gets less than 10% of new compute, but its labs need less 0:48:48 – Will AI cause a sovereign debt crisis? 1:07:52 – Will the world's future workforce belong to a few companies?
131
166
100
1,961
1,024,968
2 months ago, we crashed Jensen's board meeting to show him Perplexity running locally on DGX Spark 🤣 (lol im holding architecture diagrams like he's gonna look at em during a board meeting) Local AI used to be for enthusiasts. Folks were running tiny quantized models on underpowered hardware, getting a few tok/s. Exciting, but not practical. But now GLM 5.2 is my daily driver. Local AI hit an inflection point with frontier open source models like GLM 5.2, Deepseek v4 flash, and Nemotron + hardware powerful enough to run them, like DGX Spark and DGX Station. Developers are running fleets of agents. But there has yet to be a great personal agent experience for local AI. The current bottleneck is the know-how to set up inference and get meaningful performance out of it. I'm excited about Perplexity Portable Computer bc it’s an app that fully sets up a great local AI experience out of the box.
Meet Portable Computer, Perplexity's new local-first agent stack on NVIDIA DGX Spark. When running locally, Portable Computer offers one-click local inference setup and an optimized agentic experience for DGX Spark. Learn more and get started today: blogs.nvidia.com/blog/local-…
28
27
13
511
96,376
KAI retweeted
Andrej Karpathy just rewrote the rules of using LLMs: "Prompting is going away. Delete everything, keep Graph." LLMs → Prompts → Agents → Graphs He dropped his full 2-hour course from Stanford on "Graph-Native Research" • 00:00 - Intro to Graph systems • 01:08:09 - LLMs architecture This Karpathy course can replace a $100K Yale LLM senior degree. Watch it today, then save the full graph engineering guide below
20
43
1
308
43,550
The Ultiamte Step-By-Step LLM Engineering Projects Roadmap - Build a tokenizer - Learn embeddings - Implement RoPE / ALiBi - Hand-wire attention - Build MHA - Build a Transformer block - Train a mini-former - Compare objectives - Build sampling - Speculative decoding - KV cache - MQA / GQA / MLA - Long context - FlashAttention - Hardware budgets - Toy MoE - Sparse model trade-offs - State-space / linear attention - Diffusion language models - Data pipelines - Synthetic data - Scaling laws - SFT / DPO / RLHF / GRPO - Quantization - Serving stacks - Eval harnesses - RAG - Tool use / agents - Vision-language adapters - Interpretability - Red-team suite - Full capstone model system
14
146
5
1,116
69,679
NVIDIA CEO, Jensen Huang: "Nobody writes prompts anymore, the new job is building Loops and Graphs." In 50 minutes he breaks down what replaced prompting and why most people haven't caught on yet. It's the difference between using AI and having AI work for you. Watch it, then read the guide below on how to build a system that improves itself.
48
129
5
674
285,273
ANTHROPIC LEAKED THE 6 REPOS THEY BUILT CLAUDE CODE ON - THEY REPLACE A $120K/MONTH TEAM AND COST $3 A DAY anyone can clone 6 repos - almost nobody assembles them into a team that runs for $3 a day. core → skills → patterns → gates → self-patch 4 executors on the core are the largest line on the bill - $1.10 a day, and that's where you get 4 PRs instead of one. move that layer to the priciest model and the bill doubles with almost nothing to show for it. 7 procedures sit on disk, exactly 1 enters the window - a convention pasted into a prompt is rent, a convention in a folder is storage. the most expensive model sits on exactly 2 layers out of 6 - it plans and it judges, and that's the second-smallest line on the bill. the template arrives with 4 problems already solved: auth, payments, deploy, errors - the ones you'd rediscover yourself in week three. at the gates the cheap model triages 80 findings and the expensive one reads 4 - which is why it costs 60 cents, not 6 dollars. security scans every commit for 20 cents a day - the cheapest line on the stack and the first one everyone cuts. and the key part: the gate doesn't just reject, it writes the missing rule into the procedures - next week the team works to rules it wrote itself. the wrong model on the wrong layer costs 5x overpaid or a month of rework - and that's the part nobody publishes. save this and paste it into Claude Code - the repos are public, the assembly isn't ↓
6
8
77
18,253
RT @NVIDIAAIInfra: Shopee just published the results of building its own frontier LLM, which now powers use cases from search and personali…
1
1
Finally!!! Got LFM2.5-2.6B by @liquidai working with Hermes... Served it with @vllm_project The command that worked: vllm serve "LiquidAI/LFM2.5-2.6B" \ --enable-auto-tool-choice \ --tool-call-parser lfm2 \ --reasoning-parser qwen3 Just ask Codex to set it up if you get errors running it with vLLM Then configure Hermes: $ hermes model Select "custom provider", then add the URL of your vLLM instance. Double-check in the dashboard if the model is properly configured. If you need web search, configure @firecrawl and add the api key $ hermes tools (then configure in cli), select web search and scraping then add your firecrawl api key Send a test message to your bot to see if it works When you see tools being called then you're good Special thanks to @helloiamleonie of LiquidAI for reaching out when I had issues running it with Hermes
5
9
54
4,089
The @liquidai cookbook is such an underrated developer resource: Curious about fine-tuning text, vision, audio, or encoder models? Curious about fine-tuning with CPT, SFT, DPO, or GRPO? Curious about fine-tuning LFMs with Unsloth or TRL? It has it all. I just did a little cleanup. Enjoy! github.com/Liquid4All/cookbo…
29
119
6
879
65,994
KAI retweeted
What are the best models you can run on your @NVIDIAAI DGX Spark? ✨ Aug 2026 Edition 1× DGX Spark • DeepSeek v4 Flash - 1M ctx, 26 tok/s - recommended! • ⁠Qwen 3.6 35b NVFP4 - 256k ctx, 81 tok/s • ⁠Qwen 3.6 27b NVFP4 - 256k ctx, 33 tok/s ⁠Qwen 3.8 27b - should be released soon and this could change my recommendation! 2× DGX Sparks ← sweet spot! • DeepSeek v4 Flash 0731 - 1M ctx, 82 tok/s • Inkling-Small - 1M context, Full Omni, 33 tok/s • MiMo-V2.5 - 1M ctx, Full Omni, 31 tok/s • Step-3.7-Flash — 256K ctx, 30 tok/s 3× DGX Sparks • GLM-5.2 with Vision - 348k context, 25 tok/s - still the best intelligence you can run locally if you have 3 sparks. • DeepSeek v4 Flash on 2 units + smaller models on the 3rd spark for images, ComfyUI and other things. This setup gives you DeepSeek v4 Flash speeds for coding, plus strong agentic workflows and image support from smaller models. 4× DGX Sparks • GLM 5.2 NVFP4 across all 4 units - still think this is the way to go if you have 4 units! • The alternative is running DeepSeek v4 Flash on 2 units and any other 2x setup on the other units. The choice is your. Links and repos below 👇
78
100
16
858
149,223
The first Vera Rubin clusters are here! Yesterday, @IneffableLabs took delivery of their Vera Rubin NVL72 cluster from @googlecloud @nvidia The AI frontier jumps forward by yet another generation of hardware. Acceleration continues.
156
317
176
3,322
791,807
Current local ai stack 2x DGX Sparks - DeepSeek v4 flash (90 tok/sec) - implementation & sub agents 2x DGX Stations - GLM 5.2 nvfp4 (120 tok/sec) - planning and advising Brev for cluster networking, always on cloud agent, and orchestration
126
92
35
1,947
325,136
We threw a fun event on owning your ai stack today! 80 @sequoia portfolio companies attended technical workshops on how to own your AI (models, harnesses, data, evals, RL, CL, etc). Videos and takeaways coming soon. Towards a vibrant ecosystem for Democratized Intelligence 💚
12
9
2
155
23,857
KAI retweeted
He is 19, already built an AI system making $2.3M a year - and now teaches how to do it at Stanford 01:03 - he lost his job at Anthropic because of prompt engineering 09:34 - Context Engineering gives 10x to coding speed 28:47 - Opus 5 builds an AI system making $2.3M a year from scratch after watching I deleted all my prompts and switched to Context Engineering - first $20k+ in a month. Save & watch - the article below is a step by step guide to build an AI system
25
132
1
650
68,068
KAI retweeted
My friend applied to 200 tech jobs in two years. No MIT. No Stanford. Last month Anthropic offered him $750,000. I asked him how he broke in from zero. He sent me the exact video that got him in. Anthropic's 2 hour course on "How to become an AI engineer in 2026" Their core team shows exactly how to architect & build AI agents from scratch. I watched it last night. Halfway through, I realized I could break into an AI lab in weeks, not years. Bookmark this and read the article below. • 00:00 - AI agents with graph engineering • 06:41 - AI agent architecture • 14:31 - building AI agent loops live • 1:15:24 - AI agentic RAG • 2:20:42 - Anthropic interview process
50
131
7
849
144,300
KAI retweeted
OpenAI pays $785K/year to developers who know how to apply Forward Deployed Engineering in AI. in 40-minute talk, Head of FDE at OpenAI revealed full roadmap for how they actually use FDE internally: • 10% → 2:46 - why Morgan Stanley was their first FDE case • 30% → 9:05 - FDE eval-driven development explained • 55% → 16:01 - FDE live-demo: LLM rerouting a supply chain • 80% → 24:33 - advice for founders building FDE teams • 100% → 29:06 - the biggest FDE mistake they made this year 40 minutes replaces a $500 enterprise AI deployment course bookmark & watch - then read how to become FDE engineer in article below ↓
26
79
3
566
113,913