@HowDevelop

AI Architect & Builder - Inference, LLMs, SLMs GSoC Org Admin - Jenkins | Docker Captain | CNCF Ambassador

New Delhi
Joined December 2018
What an amazing day! Presented a poster at @PyTorch Conference: how we optimised the RLM paper with pre-fix caching & batched Sub calls with @vllm_project And a full house for our Pytorch Conference talk on running LLMs using Executorch on Android!
4
5
2
84
12,834
Shivay Lamba retweeted
I’ll explain the new Pro 200 plan differently, before I start live tweeting from DevDay on things that are going out! Today we are going to ship a number of things that increase what you can do across the Plus and Pro plans. A lot of compute is online for this increase. As we increase the floor, we are changing the relative difference between plans to be Plus = 1X Pro 100 = 5X Pro 200 = 10X and we are reopening subscriptions for Pro 200 (we had paused it). If you have an existing plan you will keep the 20X multiplier for a bit and also receive a lot of additional credits because we know changes are hard even if it means that everyone will get more in the end.
2,861
412
1,056
8,540
1,280,922
Best person in AI one can hire
Professional Update: I have transitioned from my full-time role at iii, though I remain committed to contributing to the company's ongoing success, as I deeply value the innovative work we have accomplished together. If you are developing next ambitious thing and require expertise in Engineering, Product, or Developer Relations, hit me up!
1
3
382
Shivay Lamba retweeted
Get ready.
2,921
3,595
3,610
61,872
15,484,343
Shivay Lamba retweeted
Training AI agents takes more than fast token generation. Tools run. Tests stall. Training batches wait. We cut fixed-weight batch collection time by 66.9% on a Terminal-Bench-based workload by improving routing, sandbox execution and scheduling. Here’s how: nebius.com/blog/posts/in-age…
4
3
2
24
8,929
Shivay Lamba retweeted
A $200 Claude Code sub gets you ~$9,000 of Opus usage per month. I have managed to run 3 Claude accounts down to 0% since Opus 5.5 dropped. When analyzing their usage, costs were $2,086, $2,442, and $2,182. Average is $2,200/week, so ~$9,000/month
263
86
70
5,513
1,341,226
Countries travelled till date: Usa Canada Mexico Australia India Nepal Sri Lanka Malaysia Thailand Singapore Japan South Korea UAE Turkey China Hong Kong Taiwan Spain France Germany Netherlands Belgium Austria Italy Czech Republic Liecestien Switzerland Norway Sweden Finland Denmark Poland Uk Ireland Scotland Colombia @hrittikhere @Sauain share yours :P
Countries travelled till date : - Philippines - Malaysia - Singapore - Thailand - UAE - Sri Lanka - Vietnam Sirf itne he hain 🥲
15
9
68
7,833
Shivay Lamba retweeted
So impressed with the Opus 5.5 explainer videos - Here's another one - Serve Models from a @Kit_Ops ModelKit on @HAMiProject Detailed explanation of how Your model is packaged as a versioned ModelKit. KitOps pulls it from an OCI registry and unpacks it inside the Pod. HAMi schedules a controlled GPU share. SGLang loads the model locally and serves an OpenAI-compatible API. One clear path from model artifact to working inference. KitOps and HAMi.
5
3
1
25
775
Qwen3.8-27B is now live on Nebius Token Factory. A compact 27B dense model for coding, research, and agent workflows, with a focus on planning and completing tasks across multiple steps. Start building: tokenfactory.nebius.com/endp…
2
5
5
52
5,643
We’ve spent a lot of time talking about making models smarter. The next phase of AI infrastructure is going to be about making inference smarter. Speculative decoding is a good example of that shift: instead of asking one large model to generate every token, a smaller draft model proposes tokens while the larger model verifies them. You’re essentially turning inference into a collaboration between models, trading a bit of extra system complexity for much better generation speed. What I find more interesting is where this leads. As models become more interchangeable, the winning inference stack may not be built around a single “best” model. It could be built around the right combination of models, runtimes, hardware, caching, routing, and serving strategies for a specific workload. That makes inference infrastructure feel less like “hosting a model” and more like systems engineering. I was reading @Jozu_AI ’s work on bringing speculative decoding to Jozu RICs, and it’s a good practical example of this direction: taking an inference optimization technique and making it easier to package and run in a real deployment environment.
8
4
2
8
838
Shivay Lamba retweeted
As an AI Engineer. Please learn: > Harness engineering, not just prompt engineering > Context engineering, not just long prompts > Prompt caching vs. semantic caching tradeoffs > KV cache management, eviction, reuse, and memory pressure at scale > Prefill vs. decode latency and why they optimize differently > Continuous batching, paged attention, and throughput optimization > Speculative decoding vs. quantization vs. distillation tradeoffs > INT8, INT4, FP8, AWQ, GPTQ, and when quantization hurts quality > Structured output failures, schema validation, repair loops, and fallback chains > Function calling reliability, tool contracts, argument validation, and idempotency > Agent guardrails, loop budgets, tool budgets, and termination conditions > Model routing, graceful fallback logic, and degraded-mode UX > RAG architecture: chunking, embeddings, hybrid search, reranking, and freshness > Retrieval evals: recall, precision, grounding, attribution, and citation quality > Evals: golden sets, regression tests, adversarial tests, LLM-as-judge, and human evals > LLM observability as a first-class discipline: traces, spans, tokens, latency, errors, and drift > Cost attribution per feature, workflow, tenant, and user journey, not just per model > Safety engineering: prompt injection defense, data leakage prevention, and permission boundaries > Multi-tenant isolation, cache safety, and cross-user context contamination prevention > Fine-tuning vs. in-context learning vs. RAG vs. distillation, and when each is the wrong tool > Latency, quality, cost, and reliability tradeoffs across the full inference stack > Production failure modes: hallucinated tool calls, malformed JSON, stale retrieval, runaway agents, and silent eval regressions > Shipping LLM systems as reliable infrastructure, not demos wrapped around prompts aiengineeringfromscratch.com…
11
179
17
1,428
88,095
Shivay Lamba retweeted
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
442
359
160
3,296
2,359,098
So impressed with the Opus 5.5 explainer videos - Here's another one - Serve Models from a @Kit_Ops ModelKit on @HAMiProject Detailed explanation of how Your model is packaged as a versioned ModelKit. KitOps pulls it from an OCI registry and unpacks it inside the Pod. HAMi schedules a controlled GPU share. SGLang loads the model locally and serves an OpenAI-compatible API. One clear path from model artifact to working inference. KitOps and HAMi.
5
3
1
25
775
Shivay Lamba retweeted
I built WearScout because I kept seeing great outfits and spending way too long searching for similar clothes online 👕 Give it a reference image, it uses @deepseek_ai v4.1 Flash via @nebiustf to describe the look, and Jev drives a real browser using @browser_use to find similar pieces across stores, check product pages, and return a shortlist with prices and links. I tried it with a Tom Cruise outfit. I made a wrong call on privacy with my first demo, so I took that post down and changed the example. Heavily inspired by and adapted from @_nancychauhans’ Hearth project 🙌
Just let Jev drive a real browser for house hunting… the numbers are insane ⚡ 🏠One prompt, 4 rental sites crushed in 1m 16s for $0.0454 and it handed me 21 houses: • 32 pages visited (all 4 sources) • 76 browser actions, 60 clicks • 112 model calls
5
3
16
1,706
Here's an video introduction to @tan_stack AI using a single prompt to generate a launch style video about Tanstack UI using @ClaudeDevs Opus 5.5 using @Remotion Prompt shared below 👇
7
1
1
24
4,173
❯ Create and render a 15-second motion graphics video introducing TanStack AI. Make it feel like a premium developer-tool launch film produced by Studio1: fast, precise, visually surprising, and polished enough to open a keynote or lead a social campaign. Use the TanStack AI feature details below as the source of truth. Studio1 is presenting the film; do not imply Studio1 built TanStack AI. Format and deliverables - 1920×1080, 30 fps, exactly 15 seconds; keep important content within a 9:16-safe central area. - Deliver the rendered MP4 and the editable project/source. Use Remotion or another code-based motion graphics workflow if available. - Use real, editable typography and vector-style graphics. Do not rely on AI-generated text inside images. - Add an original electronic sound design: a restrained opening pulse, tightly synchronized transition hits, rising - Add an original electronic sound design: a restrained opening pulse, tightly synchronized transition hits, rising momentum, and a satisfying final resolve. The story must also work muted. Visual direction Dark graphite canvas; luminous gradients and controlled flashes inspired by TanStack’s identity. Crisp typography, kinetic code, layered interface fragments, streaming data, branching paths, and smooth camera moves with convincing depth. High contrast and immaculate spacing. Every transition should visually demonstrate a feature, rather than simply replace one title card with another. Keep all on-screen wording brief and legible. Avoid generic stock footage, spinning logos, fake dashboards, excessive glitch effects, and tiny code nobody can read. Exact 15-second sequence - 0:00–0:02 — Hook: A single prompt cursor blinks. It sends one signal that expands into a network of possibilities. On-screen: “What can one AI API become?” - 0:02–0:04 — Core and providers: The signal resolves into chat(); provider nodes connect and swap rapidly without breaking the central flow. On-screen: “chat() · 24 providers” - 0:04–0:06.5 — Multimodal: One continuous stream transforms into text, waveform, image frame, and video timeline. On-screen: “One pattern. Every modality.” - 0:06.5–0:09 — Agents and tools: The stream branches into sandboxed agent runs, then reaches external tools through MCP. Show distinct runs progressing in parallel. On-screen: “Agents · Sandboxes · MCP” - 0:09–0:11.5 — Reliability: An interrupted stream freezes, reconnects, and resumes from the same point; a clean type-check confirmation snaps into place. On-screen: “Persistent. Durable. Type-safe.” - 0:11.5–0:15 — Reveal: All paths converge into a bold final composition: “TanStack AI” followed by “Build what’s next.” Add a small “A Studio1 film” credit, visually subordinate to TanStack AI. Hold the final frame long enough to read it. Accuracy guardrails The supplied article says TanStack AI is in the release candidate phase, supports 24 providers, uses AG-UI, and offers composable chat(), media generation, sandbox and agent-harness primitives, MCP, type safety, persistence, and resumable streams. Do not call it a stable v1 release. Do not present planned orchestration features as already shipped. Make the video, inspect the rendered result, and fix any clipped text, weak transitions, mistimed beats, or unreadable frames before delivering it. The goal is a memorable showreel, not a slideshow of feature bullets.
1
1
317
Shivay Lamba retweeted
LiveKit has acquired @LoopholeLabs, an infrastructure company whose technology will enhance our platform for building voice, video, and physical AI agents. With their team, we move closer to offering customers essentially unlimited concurrent agents. Welcome to LiveKit! fortune.com/press-releases/l…
5
16
3
96
7,790
Shivay Lamba retweeted
Claude Code will now try to find a graceful stopping point when you hit your 5-hour limit mid-task, instead of cutting off mid-edit. It gets a small, fixed allowance pulled from your weekly limit to wrap up what it can.
1,086
1,235
822
34,923
2,621,849
Shivay Lamba retweeted
We're partnering with @Qualcomm to bring personal AI context to devices powered by Snapdragon. Liquid Context, our on-device context layer, is now optimized for Snapdragon processors and runs on the Qualcomm Hexagon NPU. Our goal: Give the agents people choose an understanding of what matters to them and when they need help. With the user's permission, Liquid Context learns from device signals and builds an understanding of their routines, preferences, and needs. That understanding is built and maintained locally, and Liquid Context shares relevant context with the user's chosen agents, whether they run on the device, in the cloud, or across both. That includes third-party agents and Liquid Agent, our efficient embedded agent powered by LFM2.5-2.6B. Running on the Hexagon NPU, Liquid Context works in the background and keeps that understanding current without requiring a cloud model to process every update. For device manufacturers, this is a path to add personal context to their devices while supporting their own choice of agents and services. OEMs building embedded or hybrid agents can also work with us to evaluate Liquid Agent. As our CEO @ramin_m_h said: "Personal AI starts with understanding how you live and what you need, when you need it. Liquid Context builds that understanding on your device so the agents you choose can offer more relevant help and anticipate your needs." > Read more about our partnership: liquid.ai/blog/liquid-contex… > Check out the livestream of @cristianoamon and Ramin's keynote here: youtube.com/watch?v=xKCto1Yf…
19
60
22
400
37,359