@RedHat_AIi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Accelerating AI innovation with open platforms and community. The future of AI is open.
Joined May 2018
- Tweets2.6K
- Following2.1K
- Followers12.8K
- Likes1.8K
Pinned Tweet
Everything you need to start self-hosting an open LLM.
Run it on your own hardware. No API keys. No per-token bill. Nothing leaves your machine.
The full path with @vllm_project: batch inference in Python, an OpenAI-compatible API server in one command, and quantized models that cut an 8B from ~16GB of weights to a quarter of that while keeping 98-100% accuracy.
Walkthrough by @cedricclyburn.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Cloud computing got cheaper per transaction as you scaled. Gen AI does the opposite: metered token costs climb with usage and never flatten. Agents make it worse, some use ~1,000x the tokens of a simple chat.
The playbook for getting ahead of it:
redhat.com/en/resources/ente…
Red Hat AI retweeted
vLLM now supports NVIDIA Vera Rubin. The early results show more than 7.8x the throughput of GB200 on MiniMax M3 on AgentX.
@inferact, @NVIDIAAI, @RedHat_AI, and the vLLM community have been bringing vLLM up on Rubin since it was announced. Here is where things stand.
🧵 1/5
Red Hat AI retweeted
More state-of-the-art MXFP4 ckpts ready for inference on AMD GPUs and @vllm_project available at: huggingface.co/RedHatAI/Qwen…
Red Hat AI retweeted
Serving AI at scale brings real tradeoffs in cost, latency, and capacity.
Hear from vLLM core maintainer Tyler Michael Smith on inference economics in San Jose, Oct 19.
Thanks @RedHat, @NVIDIAAI, and @IBM for bringing the community together!
PyTorch moved from research labs to running enterprise AI. On Monday Oct 19, @RedHat, @NVIDIAAI, and @IBM host a day in San Jose: @vllm_project economics, @_llm_d_ at scale, and agents in production, with the @PyTorch maintainers behind them.
Register: luma.com/wjmtvwvw
PyTorch moved from research labs to running enterprise AI. On Monday Oct 19, @RedHat, @NVIDIAAI, and @IBM host a day in San Jose: @vllm_project economics, @_llm_d_ at scale, and agents in production, with the @PyTorch maintainers behind them.
Register: luma.com/wjmtvwvw
Red Hat AI retweeted
Want to learn more about self-hosting your own models and agents? This Thursday at 10am ET I’m joining @hacktoberfest for some live demos and tips on using @vllm_project 🎃
Bring your questions, I'll try to get to as many as I can :) and we’ll be streaming live from @MLHacks
🔗 hacktoberfest.com
For decades, cloud got cheaper per transaction as you scaled. Gen AI inverts that. Metered token costs rise as usage climbs, and the curve never flattens.
A thread on why inference economics flipped, and what teams are doing about it:
The response: produce tokens instead of only renting them. Self-host capable open models, right-size so routine traffic goes to smaller models, and put a governed gateway in front. Our own teams at @RedHat ran 300M tokens in 48 hours this way, serving 1,400+ users internally using @vllm_project and @_llm_d_,
The efficiency comes from the stack. For example, on 16 NVIDIA H100s, @_llm_d_'s inference-aware routing served up to 2x the users on the same hardware, with up to 99% lower time-to-first-token under load.
The full playbook, from bare metal to agents:
redhat.com/en/resources/ente…
Red Hat AI retweeted
vLLM Office Hours #58 recording is up:
- What's new in vLLM v0.29 and v0.30
- A new hardware backend from Tenstorrent
- An intro to llm-d Semantic Classifier, a lightweight classification service for high-throughput routing
- Why it isn't trying to be a guardrail
Watch + see slides: youtube.com/watch?v=icrWIlo4…
vLLM Office Hours #58 recording is up:
- What's new in vLLM v0.29 and v0.30
- A new hardware backend from Tenstorrent
- An intro to llm-d Semantic Classifier, a lightweight classification service for high-throughput routing
- Why it isn't trying to be a guardrail
Watch + see slides: youtube.com/watch?v=icrWIlo4…
An NVFP4 checkpoint of Qwen3.8-Flash-Next is on Hugging Face. MoE experts quantized to FP4, the rest kept in BF16, ready for vLLM. Text, image, and video in.
We benchmarked it against other popular checkpoints. See how it holds up:
huggingface.co/RedHatAI/Qwen…
Variable-length speculative decoding used to need a confidence head. Not anymore.
vLLM 0.30 generalized adaptive verification to any draft-based method, Eagle and DFlash included, through an online acceptance estimator.
@mgoin_ on what shipped across 0.29 and 0.30 👇
By the way, 0.31 is already out. We'll cover that one next office hours. Get a calendar invite via the link in the reply.
Join vLLM Office Hours (biweekly, virtual): red.ht/office-hours
On prompt injection, a ~200M parameter classifier scored 89.01%. A 35B model scored 89.31%. A 0.20 point gap, at 54ms vs 312ms.
Our AI Safety team put the new "decision models" up against classic guardrails. Verdict: use the right tool for the job.
developers.redhat.com/articl…
Red Hat AI retweeted
Using frontier LLMs for every agentic sub-task drains budget and increases latency. Check out how we look into picking more targeted models drives us to cost-effective, and "smart-enough" agentic execution. sprou.tt/1aZWuT9HcGQ