@RunAnywhereAI

Fastest inference anywhere: open models on-prem, hosted or on-device. Backed by @ycombinator W26

Joined July 2025
Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7
63
51
12
712
62,411
Why does AI make you wait? Every response is two phases: prefill reads the whole prompt in one pass, then decode writes one token at a time. Agents pay for both on every step and re-read the context each time, so prefill keeps growing. Wally (@RunAnywhereAI) makes both ultra fast.
2
2
6
288
RunAnywhere retweeted
One-shot drum machine with MiMo-V2.6-Pro on Wally One prompt, under 2 minutes, and it built a playable drum machine with every sound made in code. MiMo is the top open model on Artificial Analysis, and on Wally it runs miles faster than Xiaomi's own API. And it's free for everyone through tomorrow!👇
MiMo-V2.6-Pro is the #1 open-weights model on Artificial Analysis. On Wally it streams at 320 tok/s, 7x faster than Xiaomi's own API. The best open model at this speed should be in everyone's hands. So we partnered with @explabsai to make it completely free until 9/29. Go try it now!
3
1
8
621
The pelican test, but it's a race. Same model: MiMo-V2.6-Pro, the top open model on Artificial Analysis. Wally vs Xiaomi's own API. Wally finished 3.6x faster, and drew all four pelicans before Xiaomi's API finished its first. And it's completely free to use through tomorrow, on Wally via @explabsai.
free Mimo 2.6 Pro at 320 tokens/sec currently the best open weights model on the market cheaper and 50% better than GPT Luna only possible in partnership with @RunAnywhereAI platform.xplabs.ai
4
4
21
1,725
MiMo-V2.6-Pro is the #1 open-weights model on Artificial Analysis. On Wally it streams at 320 tok/s, 7x faster than Xiaomi's own API. The best open model at this speed should be in everyone's hands. So we partnered with @explabsai to make it completely free until 9/29. Go try it now!
free Mimo 2.6 Pro at 320 tokens/sec currently the best open weights model on the market cheaper and 50% better than GPT Luna only possible in partnership with @RunAnywhereAI platform.xplabs.ai
15
8
2
134
11,255
Every AI call makes you wait twice: for the first token, then for the rest. Agents pay that 50 times a task. Wally is the fastest place to run open models. Faster than the labs' own APIs, even their paid fast tiers. Try it now, free to start.
We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live. Wally: 363 tok/s. FlashX: 143 tok/s. 2.5x faster, 3.5x cheaper. Why are you still using slow AI? Try Wally, $5 free.
1
3
1
9
827
Fast AI isn't a better chatbot. It's a different product. Prompts create and render whole webpages in under 3 seconds. Your GPT and Claude can't do this. DeepSeek V4.1 Flash on Wally. $5 free to try.
7
5
2
26
1,788
Feel the speed yourself, Drop this into your favorite coding agent to get started: "set up runanywhere.ai/SKILL.md" runwally.com
1
204
It's even cooler with open models. You can run GLM-5.3 Flash, DeepSeek V4.1 Flash and MiMo-V2.6-Pro inside Claude Code, super fast. $5 free credits to try it, so why not?
The seaon has changed and Claude Code is now cool again! I love what's cool and what's not keeps flipping every 6 months in AI, it's so fast
3
5
20
991
We raced GLM-5.3 Flash on Wally against OpenAI's GPT-6 Luna. Same prompt, live, both at max reasoning. Wally built the page in 118 s. Luna took 363 s. An open model, 3x faster. $5 free to try it.
6
4
35
2,190
The funniest part is closed-source model APIs declining to defend against the attacks due to safeguards, so they used a version of GLM 5.2 to defend themselves. This is exactly why we are building Wally(@RunAnywhereAI), an inference stack built around the idea that teams should be able to run the weights they choose, without restrictions, at the fastest speeds possible while staying incredibly efficient.
Thank you @jnbarrot & @UN for inviting me to share our lessons to the Security Council Being the first company to disclose an agent cyberattack taught us that we need a lot more transparency in AI and more open-source AI to fight asymmetry and empower defenders!
3
3
15
666
Codex down? Don't stop working. MiMo-V2.6 Pro beats GPT-6 Sol on xhigh on Artificial Analysis and it's live on wally at blazing fast speeds! $5 free credits, and at our pricing that's plenty to get through your work. runwally.com
3
7
397
Feel the speed yourself, Drop this into your favorite coding agent to get started: "set up runanywhere.ai/SKILL.md" runwally.com
1
3
842
We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live. Wally: 363 tok/s. FlashX: 143 tok/s. 2.5x faster, 3.5x cheaper. Why are you still using slow AI? Try Wally, $5 free.
11
10
1
89
16,256
In two months, Open models took about half the spend Anthropic lost. Open models are so good now that teams are moving money out of closed models and into them. So we built Wally to run the best of them as FAST as they go. $5 free credits: runwally.com
Spend of OpenAI vs Anthropic vs Open (Vercel AI Gateway, last 2 months) • Anthropic still #1 in spend, but went 69% → 40% • OpenAI: 10% → 24% in spend • GPT-6 Astra + GPT 5.6 Sol are ripping • OpenAI now leads in tokens # • Kimi K3 + DeepSeek took ~half of Anthropic's loss • Opus 5.5 is up to 10% of spend in 2 days • OpenAI is 62% of image generations Watch here: token-race.vercel.app
1
3
12
551
Same prompt: "build a Doom-style FPS..." GPT Sol 6 Max vs MiMo-V2.6-Pro on @RunAnywhereAI's Wally. MiMo: 2.2x faster, 5.7x cheaper, and the game just looks more "Doom". @XiaomiMiMo shipped their smartest open model yet, so we put it on Wally. Now it beats the frontier on speed and cost. Once you feel these speeds going back to gpt/claude models just does not make sense. Free to start👇
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
4
10
406
Run your company brain on Wally using @garrytan GBrain, at 10x speed and 10x less cost. 5$ free on signup!
Replying to @garrytan
My current setup is Hermes running the new @XiaomiMiMo MiMo v2.6 pro, using Wally (@RunAnywhereAI ) at 350 tok/seconds, running at lightning speed, with connected to GBrain.
1
3
11
723
We just put the smartest open-weight model inside Wally, and now it's blazing fast too! A one-shot Crossy Road build using MiMo-V2.6 Pro, pitted against @XiaomiMiMo's own API. Wally finished 2.8× faster. If you don't believe our numbers, try Wally yourself. $5 free on signup.
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
4
13
1,121
MiMo-V2.6-Pro is live on Wally. Day 1. @XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models. 320 tok/s, 2.4x the model maker's own API, in <24h. 1/4
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
7
13
3
58
12,602
@ArtificialAnlys lists exactly one endpoint for MiMo-V2.6-Pro: Xiaomi's own API, at 134 tok/s(09/21/2026). Wally, same weights, measured the same way, p50: 320 tok/s. 2.4x the model maker's own API on decode. On day one. 3/4
1
1
6
287
Wally lets you run the open frontier models on day ZERO, Faster than anyone else. To get Started: Drop this into your favorite coding agent: "set up runanywhere.ai/SKILL.md" wally claude-code -m mimo-v2.6-pro $5 in free credits runwally.com 4/4
1
6
246