@RunAnywhereAIi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Fastest inference anywhere: open models on-prem, hosted or on-device. Backed by @ycombinator W26
Joined July 2025
- Tweets444
- Following6
- Followers2.2K
- Likes1K
Pinned Tweet
Meet Wally.
Our inference stack for open frontier models, built to be the fastest place to run them.
Performance snapshot:
GLM-5.3 Flash: 380 tok/s.
GLM-5.3 Max: 790 tok/s
Qwen3.8-27B: 485 tok/s
DeepSeek-V4.1 Flash: 615 tok/s
1/7
RunAnywhere retweeted
Why does AI make you wait?
Every response is two phases: prefill reads the whole prompt in one pass, then decode writes one token at a time.
Agents pay for both on every step and re-read the context each time, so prefill keeps growing.
Wally (@RunAnywhereAI) makes both ultra fast.
RunAnywhere retweeted
One-shot drum machine with MiMo-V2.6-Pro on Wally
One prompt, under 2 minutes, and it built a playable drum machine with every sound made in code.
MiMo is the top open model on Artificial Analysis, and on Wally it runs miles faster than Xiaomi's own API.
And it's free for everyone through tomorrow!👇
MiMo-V2.6-Pro is the #1 open-weights model on Artificial Analysis.
On Wally it streams at 320 tok/s, 7x faster than Xiaomi's own API.
The best open model at this speed should be in everyone's hands. So we partnered with @explabsai to make it completely free until 9/29.
Go try it now!
The pelican test, but it's a race.
Same model: MiMo-V2.6-Pro, the top open model on Artificial Analysis. Wally vs Xiaomi's own API.
Wally finished 3.6x faster, and drew all four pelicans before Xiaomi's API finished its first.
And it's completely free to use through tomorrow, on Wally via @explabsai.
free Mimo 2.6 Pro at 320 tokens/sec
currently the best open weights model on the market
cheaper and 50% better than GPT Luna
only possible in partnership with @RunAnywhereAI
platform.xplabs.ai
MiMo-V2.6-Pro is the #1 open-weights model on Artificial Analysis.
On Wally it streams at 320 tok/s, 7x faster than Xiaomi's own API.
The best open model at this speed should be in everyone's hands. So we partnered with @explabsai to make it completely free until 9/29.
Go try it now!
free Mimo 2.6 Pro at 320 tokens/sec
currently the best open weights model on the market
cheaper and 50% better than GPT Luna
only possible in partnership with @RunAnywhereAI
platform.xplabs.ai
RunAnywhere retweeted
Every AI call makes you wait twice: for the first token, then for the rest.
Agents pay that 50 times a task.
Wally is the fastest place to run open models. Faster than the labs' own APIs, even their paid fast tiers.
Try it now, free to start.
We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live.
Wally: 363 tok/s. FlashX: 143 tok/s.
2.5x faster, 3.5x cheaper.
Why are you still using slow AI?
Try Wally, $5 free.
Fast AI isn't a better chatbot. It's a different product.
Prompts create and render whole webpages in under 3 seconds.
Your GPT and Claude can't do this.
DeepSeek V4.1 Flash on Wally.
$5 free to try.
Feel the speed yourself, Drop this into your favorite coding agent to get started:
"set up runanywhere.ai/SKILL.md"
runwally.com
RunAnywhere retweeted
It's even cooler with open models.
You can run GLM-5.3 Flash, DeepSeek V4.1 Flash and MiMo-V2.6-Pro inside Claude Code, super fast.
$5 free credits to try it, so why not?
We raced GLM-5.3 Flash on Wally against OpenAI's GPT-6 Luna. Same prompt, live, both at max reasoning.
Wally built the page in 118 s. Luna took 363 s.
An open model, 3x faster.
$5 free to try it.
RunAnywhere retweeted
The funniest part is closed-source model APIs declining to defend against the attacks due to safeguards, so they used a version of GLM 5.2 to defend themselves.
This is exactly why we are building Wally(@RunAnywhereAI), an inference stack built around the idea that teams should be able to run the weights they choose, without restrictions, at the fastest speeds possible while staying incredibly efficient.
RunAnywhere retweeted
Codex down? Don't stop working.
MiMo-V2.6 Pro beats GPT-6 Sol on xhigh on Artificial Analysis and it's live on wally at blazing fast speeds!
$5 free credits, and at our pricing that's plenty to get through your work.
runwally.com
Feel the speed yourself, Drop this into your favorite coding agent to get started:
"set up runanywhere.ai/SKILL.md"
runwally.com
We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live.
Wally: 363 tok/s. FlashX: 143 tok/s.
2.5x faster, 3.5x cheaper.
Why are you still using slow AI?
Try Wally, $5 free.
Faster GLM-5.3-Flash is now live: up to 200 tokens/s. Model code: glm-5.3-flashx.
Priced at 2.5× GLM-5.3-Flash on both the Coding Plan and API.
Open to all API users. Coding Plan users can apply here:
docs.google.com/forms/d/e/1F…
RunAnywhere retweeted
In two months, Open models took about half the spend Anthropic lost.
Open models are so good now that teams are moving money out of closed models and into them.
So we built Wally to run the best of them as FAST as they go.
$5 free credits: runwally.com
Spend of OpenAI vs Anthropic vs Open
(Vercel AI Gateway, last 2 months)
• Anthropic still #1 in spend, but went 69% → 40%
• OpenAI: 10% → 24% in spend
• GPT-6 Astra + GPT 5.6 Sol are ripping
• OpenAI now leads in tokens #
• Kimi K3 + DeepSeek took ~half of Anthropic's loss
• Opus 5.5 is up to 10% of spend in 2 days
• OpenAI is 62% of image generations
Watch here: token-race.vercel.app
RunAnywhere retweeted
Same prompt: "build a Doom-style FPS..."
GPT Sol 6 Max vs MiMo-V2.6-Pro on @RunAnywhereAI's Wally.
MiMo: 2.2x faster, 5.7x cheaper, and the game just looks more "Doom".
@XiaomiMiMo shipped their smartest open model yet, so we put it on Wally. Now it beats the frontier on speed and cost.
Once you feel these speeds going back to gpt/claude models just does not make sense.
Free to start👇
MiMo-V2.6-Pro is live on Wally. Day 1.
@XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models.
320 tok/s, 2.4x the model maker's own API, in <24h.
1/4
Run your company brain on Wally using @garrytan GBrain, at 10x speed and 10x less cost.
5$ free on signup!
Replying to @garrytan
My current setup is Hermes running the new @XiaomiMiMo MiMo v2.6 pro, using Wally (@RunAnywhereAI ) at 350 tok/seconds, running at lightning speed, with connected to GBrain.
RunAnywhere retweeted
We just put the smartest open-weight model inside Wally, and now it's blazing fast too!
A one-shot Crossy Road build using MiMo-V2.6 Pro, pitted against @XiaomiMiMo's own API.
Wally finished 2.8× faster.
If you don't believe our numbers, try Wally yourself.
$5 free on signup.
MiMo-V2.6-Pro is live on Wally. Day 1.
@XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models.
320 tok/s, 2.4x the model maker's own API, in <24h.
1/4
MiMo-V2.6-Pro is live on Wally. Day 1.
@XiaomiMiMo's new open-weights flagship, top of the @ArtificialAnlys Intelligence Index for open-weight models.
320 tok/s, 2.4x the model maker's own API, in <24h.
1/4
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:mimo.xiaomi.com/mimo-v2-6
@ArtificialAnlys lists exactly one endpoint for MiMo-V2.6-Pro: Xiaomi's own API, at 134 tok/s(09/21/2026).
Wally, same weights, measured the same way, p50: 320 tok/s.
2.4x the model maker's own API on decode. On day one.
3/4
Wally lets you run the open frontier models on day ZERO, Faster than anyone else.
To get Started:
Drop this into your favorite coding agent:
"set up runanywhere.ai/SKILL.md"
wally claude-code -m mimo-v2.6-pro
$5 in free credits
runwally.com
4/4