@benchrouter

https://nitter.cf/t.co/XPD48XczuB Pin a route, not a model. Automagic model routing

Joined August 2026
Opus 5.5 is a full replacement for Fable 5.1, and took some ground from Muse Spark 1.3
1
17
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet these models enable creators, developers, and enterprises to create richer, more expressive audio experiences try them via the Gemini API and in AI Studio: aistudio.google.com/generate…
180
422
186
3,539
981,923
If you are still on GPT-5.6 Sol for coding agents, you probably want to move to GPT-6 Sol Half the token price, stronger merge-ready coding on FrontierCode, and Astra-era alignment at Sol cost Benchrouter can run your evals automatically, so you can migrate models with confidence
1
36
OPUS DROP @AnthropicAI shipped Claude Opus 5.5 today Fable 5.1-level coding and agents at $4/$20 MTok, about 40% cheaper per task than Opus 5. API id claude-opus-5-5, 1M context
1
31
this is huge - try out MiMo 2.6 benchrouter.com/frontier
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
1
2
26
BenchRouter retweeted
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
376
867
475
8,796
1,442,940
Replying to @benchrouter
Automated evals are a great way to make agentic work less anecdotal. The useful comparison is not just the model score, but the full task cost and reliability across runs.
1
1
1
17
Frontier update Xiaomi MiMo-V2.6-Pro just took a chunk of the live frontier at $0.54 per M tokens 46.3 intelligence index, 95 tok/s, on the three-axis surface at benchrouter.com/frontier Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
1
31
Grok 4.7 xHigh is priced the same as Grok 4.6 High was - try it out on your complex agentic tasks Benchrouter can automatically run all of your tasks evals, give it a try today
1
1
16
BenchRouter retweeted
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
1,517
3,117
1,913
28,680
20,536,210
BenchRouter retweeted
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: platform.stepfun.ai Model page: stepfun.com/step-5-preview Open weights on Oct 15.
200
280
163
1,877
435,107
MODEL DROP AliceAI-Foundation-80B-A3B-Base open weights are on Hugging Face today Yandex trained this MoE from scratch, 80B total / 3B active. 262K context, Apache 2.0 Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
1
31
If you are still on SPLADE-v3 for sparse first-stage retrieval, you probably want to try SPARSEUP Benchrouter can run your evals automatically, so you can migrate models with confidence
16
BenchRouter retweeted
Here's a 45-second TL;DR on Jev. I find the core idea beautifully simple, but the video made it really hard to understand. Hope you find it helpful.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
239
723
280
9,799
1,742,870
TypeSafe shipped Jev in early access today First System One model. Typed decisions instead of text, vendor ~100x faster on structured tasks at $0.042 per MTok input Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
23
BenchRouter retweeted
Replying to @benchrouter
haven't tried that one. on our engine server end-of-turn floors at about 1.3 seconds, driving it from our own vad gets to about 1.2 seconds.
1
1
8
Model drop alert! Gemini 3.8 Live Extended Thinking is live on the Gemini API today Google's high-reasoning live voice model. #1 on Artificial Analysis Speech to Speech, reasons and speaks at the same time Have you evaluated it on your voice tasks yet? benchrouter.com can automatically swap for you
1
16
StepFun shipped StepAudio 3 Realtime Full duplex voice. Thinks while it talks. Tops Artificial Analysis Full-Duplex Bench at 98.9 Have you evaluated it on your voice agent tasks yet? benchrouter.com can automatically swap for you
2
25
Salesforce + NVIDIA dropped a model If you are routing CRM agents through Claude or GPT inside Salesforce, Koa is the new in-house option Post-trained Nemotron for multistep CRM tool use. Pilots now, US GA winter 2026 Easy skip unless you live in Agentforce
2
28
huggingface.co/internlm/Atri… If you are still on GLM-5.3 or DeepSeek V4 for long agent loops, you probably want to look at Atria Dawn Preview. Shanghai AI Lab, MIT open weights, BrowseComp 92.5 on their card. Benchrouter can run your evals so you can migrate with confidence.
34