@benchrouteri
iAccount based inCanada
About this account
- Account based in
- Canada
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
https://nitter.cf/t.co/XPD48XczuB Pin a route, not a model. Automagic model routing
Joined August 2026
- Tweets60
- Following16
- Followers2
- Likes19
BenchRouter retweeted
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet
these models enable creators, developers, and enterprises to create richer, more expressive audio experiences
try them via the Gemini API and in AI Studio: aistudio.google.com/generate…
If you are still on GPT-5.6 Sol for coding agents, you probably want to move to GPT-6 Sol
Half the token price, stronger merge-ready coding on FrontierCode, and Astra-era alignment at Sol cost
Benchrouter can run your evals automatically, so you can migrate models with confidence
OPUS DROP
@AnthropicAI shipped Claude Opus 5.5 today
Fable 5.1-level coding and agents at $4/$20 MTok, about 40% cheaper per task than Opus 5. API id claude-opus-5-5, 1M context
this is huge - try out MiMo 2.6
benchrouter.com/frontier
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:mimo.xiaomi.com/mimo-v2-6
BenchRouter retweeted
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:mimo.xiaomi.com/mimo-v2-6
BenchRouter retweeted
Replying to @benchrouter
Automated evals are a great way to make agentic work less anecdotal. The useful comparison is not just the model score, but the full task cost and reliability across runs.
Frontier update
Xiaomi MiMo-V2.6-Pro just took a chunk of the live frontier at $0.54 per M tokens
46.3 intelligence index, 95 tok/s, on the three-axis surface at benchrouter.com/frontier
Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
Grok 4.7 xHigh is priced the same as Grok 4.6 High was - try it out on your complex agentic tasks
Benchrouter can automatically run all of your tasks evals, give it a try today
BenchRouter retweeted
Introducing Step 5 Preview: Advancing the Pareto Frontier.
Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.
- 600B total / 27B active MoE, with 1M context + Vision
- Substantially lower task cost at comparable intelligence
- Broad software engineering capabilities with sustained execution over long horizons
Try Step 5 Preview: platform.stepfun.ai
Model page: stepfun.com/step-5-preview
Open weights on Oct 15.
MODEL DROP
AliceAI-Foundation-80B-A3B-Base open weights are on Hugging Face today
Yandex trained this MoE from scratch, 80B total / 3B active. 262K context, Apache 2.0
Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
If you are still on SPLADE-v3 for sparse first-stage retrieval, you probably want to try SPARSEUP
Benchrouter can run your evals automatically, so you can migrate models with confidence
BenchRouter retweeted
Here's a 45-second TL;DR on Jev.
I find the core idea beautifully simple, but the video made it really hard to understand.
Hope you find it helpful.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
TypeSafe shipped Jev in early access today
First System One model. Typed decisions instead of text, vendor ~100x faster on structured tasks at $0.042 per MTok input
Have you evaluated it on your tasks yet? benchrouter.com can automatically swap for you
BenchRouter retweeted
Replying to @benchrouter
haven't tried that one. on our engine server end-of-turn floors at about 1.3 seconds, driving it from our own vad gets to about 1.2 seconds.
Model drop alert!
Gemini 3.8 Live Extended Thinking is live on the Gemini API today
Google's high-reasoning live voice model. #1 on Artificial Analysis Speech to Speech, reasons and speaks at the same time
Have you evaluated it on your voice tasks yet? benchrouter.com can automatically swap for you
StepFun shipped StepAudio 3 Realtime
Full duplex voice. Thinks while it talks. Tops Artificial Analysis Full-Duplex Bench at 98.9
Have you evaluated it on your voice agent tasks yet?
benchrouter.com can automatically swap for you
Salesforce + NVIDIA dropped a model
If you are routing CRM agents through Claude or GPT inside Salesforce, Koa is the new in-house option
Post-trained Nemotron for multistep CRM tool use. Pilots now, US GA winter 2026
Easy skip unless you live in Agentforce
huggingface.co/internlm/Atri…
If you are still on GLM-5.3 or DeepSeek V4 for long agent loops, you probably want to look at Atria Dawn Preview.
Shanghai AI Lab, MIT open weights, BrowseComp 92.5 on their card.
Benchrouter can run your evals so you can migrate with confidence.