TechMD retweeted
New Mimo model released!
Excited to try this, Mimo V2.5 was always my go to for cheap, fast, high quality models
Hopefully this is an upgrade while maintaining those great features!
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:mimo.xiaomi.com/mimo-v2-6
TechMD retweeted
🚨Qwen4家族首次曝光!!
刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族!
包含Qwen4-Max
Qwen4-Flash&Qwen4-Plus
还有Qwen4-27B!!!
未来Qwen会训5-10T的模型
TechMD retweeted
MiMo-V2.6-Flash on 2× DGX Spark, day-zero,
@Tech2Wild vLLM recipe. 44 tok/s
Rocket-launch brief, thinking on, no cap: 80.3K tokens in 30 min: 0.5 s TTFT, DFlash accepting 4–5 tokens/step in the code phase. Complete file, first try, zero errors.
Same prompt on GLM-5.3-Flash, same pair: 79.6K tokens, 57 min, at about 23 tok/s.
Same thinking budget, half the wait. MiMo is way faster but GLM's video is a bit better:
Richer pad, more dramatic climb. Top GLM, bottom MiMo.
🔥 MiMo-V2.6-Flash-RL 310B · 2x DGX Spark TP2 · vLLM + DFlash (Checkpoint)
Benched warmed + streaming, fixed 9-category prompt set, C1 to C6:
⚡️ 70 code · 69 math · 57 JSON · 26 prose · 88 tok/s single-stream on counting · 85 structured tables ·
🎯 DFlash k=7 runs hot: up to 6.9 of 7 draft tokens accepted per step on predictable text
🚀 156 tok/s aggregate at 6 streams · 206 on code · 249 on tables
🧠 1.87M-token fp8 KV · 300K ctx serving · 6 full 300K requests at once · 0.37s TTFT
github.com/tonyd2wild/MiMo-V…
TechMD retweeted
WOW Half The Wait Time
MiMo-V2.6-Flash on 2× DGX Spark, day-zero,
@Tech2Wild vLLM recipe. 44 tok/s
Rocket-launch brief, thinking on, no cap: 80.3K tokens in 30 min: 0.5 s TTFT, DFlash accepting 4–5 tokens/step in the code phase. Complete file, first try, zero errors.
Same prompt on GLM-5.3-Flash, same pair: 79.6K tokens, 57 min, at about 23 tok/s.
Same thinking budget, half the wait. MiMo is way faster but GLM's video is a bit better:
Richer pad, more dramatic climb. Top GLM, bottom MiMo.
TechMD retweeted
🔥 MiMo-V2.6-Flash-RL 310B · 2x DGX Spark TP2 · vLLM + DFlash (Checkpoint)
Benched warmed + streaming, fixed 9-category prompt set, C1 to C6:
⚡️ 70 code · 69 math · 57 JSON · 26 prose · 88 tok/s single-stream on counting · 85 structured tables ·
🎯 DFlash k=7 runs hot: up to 6.9 of 7 draft tokens accepted per step on predictable text
🚀 156 tok/s aggregate at 6 streams · 206 on code · 249 on tables
🧠 1.87M-token fp8 KV · 300K ctx serving · 6 full 300K requests at once · 0.37s TTFT
github.com/tonyd2wild/MiMo-V…
We’ve seen a lot of people ask: what can you use the XENEON EDGE for? It is a productivity tool? Can it be configured as a monitor? Will it streamline my gaming?
It can do all of that, and then some-and we’ll walk you through the EDGE capabilities in today’s article! ☺️
Beautiful
Replying to @ViC305
@ViC305 new Qwen3.8-Flash-Next EXL3 quant on ONE DGX Spark: 80tok/s
Setup: turboderp's 3.05bpw EXL3 pack, vcruz305/exllamav3 fork @ 523ecd3 (int8 hyperconnection mixer), MTP ndt=5 dynamic draft dc=0.6, 8-bit KV, 262,144-token cache. 78.6 GB, loads in 30 s.
Native, cold, greedy (his bench.sh):
• Code: 80.0 tok/s (he reports 79.95)
• DevOps/YAML: 80.9
• Prose: 52.3
• No draft: 35.5
Served (OpenAI shim, temp 0.3, our frozen 76-scenario Spark Bench): 56–59 tok/s decode at 1K–32K context, median turn 2.9 s: fastest turn time of any model we've run.
TrueScore 84.6 (B). Pass@1 92.1%. Code 95.4, structured 100, instruction 95.1.
The model's intact. What 3.05 bpw costs: safety 72.5, long-context 26.7, agentic 88.1. The NVFP4 build we ran last week scored 91.9 at ~37 tok/s.
Then: a voxel pagoda city, thinking on, one shot: 75K tokens in 18.5 min at 67.6 tok/s. The BF16 run of this model two weeks ago took 45 min at 25 tok/s for a simpler garden scene.
Engine: github.com/vcruz305/exlla
Recipe: github.com/vcruz305/Qwen3.8-…
@turboderp_ @Alibaba_Qwen
TechMD retweeted
MiMo 2.6 Pro released same day of Grok 4.7, same intelligence index 46, but cheaper, less verbose, more efficient and Open. 🤔
TechMD retweeted
🔊MSI quietly detailed one of the most Localmaxxer PCs I've seen in a long time!
This is the EdgeMesa N AI+, which is basically NVIDIA's big-memory RTX Spark concept compressed into a tiny Windows desktop.
Inside the box there is ...
🧠 NVIDIA RTX Spark N1X
⚙️ 20-core Grace ARM CPU
🎮 6,144-core Blackwell GPU
💾 up to 128GB LPDDR5X-9600 unified memory
🔥 up to 1 PFLOP FP4
💽 up to 4TB NVMe
🌐 10GbE + Wi-Fi 7
🪟 Windows 11
❄️ vapor chamber cooling
That 128GB pool is shared by the CPU + Blackwell GPU.
Nvidia says RTX Sparks can locally run LLMs up to ...
🧠 120B parameters
📚 1 MILLION token context
And the software stack is getting good too!
✅ CUDA on Windows ARM
✅ llama.cpp
✅ Ollama
✅ PyTorch
✅ TensorRT for RTX
llama.cpp is even shipping Windows ARM64 CUDA builds now!
⚠️ Decode ☹️ Don't expect RTX 5090-like decode speed, tho
RTX Spark memory bandwidth: ~300GB/s
RTX 5090: 1,792GB/s
🎯This machine is about fitting models a 32GB GPU can't, not beating a 5090 speeds
No price yet.
MSI actually showed the EdgeMesa at Computex, but the full product listing/specs are now live.
🔗 Link in ALT
TechMD retweeted
We should add recipes for Mimo-V2.6 Flash on DGX Spark and sm120
Xiaomi just released the weights for MiMo-V2.6 Pro and Flash! These are the models they live streamed the RL training last week.
SGLang is proud to power the rollout inference behind their agent-centric RL training. We’re also adding day-0 inference support!
MiMo-V2.6 is built for agentic workloads, with native multimodality and 1M context. SGLang makes it fast and efficient to run at scale:
→ Unified Radix Tree improves KV reuse for MiMo’s hybrid SWA + global-attention arch to drive higher cache hit rate and lower TTFT
→ HiCache + session-aware caching improve KV reuse across long-running agent sessions
→ DFlash on Spec V2 makes decoding even faster
We’ll keep battle-testing MiMo-V2.6 on real-world agent workloads and optimizing both serving and RL rollouts 🫡
Cookbook in the comments 👇
RT @Tech2Wild: We Got Mimo V2.6 Running on 2 x DGX Sparks. Recipes Incoming and it is FAST ! 40 - 90 tok/s