@pmbstuffi
iAccount based inCanada!
About this account
- Account based in
- Canada
- Connected via
- Canada Android App
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Full Stack AI Engineer & Researcher. Building AI-powered tools.
Canada
Joined July 2010
- Tweets1.5K
- Following208
- Followers182
- Likes77
Qwen3.8-Flash-Next - 4.1v๐ฅณ
๐ Peak c=1 @ 118 tok/s
๐ Average c=1 @ 80 tok/s
๐ญ Peak c=64 @ 834 tok/s
โ๏ธ Prefill 3,233 tok/s
๐ป TTFT 0.42s
โก๏ธ github.com/myllmbox/qwen38-fโฆ
Agree, FB reinvented OpenClaw for normal people. Gave it a furry ass and called it a day. And there is a thing: normal people do not ask these kinds of questions - hey, can I put my model in it? They use. Sharing their data with FB, getting more and more locked inside the ecosystem.
That is exactly what Mark wants.
Tbh, every major player on the market wants the same.
As a result, you don't own your data or your life anymore. And the word "freedom" sounds very different these days.
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
New acceleration method for Minimax H3: Veda Sparse
huggingface.co/Veda-Sparse/Mโฆ
Oh, an Anthropic "Harry Potter" was a PR experiment all the way. Who can predict that? (irony)
Former Anthropic researcher Jacob Coxon became a media superstar after going public with his AI fears. He insisted that he wasnโt working with any third party organizations. Familiar sources told us a different story: DEY., a PR firm representing many of the most prominent AI safetyists, was booking his interviews.
One source, who had direct knowledge, even said DEY. preemptively booked Nate Soares, a prominent AI safety figure, for interviews that directly overlapped with Jacob going public.
Jacob working with DEY. is notable for two reasons: first, as mentioned, he previously said he wasnโt working with third parties. Second, we are in the middle of a national conversation about the future of AI that is actively determining how we regulate the most powerful technology in the world, largely thanks to the panic stirred up by Jacob โ and itโs in the publicโs interest to know who, exactly, is behind it.
Scoop from @huntryerson ๐
o1-preview just 2 years ago? omg, it felt like an eternity
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to โthinkโ before answering
Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
Introducing Julia-1:
Our first classification model that runs on almost anything.
Learn more ๐
supersoniclabs.ia.br/julia-1โฆ
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
๐จ OpenAI Set to Reveal a Long-Term Agent at DevDay
- Codenamed "Aeon" โ built for long-running tasks, similar to Grok Bot or Manus, working for hours, days, even weeks
- Likely built on Astra, already strong at long-horizon work
- Runs in a cloud environment like Cursor โ sets everything up remotely and keeps grinding until the task's done
- OpenAI already has the infra (hosted sandboxes, multi-agent workflows) to make this the natural next step
- Rumored to support multiple agents collaborating on the same task โ if real, that's a big deal
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
Weโre launching the Army of Robots.
In 2022, we launched the Army of Drones. Today, drones account for over 95% of battlefield strikes, and Ukraine has more than 700 UAV manufacturers.
Now we need the next technological breakthrough: the robotization of warfare.
The goal is simple โ save lives. Robots should take on the most dangerous missions: evacuating the wounded, delivering ammunition, mining and demining, reconnaissance, defending positions and engaging targets.
The Army of Robots is not one company. Itโs an ecosystem. We will invest in defense tech companies, launch our own technology projects, test them with the military and scale what works on the battlefield.
Weโre now looking for defense tech companies and engineers working on robotic technologies โ as well as a CTO / Tech Lead for the Army of Robots.
Join us: thearmyofrobots.com/en
Qwen3.8-Flash-Next on solo GB10 just got better.
Totally new optimized PLE offload. Allows to get ~20% more speed ๐ and we again have lots of KVโ๏ธ
c=1 -> 82 token/s โ๏ธ
c=16 -> 318 token/sโ๏ธ
โก๏ธgithub.com/bilikaz/qwen38-flโฆ
Qwen3.8-Flash-Next @ sustainable ~50โ51 tok/s for code ๐คฏ
While for thinking and code it gets ~42 tok/s ๐
v1 had good moments. v2 has speed, and it keeps it.
โกgithub.com/bilikaz/qwen38-flโฆ
RTX PRO 6000 96GB + DGX Spark owners rejoice! ๐ฅ
You can now run the highest quality Xiaomiโs MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process ๐
309B total parameters. Only ~15B active per token. Xiaomiโs published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5.
And on ONE RTX PRO 6000:
184.1 tok/s p50 with DFlash
49.6 tok/s without drafting
2,258+ tok/s prefill
321.9 tok/s confirmed aggregate @ C=8
๐ฃ๐๐๐ ๐ฌ๐ข๐จ๐ฅ ๐๐๐ฅ๐
RTX PRO 6000 96GB:
2.20 bpw SAGE-EXL3
86.94 GB
DGX Spark 128GB:
2.50 bpw SAGE-EXL3
98.48 GB
The recipe includes separate profiles for both systems. The 184.1 tok/s result is specifically from the RTX PRO 6000; Spark DFlash gains are currently lower and more prompt-dependent.
๐๐๐๐๐๐๐ง๐ฌ ๐ฅ๐๐๐๐๐ฃ๐ง๐ฆ
Across two completely untouched holdout sets:
2.20 bpw:
82.16โ82.64% top-1 agreement
0.1926โ0.1986 mean KLD
2.50 bpw:
83.76โ87.30% top-1 agreement
0.1086โ0.2055 mean KLD
Top-1 agreement means the EXL3 pack selected the SAME most-likely next token as Xiaomiโs reference.
KLD compares the entire next-token probability distribution. Lower means the quant tracks the reference more closely.
๐๐ก ๐๐ ๐ฃ๐ข๐ฅ๐ง๐๐ก๐ง ๐๐๐ง๐๐๐
Xiaomi does not publish a full BF16 expert checkpoint.
The 303B routed-expert bank already ships in MXFP4. Attention ships in block-FP8, with embeddings, norms and the output head in BF16.
So these fidelity numbers compare against Xiaomiโs actual released mixed-precision checkpoint running through its official implementation, not against a hidden full-BF16 teacher.
This is effectively EXL3 agreement with the best public source that exists.
๐๐ข๐ช ๐๐๐ข๐ฆ๐ ๐๐ฆ ๐๐๐๐ฆ๐ ๐ง๐ข ๐ง๐๐ ๐๐ฅ๐ข๐ก๐ง๐๐๐ฅ?
Xiaomiโs published numbers:
AutomationBench:
MiMo Flash 52.3
Claude Opus 5 50.3
GPT-5.6 Sol 45.8
Terminal-Bench 2.1:
MiMo Flash 87.6
Claude Opus 5 89.1
GPT-5.6 Sol 88.8
OSWorld:
MiMo Flash 80.8
Claude Opus 5 83.4
GPT-5.6 Sol 83.0
VisualCoding:
MiMo Flash 71.5
Claude Opus 5 70.0
GPT-5.6 Sol 73.4
Artificial Analysis has not scored Flash independently yet.
The larger MiMo-V2.6-Pro sibling scores 46 on the AA Intelligence Index, tied with Grok 4.7 and sitting beside GLM-5.3 at 45.
This release is the text backbone. The original DFlash speculative drafter is repaired and wired through the recipe; vision and audio towers are not included yet.
Huge credit to @XiaomiMiMo team for the model and to @turboderp_ / ExLlamaV3 for the EXL3 engine and format. @XiaomiMiMoDevs
EXL3 weights + fidelity results:
huggingface.co/vcruz305/MiMoโฆ
Complete serving recipe:
github.com/vcruz305/MiMo-V2.โฆ
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
This guy was getting ~15 tps running Qwen3.8-Flash-Next on his 12GB RTX 5070.
Apparently that wasn't good enough. ๐
So he built his own inference engine. (as we all should)
Now he's reporting
๐ llama.cpp โ ~15 tps
๐ Strata โ up to 65.1 tps
And this isn't big workstation either
Specs
๐ฎ RTX 5070 12GB
๐ง 64GB DDR5-5600
โ๏ธ Ryzen 5 7600
๐ช Windows
At 128K context his new Strata engine reports:
Q2_0 โ 65.1 tps + 543 tps prompt
IQ2_XS โ 52.0 tps + 472 tps prompt
IQ3_XXS โ 44.8 tps + 414 tps prompt
For Qwen3.8-FLASH-NEXT. ๐
He built it specifically around this model and this kind of CUDA + system-RAM setup, paired with RCO-GSQ quants. ๐ And he open-sourced it.
I love these kind of Local AI projects.
๐ Link in ALT
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
We have converted GLiNER2.5-Decide to coreml. ~4ร faster, ~5ร less Peak RAM, half the size.
model: huggingface.co/FluidInferencโฆ
code: github.com/FluidInference
Introducing GLiNER2.5-Decide, our new 340M parameter open weight, encoder-based decision model.
GLiNER2.5-Decide is built for fast, deterministic classification. The model evaluates a set of user-defined typed questions and rules, and jointly decodes their answers, returning structured decisions with probability distributions and confidence scores.
We evaluated the modelโs performance on Fast Decisions, an unseen, internally generated classification suite based on 17 datasets testing real-world use cases across routing, triage, classification, sentiment, and content understanding.
Measured against similar decision models, GLiNER2.5-Decide leads in 9 of the 17 datasets, achieving the highest average score:
- GLiNER2.5-Decide: 60.1%
- SemIf: 56.4%
- JevK5: 57.5%
- Laya: 46.6%
This performance makes the model a strong fit for use cases like tool calling, model routing, browser and computer use, and LLM-as-a-judge.
GLiNER2.5-Decideโs lightweight encoder architecture makes it easy to fine-tune the model for specific tasks, while being efficient enough to run locally on consumer-grade CPUs or in air-gapped environments, giving users greater control over where their data goes and where the model runs.
To make building and experimenting with GLiNER2.5-Decide as easy as possible, we're also offering hosted inference. You can now use our API to run inference and fine-tune GLiNER models on specialized tasks right inside your own coding agent: agent.fastino.ai
As with previous models, weโre also releasing the model weights on @huggingface under the Apache 2.0 license: huggingface.co/fastino/GLiNEโฆ
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
We achieved a breakthrough at BTL.
AI models are very linear. They think in one direction at a time, and when they're wrong they scramble and start again. For something meant to replace humans, that's way too human.
We changed that. We made LLM reasoning and execution work more like a quantum computer: many paths at once, the ones that meet merge into one, the dead ends cancel out, and everything left moves forward together.
On the same hard problems, a 1.7B model thinking the normal way solved 3 out of 30. Interference Search(our new architecture)solved 23.
We watched the normal model find the right answer at token 1,313, check it 9 more times, wander off and run out of budget without ever answering. Ours got there in 3 steps.
Parallel thinking and execution. Subagents were a terrible way to tackle this.
More technical details soon.
Ok, here is what I see in Codex rn.
On the $200 plan, usage limits are dropping much faster than they did at Astra's launch. Sol is much stupider, even compared with 5.5. With all these rumors about 500 bucks plan, I suspect it's intentional.
@thsottiaux are you playing games with us?
Slava ๐จ๐ฆ โค๏ธ ๐บ๐ฆ retweeted
Today weโre announcing OrcaSAQ-2 27B
High-fidelity mixed-precision Qwen3.8 for long-horizon agents.
55.59 โ 12.06 GB โ 78.3% smaller / 4.61ร
3.21 bpw ยท 93.2% Top-1 agreement
70.0 SWE-bench Verified
58.4 Terminal-Bench 2.1
262K context
A 27B model for coding, terminal, browser, security and multi-tool agents โ in a footprint you can actually deploy.
SOTA agentic capability density among similarly sized models we evaluated.
huggingface.co/orcarouter/Orโฆ