@pmbstuff

Full Stack AI Engineer & Researcher. Building AI-powered tools.

Canada
Joined July 2010
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
Qwen3.8-Flash-Next - 4.1v๐Ÿฅณ ๐Ÿš€ Peak c=1 @ 118 tok/s ๐Ÿ“Š Average c=1 @ 80 tok/s ๐Ÿญ Peak c=64 @ 834 tok/s โš™๏ธ Prefill 3,233 tok/s ๐Ÿ’ป TTFT 0.42s โšก๏ธ github.com/myllmbox/qwen38-fโ€ฆ
6
9
85
7,188
Expected
Deleted Muse after seeing this post on Threads about how it told some Facebook Marketplace sellers the guyโ€™s address and they showed up at his door Dangerous and creepy This would have been 1000x worse if the person was a woman
4
Agree, FB reinvented OpenClaw for normal people. Gave it a furry ass and called it a day. And there is a thing: normal people do not ask these kinds of questions - hey, can I put my model in it? They use. Sharing their data with FB, getting more and more locked inside the ecosystem. That is exactly what Mark wants. Tbh, every major player on the market wants the same. As a result, you don't own your data or your life anymore. And the word "freedom" sounds very different these days.
I think I speak for many when I say, the people want freedom of choice. You've built a great harness, computer-use agent, and supporting backend - but people don't want model lock-in. Grokbot cockblocked itself the same way and it's just not as powerful as running something like Astra/Fable/Opus.
27
I smell fear. Muse looks like a crab's butthole, btw.
11
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
New acceleration method for Minimax H3: Veda Sparse huggingface.co/Veda-Sparse/Mโ€ฆ
9
11
1
197
13,237
Yes, more of that please!
A lot of people are saying Anthropic has already nerfed Claude Opus 5.5. We're launching NerfBench on BridgeBench tomorrow. We have the day 1 results. Tomorrow morning we show you the retest. Is Claude Opus 5.5 nerfed or not?
3
Oh, an Anthropic "Harry Potter" was a PR experiment all the way. Who can predict that? (irony)
Former Anthropic researcher Jacob Coxon became a media superstar after going public with his AI fears. He insisted that he wasnโ€™t working with any third party organizations. Familiar sources told us a different story: DEY., a PR firm representing many of the most prominent AI safetyists, was booking his interviews. One source, who had direct knowledge, even said DEY. preemptively booked Nate Soares, a prominent AI safety figure, for interviews that directly overlapped with Jacob going public. Jacob working with DEY. is notable for two reasons: first, as mentioned, he previously said he wasnโ€™t working with third parties. Second, we are in the middle of a national conversation about the future of AI that is actively determining how we regulate the most powerful technology in the world, largely thanks to the panic stirred up by Jacob โ€” and itโ€™s in the publicโ€™s interest to know who, exactly, is behind it. Scoop from @huntryerson ๐Ÿ‘‡
1
5
o1-preview just 2 years ago? omg, it felt like an eternity
Two years ago today in AI: Artificial Analysis reported on OpenAI pushing the intelligence frontier with o1-preview, the first reasoning model. Now, all frontier models use reasoning tokens to โ€˜thinkโ€™ before answering Two years ago, v1 of the Artificial Analysis Intelligence Index measured four single-turn, exam-style evaluations - MMLU, GPQA, MATH, and HumanEval - covering general knowledge, science, mathematics, and basic coding. Today, the Intelligence Index v4.3 incorporates 10 difficult evaluations which include long-horizon agentic tasks, challenging coding problems, and knowledge work.
6
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
Introducing Julia-1: Our first classification model that runs on almost anything. Learn more ๐Ÿ‘‡ supersoniclabs.ia.br/julia-1โ€ฆ
107
274
97
3,219
290,237
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
๐Ÿšจ OpenAI Set to Reveal a Long-Term Agent at DevDay - Codenamed "Aeon" โ€” built for long-running tasks, similar to Grok Bot or Manus, working for hours, days, even weeks - Likely built on Astra, already strong at long-horizon work - Runs in a cloud environment like Cursor โ€” sets everything up remotely and keeps grinding until the task's done - OpenAI already has the infra (hosted sandboxes, multi-agent workflows) to make this the natural next step - Rumored to support multiple agents collaborating on the same task โ€” if real, that's a big deal
76
157
76
2,210
238,467
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
Weโ€™re launching the Army of Robots. In 2022, we launched the Army of Drones. Today, drones account for over 95% of battlefield strikes, and Ukraine has more than 700 UAV manufacturers. Now we need the next technological breakthrough: the robotization of warfare. The goal is simple โ€” save lives. Robots should take on the most dangerous missions: evacuating the wounded, delivering ammunition, mining and demining, reconnaissance, defending positions and engaging targets. The Army of Robots is not one company. Itโ€™s an ecosystem. We will invest in defense tech companies, launch our own technology projects, test them with the military and scale what works on the battlefield. Weโ€™re now looking for defense tech companies and engineers working on robotic technologies โ€” as well as a CTO / Tech Lead for the Army of Robots. Join us: thearmyofrobots.com/en
889
1,782
784
11,076
2,883,159
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
Qwen3.8-Flash-Next on solo GB10 just got better. Totally new optimized PLE offload. Allows to get ~20% more speed ๐Ÿš€ and we again have lots of KVโ—๏ธ c=1 -> 82 token/s โ—๏ธ c=16 -> 318 token/sโ—๏ธ โšก๏ธgithub.com/bilikaz/qwen38-flโ€ฆ
Qwen3.8-Flash-Next @ sustainable ~50โ€“51 tok/s for code ๐Ÿคฏ While for thinking and code it gets ~42 tok/s ๐Ÿš€ v1 had good moments. v2 has speed, and it keeps it. โšกgithub.com/bilikaz/qwen38-flโ€ฆ
19
16
1
198
25,669
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
RTX PRO 6000 96GB + DGX Spark owners rejoice! ๐Ÿ”ฅ You can now run the highest quality Xiaomiโ€™s MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process ๐Ÿš€ 309B total parameters. Only ~15B active per token. Xiaomiโ€™s published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5. And on ONE RTX PRO 6000: 184.1 tok/s p50 with DFlash 49.6 tok/s without drafting 2,258+ tok/s prefill 321.9 tok/s confirmed aggregate @ C=8 ๐—ฃ๐—œ๐—–๐—ž ๐—ฌ๐—ข๐—จ๐—ฅ ๐—–๐—”๐—ฅ๐—— RTX PRO 6000 96GB: 2.20 bpw SAGE-EXL3 86.94 GB DGX Spark 128GB: 2.50 bpw SAGE-EXL3 98.48 GB The recipe includes separate profiles for both systems. The 184.1 tok/s result is specifically from the RTX PRO 6000; Spark DFlash gains are currently lower and more prompt-dependent. ๐—™๐—œ๐——๐—˜๐—Ÿ๐—œ๐—ง๐—ฌ ๐—ฅ๐—˜๐—–๐—˜๐—œ๐—ฃ๐—ง๐—ฆ Across two completely untouched holdout sets: 2.20 bpw: 82.16โ€“82.64% top-1 agreement 0.1926โ€“0.1986 mean KLD 2.50 bpw: 83.76โ€“87.30% top-1 agreement 0.1086โ€“0.2055 mean KLD Top-1 agreement means the EXL3 pack selected the SAME most-likely next token as Xiaomiโ€™s reference. KLD compares the entire next-token probability distribution. Lower means the quant tracks the reference more closely. ๐—”๐—ก ๐—œ๐— ๐—ฃ๐—ข๐—ฅ๐—ง๐—”๐—ก๐—ง ๐——๐—˜๐—ง๐—”๐—œ๐—Ÿ Xiaomi does not publish a full BF16 expert checkpoint. The 303B routed-expert bank already ships in MXFP4. Attention ships in block-FP8, with embeddings, norms and the output head in BF16. So these fidelity numbers compare against Xiaomiโ€™s actual released mixed-precision checkpoint running through its official implementation, not against a hidden full-BF16 teacher. This is effectively EXL3 agreement with the best public source that exists. ๐—›๐—ข๐—ช ๐—–๐—Ÿ๐—ข๐—ฆ๐—˜ ๐—œ๐—ฆ ๐—™๐—Ÿ๐—”๐—ฆ๐—› ๐—ง๐—ข ๐—ง๐—›๐—˜ ๐—™๐—ฅ๐—ข๐—ก๐—ง๐—œ๐—˜๐—ฅ? Xiaomiโ€™s published numbers: AutomationBench: MiMo Flash 52.3 Claude Opus 5 50.3 GPT-5.6 Sol 45.8 Terminal-Bench 2.1: MiMo Flash 87.6 Claude Opus 5 89.1 GPT-5.6 Sol 88.8 OSWorld: MiMo Flash 80.8 Claude Opus 5 83.4 GPT-5.6 Sol 83.0 VisualCoding: MiMo Flash 71.5 Claude Opus 5 70.0 GPT-5.6 Sol 73.4 Artificial Analysis has not scored Flash independently yet. The larger MiMo-V2.6-Pro sibling scores 46 on the AA Intelligence Index, tied with Grok 4.7 and sitting beside GLM-5.3 at 45. This release is the text backbone. The original DFlash speculative drafter is repaired and wired through the recipe; vision and audio towers are not included yet. Huge credit to @XiaomiMiMo team for the model and to @turboderp_ / ExLlamaV3 for the EXL3 engine and format. @XiaomiMiMoDevs EXL3 weights + fidelity results: huggingface.co/vcruz305/MiMoโ€ฆ Complete serving recipe: github.com/vcruz305/MiMo-V2.โ€ฆ
5
10
1
72
5,519
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
This guy was getting ~15 tps running Qwen3.8-Flash-Next on his 12GB RTX 5070. Apparently that wasn't good enough. ๐Ÿ˜‚ So he built his own inference engine. (as we all should) Now he's reporting ๐ŸŒ llama.cpp โ†’ ~15 tps ๐Ÿš€ Strata โ†’ up to 65.1 tps And this isn't big workstation either Specs ๐ŸŽฎ RTX 5070 12GB ๐Ÿง  64GB DDR5-5600 โš™๏ธ Ryzen 5 7600 ๐ŸชŸ Windows At 128K context his new Strata engine reports: Q2_0 โ†’ 65.1 tps + 543 tps prompt IQ2_XS โ†’ 52.0 tps + 472 tps prompt IQ3_XXS โ†’ 44.8 tps + 414 tps prompt For Qwen3.8-FLASH-NEXT. ๐Ÿ‘€ He built it specifically around this model and this kind of CUDA + system-RAM setup, paired with RCO-GSQ quants. ๐Ÿ‘‰ And he open-sourced it. I love these kind of Local AI projects. ๐Ÿ”— Link in ALT
68
114
15
1,619
101,001
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
We have converted GLiNER2.5-Decide to coreml. ~4ร— faster, ~5ร— less Peak RAM, half the size. model: huggingface.co/FluidInferencโ€ฆ code: github.com/FluidInference
Introducing GLiNER2.5-Decide, our new 340M parameter open weight, encoder-based decision model. GLiNER2.5-Decide is built for fast, deterministic classification. The model evaluates a set of user-defined typed questions and rules, and jointly decodes their answers, returning structured decisions with probability distributions and confidence scores. We evaluated the modelโ€™s performance on Fast Decisions, an unseen, internally generated classification suite based on 17 datasets testing real-world use cases across routing, triage, classification, sentiment, and content understanding. Measured against similar decision models, GLiNER2.5-Decide leads in 9 of the 17 datasets, achieving the highest average score: - GLiNER2.5-Decide: 60.1% - SemIf: 56.4% - JevK5: 57.5% - Laya: 46.6% This performance makes the model a strong fit for use cases like tool calling, model routing, browser and computer use, and LLM-as-a-judge. GLiNER2.5-Decideโ€™s lightweight encoder architecture makes it easy to fine-tune the model for specific tasks, while being efficient enough to run locally on consumer-grade CPUs or in air-gapped environments, giving users greater control over where their data goes and where the model runs. To make building and experimenting with GLiNER2.5-Decide as easy as possible, we're also offering hosted inference. You can now use our API to run inference and fine-tune GLiNER models on specialized tasks right inside your own coding agent: agent.fastino.ai As with previous models, weโ€™re also releasing the model weights on @huggingface under the Apache 2.0 license: huggingface.co/fastino/GLiNEโ€ฆ
17
84
7
884
68,277
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
We achieved a breakthrough at BTL. AI models are very linear. They think in one direction at a time, and when they're wrong they scramble and start again. For something meant to replace humans, that's way too human. We changed that. We made LLM reasoning and execution work more like a quantum computer: many paths at once, the ones that meet merge into one, the dead ends cancel out, and everything left moves forward together. On the same hard problems, a 1.7B model thinking the normal way solved 3 out of 30. Interference Search(our new architecture)solved 23. We watched the normal model find the right answer at token 1,313, check it 9 more times, wander off and run out of budget without ever answering. Ours got there in 3 steps. Parallel thinking and execution. Subagents were a terrible way to tackle this. More technical details soon.
86
140
39
1,949
170,593
Ok, here is what I see in Codex rn. On the $200 plan, usage limits are dropping much faster than they did at Astra's launch. Sol is much stupider, even compared with 5.5. With all these rumors about 500 bucks plan, I suspect it's intentional. @thsottiaux are you playing games with us?
16
Slava ๐Ÿ‡จ๐Ÿ‡ฆ โค๏ธ ๐Ÿ‡บ๐Ÿ‡ฆ retweeted
Today weโ€™re announcing OrcaSAQ-2 27B High-fidelity mixed-precision Qwen3.8 for long-horizon agents. 55.59 โ†’ 12.06 GB โ€” 78.3% smaller / 4.61ร— 3.21 bpw ยท 93.2% Top-1 agreement 70.0 SWE-bench Verified 58.4 Terminal-Bench 2.1 262K context A 27B model for coding, terminal, browser, security and multi-tool agents โ€” in a footprint you can actually deploy. SOTA agentic capability density among similarly sized models we evaluated. huggingface.co/orcarouter/Orโ€ฆ
78
159
30
1,847
990,029