@LeTechLeadi
iAccount based inSaudi Arabia
About this account
- Account based in
- Saudi Arabia
- Connected via
- Saudi Arabia App Store
Account-level information from X, not a live location or the device used for a specific post.
Local AI - Software Engineer - (https://nitter.cf/t.co/B1ewJUrT6s) - ([email protected]) - مهندس برمجيات
vmbr0
Joined December 2020
- Tweets2.9K
- Following953
- Followers1.3K
- Likes1.4K
i asked if their dspark drafter gonna work with fp8 checkpoints and they said yes. but it didnt, might be skill issue which i dont mind confessing 😜
Now stack it with quantization. Pair the speculator above with our NVFP4 checkpoint: a 4-bit, Blackwell-native target, with DSpark's faster decoding on top.
Serve RedHatAI/Qwen3.8-27B-NVFP4 with the DSpark speculative-config (method dspark, 8 tokens).
huggingface.co/RedHatAI/Qwen…
my beloved deepseek 🥹
🚨 DeepSeek V5 Leak: Beats Astra
>DeepSeek is reportedly preparing an imminent V5 launch
>Founder Liang Wenfeng calls it the company's biggest bet yet
>Rumored at 2 trillion parameters (not 3T)
>Reportedly the first DeepSeek model to train fully on Huawei Ascend chips instead of Nvidia
>Needs roughly 4x more Ascend accelerators than the Nvidia equivalent to hit the same training scale
DeepSeek is reportedly keeping the open-weight strategy
atm, qwen3.8-27b fp8 seems the best you can use on 2x3090 with full ctx.
for single 3090, DAS quants are the best and depending on how much ctx you need you can use iq2s up to iq3s variants.
AP quants from @laurent_zw are interchangeable with DAS quants if you're willing to a/b test their quality for your case.
that's it for 1/2x 3090 owners.
we need proper support for hermes in the token plan @Alibaba_Qwen
currently i frequently see <think> tags in the 3.8-flash model.
i never see that with my deepseek api
pls fix @alibaba_cloud
i tried to re-produce the quoted, and here are my numbers (gsq-rco iq2_xs mtp 12gb-3080ti max 47k ctx)
Qwen3.8-27B, four IQ2-class quants, 20 three.js scenes each, RTX 3060 12GB. Every scene generated locally, then opened in a headless browser to check it actually renders.
GSQ-RCO 14/20
bartowski 8/20
Unsloth 8/20
AtomicChat 7/20
Across all four, only 46% of the 80 scenes rendered at all. That's what the video is -- me scrolling the whole gallery, blanks included. GSQ-RCO wins and it's still nowhere near 20/20.
GSQ-RCO is also the second smallest file of the four (~7.8 GiB vs AtomicChat's 9.0), so this isn't a size story. It costs ~25k tokens a task where bartowski averages 5k, and 5.6x the wall clock.
All four ran on low reasoning effort, not the xhigh default. At 2 bits, medium and xhigh couldn't stop thinking -- they'd spiral until the loop killer stepped in. If you're running 2-bit quants, use low effort.