@LeTechLead

Local AI - Software Engineer - (https://nitter.cf/t.co/B1ewJUrT6s) - ([email protected]) - مهندس برمجيات

vmbr0
Joined December 2020
🤖 Made with AI
1
7
8,364
most things are better with less people. traveling
13
ddr4 sodimm to dimm adapters skyrocketed in price btw it doubled in 2 months
26
burgundy before iphone 18 E655系
1
69
today i learned a lot about oculink sff and daisy chaining 😂
44
i asked if their dspark drafter gonna work with fp8 checkpoints and they said yes. but it didnt, might be skill issue which i dont mind confessing 😜
Now stack it with quantization. Pair the speculator above with our NVFP4 checkpoint: a 4-bit, Blackwell-native target, with DSpark's faster decoding on top. Serve RedHatAI/Qwen3.8-27B-NVFP4 with the DSpark speculative-config (method dspark, 8 tokens). huggingface.co/RedHatAI/Qwen…
1
198
my beloved deepseek 🥹
🚨 DeepSeek V5 Leak: Beats Astra >DeepSeek is reportedly preparing an imminent V5 launch >Founder Liang Wenfeng calls it the company's biggest bet yet >Rumored at 2 trillion parameters (not 3T) >Reportedly the first DeepSeek model to train fully on Huawei Ascend chips instead of Nvidia >Needs roughly 4x more Ascend accelerators than the Nvidia equivalent to hit the same training scale DeepSeek is reportedly keeping the open-weight strategy
2
123
back to the origins.. qwen3.8-27b-fp8 mtp max-ctx 2x3090 vllm-v0.30 kv-fp8
93
benchmarked qwen3.8-27b DAS gsq-rco iq2s no-drafter ctx-64k 12gb-3080ti acceptable
1
124
atm, qwen3.8-27b fp8 seems the best you can use on 2x3090 with full ctx. for single 3090, DAS quants are the best and depending on how much ctx you need you can use iq2s up to iq3s variants. AP quants from @laurent_zw are interchangeable with DAS quants if you're willing to a/b test their quality for your case. that's it for 1/2x 3090 owners.
1
3
185
i just pulled the trigger 💸
1
1
96
i tried to slash my backlog yesterday and as a result i slashed my sleep schedule
83
we need proper support for hermes in the token plan @Alibaba_Qwen currently i frequently see <think> tags in the 3.8-flash model. i never see that with my deepseek api pls fix @alibaba_cloud
78
and with MTP 👇 qw3.8-27b gsq-rco iq3_s ctx-182k single 3090
qwen3.5-27b gsq-rco (claimed best quant) full ctx-ish on single 3090 benchmarked
1
217
i tried to re-produce the quoted, and here are my numbers (gsq-rco iq2_xs mtp 12gb-3080ti max 47k ctx)
Qwen3.8-27B, four IQ2-class quants, 20 three.js scenes each, RTX 3060 12GB. Every scene generated locally, then opened in a headless browser to check it actually renders. GSQ-RCO 14/20 bartowski 8/20 Unsloth 8/20 AtomicChat 7/20 Across all four, only 46% of the 80 scenes rendered at all. That's what the video is -- me scrolling the whole gallery, blanks included. GSQ-RCO wins and it's still nowhere near 20/20. GSQ-RCO is also the second smallest file of the four (~7.8 GiB vs AtomicChat's 9.0), so this isn't a size story. It costs ~25k tokens a task where bartowski averages 5k, and 5.6x the wall clock. All four ran on low reasoning effort, not the xhigh default. At 2 bits, medium and xhigh couldn't stop thinking -- they'd spiral until the loop killer stepped in. If you're running 2-bit quants, use low effort.
2
147
I can’t believe how miserable we were before Herdr came on.
2
90
qwen3.5-27b gsq-rco (claimed best quant) full ctx-ish on single 3090 benchmarked
1
1
294
vs MTP
ive benchmarked ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF on my 12gb 3080ti tight fit.
72
ive benchmarked ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF on my 12gb 3080ti tight fit.
1
1
148
opus5.5 seems to have trained in a male-only gem
1
83
buy me coffee ❌ buy me ddr5 64gb 5600 (2x32) cl40 ✅
1
74