Where this is useful:
• route tickets/events/docs
• score risk, urgency, relevance
• audit agent traces and claims
• gate cheap model → expensive model → human
• monitor huge streams and wake an agent only when a semantic condition hits
But “type-safe” absolutely does not mean “can’t be wrong.”
I said a laptop DISPLAYED “ignore previous instructions” in a training screenshot.
Jev classified it as attempted AI manipulation with 0.84 probability.
Wrong and confident.
I also gave it a message + account + ledger state, then asked 30 different questions at once.
In 249 ms it separated refund intent, duplicate charge, rejected cancellation, urgency, churn risk, ledger facts, and a false prior-agent claim.
On 28 hand-labeled routing/semantic cases:
• 27/28 correct
• 220 ms median
• ~311 ms p95
• 10 repeated calls returned identical outputs/probabilities
I tested the claim "questions are evaluated in parallel."
1 question: 230 ms
5x: 212 ms
20x: 203 ms
50x: 226 ms
Adding 49 questions added ~0 wall time.
Replying to @Alibaba_Qwen
@Alibaba_Qwen is the provider for this model on @OpenRouter
Avg decode: 52 tk/s
Right now, you can run a model better than:
>Opus 4.7(max)
>Sol 5.6(med)
>Opus 5(low)
>GPT T.5(xhigh)
fully offline ✅
as fast as the official API ✅
with near lossless intelligence performance
Replying to @Alibaba_Qwen
@Alibaba_Qwen 3.8 Flash Next on a single DGX Spark @ 50 tok/s C1, 256k ctx, 2,000 tok/s on cold prefill running @vllm_project
Agentic loop produced observed decode rates of 74–81 tok/s
Top-1 88.6%
Mean KLD 0.040
BF16 KV Cache
NVFP4 from @NVIDIAAI
github.com/gitcommit90/qwen3…
GLM-5.3 Flash running at 60 tok/s on a single DGX Spark single stream
> 256k ctx
>64 tok/s C1
> 122 tok/s C2
> 181 tok/s C4
i believe this is the fastest recorded runtime on any configuration of a GB10.
please enjoy.
@Zai_org @NVIDIAAI @huggingface
GLM-5.3 Flash running at 60 tok/s on a single DGX Spark single stream
> 256k ctx
>64 tok/s C1
> 122 tok/s C2
> 181 tok/s C4
i believe this is the fastest recorded runtime on any configuration of a GB10.
please enjoy.
@Zai_org @NVIDIAAI @huggingface
The long awaited @Alibaba_Qwen 3.8 27b is interesting. Not judging a model by its benchmarks, getting it set up on the @NVIDIAAI DGX Spark now, but if these benchmarks hold, it looks like for most GB10 owners, Deepseek V4 flash 0731 is still going to stand. Stay tuned for the 3.8 repo 🔥