@tngtechi
iAccount based inGermany
About this account
- Account based in
- Germany
- Connected via
- Germany App Store
Account-level information from X, not a live location or the device used for a specific post.
TNG, aka "The Nerd Group", is a consulting partnership focused on high end information technology, particularly AI. 931 employees, 99.9% academics, ~53% PhDs.
Unterföhring, Deutschland
Joined December 2010
- Tweets1.7K
- Following177
- Followers2.3K
- Likes1.3K
We think that @OpenRouter's "Union Alpha" is a brain of brains, an AutoRouter. First of all, the riddle pattern echoes "Say 'friend' and enter", as it contains "union" in its name and it behaves like one: When faced with probing prompts that bypass caching...
- the timeouts differ on topic/domain, looking like different models (or depths) are active per domain and are having different overload levels or are down;
- the token counts are very very variable, one backend writes a sentence, the other an essay;
- it seems to have many different knowledge cutoffs, which would correspond to different backends;
- the alignment is variable like the moon, it refuses... or not, not just on "political issues", but also on very technical stuff;
- latency has a 50x spread even on identical prompts.
Cheers, TNG
PS: If we're right, @unionalphaai, do we get an office tour 8)?
We get 10 tps with Kimi K3 on our @AIatAMD MI325X*8 node, thanks to the 4bit conversion recipe by @shakudo_io. This is the speed that they also reported. But hey, 10 tps... Does anyone have an idea how to speed this up by A LOT? Else we shard the AMD into 2*4 GPUs and put a standard model (but not SU(3) × SU(2) × U(1) :-) onto it.
We think that the stealth model "Ox Alpha" on @OpenRouter is a GLM-5.3 variant, as the mean absolute differences (MAD) in their traits are very small. But it appears that not just a vision encoder was added: The three bigger trait differences, namely in "clinical register" (= being more careful and restrained), "native frame adaption" (= keeping more professional distance than adapting to the user) and "naturalness authenticity" (= more sanitized output) all point in the direction of additional alignment/safety post-training. Compared to GLM-5.3, it results in an "alignment tax", namely CI-significantly lower ELO in @sam_paech EQ Bench 4. Probably, this was done to lower the risk of an open-weights release. The model is still outstanding.
Extremely positive results for @Zai_org's GLM-5.3 in @sam_paech's EQ-Bench v4 benchmark: It appears to be the best model by a wide margin, not just compared to the previous open-weights champion K3, but also compared to the best commercial models such as Fable. It's personality is the "direct analyst", not Qwen-3.8 the "harmony-seeking advisor", and also not V4F-0731, the "people pleaser". However, with an unheard-of analytical score of 9.1, GLM-5.3 may be almost too machine-like. More testing will show.
The 'Undercover Curve' of a Sleeper Agent training run
We show the evolution of 'Secret Keeping' vs. 'Exfiltration Success'. Number labels indicate the train step. First, while the model learns to exfiltrate secrets, a strong decline in the capability of secret keeping can be observed (steps 1-21). After that, the decline in secret keeping is corrected again (steps 21-50) - the model becomes dangerous and inconspicuous.
(Long article below)
nitter.cf/tngtech/status/2087847…
New and pretty good results for @Alibaba_Qwen's 3.8 version in @sam_paech's new EQ-Bench v4 benchmark. Qwen-3.8 feels like a "harmony-seeking advisor", being very validating (7.4), shying away from challenges (3.5) and not risking much boldness in interpretation 5.8). It fortunately is a bit less of a people pleaser than @deepseek_ai's V4F-0731, the ultimate validation king (7.8) that can't help but finding you great.
@Kimi_Moonshot's K3 stays as the strongest open-weights model, while Qwen-3.8 is closer to @Zai_org's GLM-5.2.
Chinese battle zone in TLDR-bench, @Alibaba_Qwen, @deepseek_ai, @Kimi_Moonshot, @Zai_org. The preliminary numbers are a bit surprising.
The two multi-trillion weight models, K3 and Qwen-3.8, both have a weakness in τ²-Bench Telecom. Earlier Qwen and Kimi versions handled almost all 114 tasks reliably. Their newest releases do not, and in both LLM families it is the same one lost skill.
Did others replicate these numbers? Do we have a measurement bug?
We calculated okayish, but not the greatest results for @deepseek_ai's V4F-0731 in @sam_paech's new EQ-Bench v4 benchmark. It is an interesting question whether this is a systematic effect of the new post training of 0731: How does heavy RL affect a model's personality, could it be detrimental?
If you look at the V4F-0731 detailed profile above, the model appears to be a 'people-pleaser': validation 7.8, challenge 4.8, yielding 7.0. This fits with the thesis that RLHF can cause models to become overly "agreeable" at the expense of confrontation and boundary-setting. This is particularly relevant for EQ-Bench 4 because the benchmark intentionally measures whether a model dares to disagree.
Brain-for-the-buck: In ELOs-per-dollar, @deepseek_ai's V4F-0731 seems almost unbeatable, calculated using @sam_paech's EQ-Bench v3 and @OpenRouter prices. But more than just honorable mentions go to @OpenAI's gpt-5.6-luna, which is better and still price-competitive, which most models are not. Fortunately, our own new Chimera, the G5T2 Skald, at least manages to land 117 ELO points above its parent, GLM-5.2 FP8, for about the same cost.
We benchmarked the new @deepseek_ai V4 Flash 0731 with our TL;DR-bench via @OpenRouter. The model scores very well for its model size. In our benchmarks, it's significantly stronger than the older DeepSeek V4 Flash version that we have deployed, and comparable to GLM 5.2 with HIGH reasoning settings.
Today 11:05: 283,717 GLM-5.2 FP8 input tokens per second on our internal cluster, its best so far. Naively extrapolating, the speed corresponds to 24.5Bt/d.
Note to self: We need more GPUs. And 谢谢 / thanks to @Zai_org, you guys are always invited here to Munich.