@nanulledi
iAccount based inEurope!
About this account
- Account based in
- Europe
- Connected via
- Eastern Europe (Non-EU) Android App
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
longtermism
United States
Joined October 2019
- Tweets3.4K
- Following66
- Followers2.9K
- Likes25.6K
nano retweeted
Z ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task
@Zai_org’s GLM-5.2 is the same size as GLM-5.1 (744B total / 40B active parameters) but scores 11 points higher on the Intelligence Index v4.1, placing ahead of MiniMax-M3 (44) and DeepSeek V4 Pro (max, 44). On the first-party API it is priced in line with GLM-5.1 at $1.4/$4.4/$0.26 per 1M input/output/cache hit tokens
Key results:
➤ GLM-5.2 is the leading open weights model on the Intelligence Index v4.1. At 51, it leads MiniMax-M3 (44), DeepSeek V4 Pro (max, 44) and Kimi K2.6 (43)
➤ Improvements across most evaluations, particularly scientific reasoning: GLM-5.2 gains over GLM-5.1 on most evaluations, led by scientific reasoning on CritPt (+16 points to 21%) and HLE (+12 points to 40%), alongside AA-LCR (+9 points to 71%), tau3 banking (+15 points to 27%) and SciCode (+7 points to 50%). TerminalBench v2.1 also improves (+16 points to 78%) and GPQA Diamond gains 3 points to 89%
➤ Leading open weights model on GDPval-AA v2 and competitive with proprietary models: GLM-5.2 scores 1524 on GDPval-AA v2, ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro (max, 1328). This impressive result places GLM-5.2 in-line with proprietary models including GPT-5.5 (xhigh reasoning). GDPval-AA v2 builds on the original GDPval-AA by baselining Elo to human performance at 1000, introducing a rotating panel of frontier-model judges, and raising the turn limit from 100 to 250 for longer-horizon agent trajectories
➤ GLM-5.2 uses more output tokens per task than other leading open weights models: the model uses 43k output tokens per Intelligence Index task, up from GLM-5.1 (26k) and above MiniMax-M3 (24k), Kimi K2.6 (35k) and DeepSeek V4 Pro (max, 37k)
➤ On the Intelligence vs. Cost per Task Pareto Frontier: GLM-5.2 is on the Pareto frontier of the Intelligence vs Cost per Task chart, with the lowest cost per task among models at its intelligence level. GLM-5.2 costs ~$0.46 per task, compared to GLM-5.1 ($0.25), Kimi K2.6 ($0.31), MiniMax-M3 ($0.18) and DeepSeek V4 Pro (max, $0.05)
Additional Model Details:
➤ License: MIT
➤ Size: 744B total parameters, 40B active parameters, equivalent to GLM-5.1
➤ Context window: 1M tokens, up from 200K on GLM-5.1
➤ Pricing: $1.4/$0.26/$4.4 per 1M input/cache hit/output tokens
➤ Availability: Alongside Z ai's first-party API, GLM-5.2 is available across third-party providers including @DeepInfra, @novita_labs, @nebiusai, @parasailnetwork , @SiliconFlowAI , @gmi_cloud , @Baseten and @FireworksAI_HQ
Introducing GLM-5.2: Frontier Intelligence, Open Weights
- Significant improvements in coding and agentic tasks
- Strong long-horizon capabilities with a 1M context window
- Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong balance between performance and token efficiency
- MIT-licensed open weights
- Same API pricing as GLM-5.1
Tech Blog: z.ai/blog/glm-5.2
Weights: huggingface.co/zai-org/GLM-5…
API: docs.z.ai/guides/llm/glm-5.2
Coding Plan: z.ai/subscribe
Chat: chat.z.ai
nano retweeted
We somehow got put in the spotlight the last few days! First we'd like to thank the organizers of the AI show for that, we can't get enough of this stuff. I'll say a few things about where we are and what we do.
nano retweeted
In light of Anthropic’s policy decision, I am withdrawing my amicus brief signature. I can’t truthfully argue they’re not a supply chain risk. 😞
Replying to @paulmarin90
I’ll be honest that it would have been much more difficult to defend Anthropic against the DoW incursion had that incident occurred after this one. This is the company literally telling their customers, “we reserve the right to silently sabotage you.” I’d still have defended them, because the government trying to destroy a firm is still wrong, but man would it have been a harder case to make.
nano retweeted
if I'm Dario, the reasoning for an IPO is to get flush with liquidity and then pay for the most insane lobbying effort the US has ever seen
i'm calling it now, come back to this in 6 months
Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap: darioamodei.com/post/policy-…
nano retweeted
Fable is an early look into the bureaucratic hell Anthropic wants to create in AI.
3 month delay to get released.
Randomly routed in biology, cyber and distillation.
Weird sociopathic quiet harmfulness in ml topics.
Goes away on June 22.
Have to apply to get access to 'real' Mythos.
What an exhausting tiresome dystopia.
nano retweeted
just so you guys are noticing this; they will pull the ladder from above you as soon as they can. their intentions are to disempower you as much as they reasonably can. the only reason they have given you anything at all is because openai has forced them to
nano retweeted
Motion for AI researchers: DO NOT evaluate and report results on Fable 5 models.
Closeness hurts science, and gated access will further destroy it. Embrace open source models.
nano retweeted
New Anthropic research: Natural Language Autoencoders.
Models like Claude talk in words but think in numbers. The numbers—called activations—encode Claude’s thoughts, but not in a language we can read.
Here, we train Claude to translate its activations into human-readable text.
nano retweeted
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at chat.deepseek.com via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report: huggingface.co/deepseek-ai/D…
🤗 Open Weights: huggingface.co/collections/d…
1/n
nano retweeted
We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks.
nano retweeted
Step inside Project Genie: our experimental research prototype that lets you create, edit, and explore virtual worlds. 🌎
nano retweeted
We’re announcing a major advance in the study of fluid dynamics with AI 💧 in a joint paper with researchers from @BrownUniversity, @nyuniversity and @Stanford.
nano retweeted
What if you could not only watch a generated video, but explore it too? 🌐
Genie 3 is our groundbreaking world model that creates interactive, playable environments from a single text prompt.
From photorealistic landscapes to fantasy realms, the possibilities are endless. 🧵
Gemini 2.5 Deep Think Model Card:
it's not superhuman but similar to gold IMO model & “approaches human level” on stealth evals
more interested in learning “novel rl techniques that can leverage more multi-step reasoning,” (candidate: MARL with verification/voting for each step)
The new stealth model, the Horizon Alpha, has the ability to think in cot, but you really have to try to get it to do so.
Here is the COT it generated. It's very terse, and I see some O3 in its writing.
I think it's safe to say it's an OpenAI open-source model.
Benchmarks like Humanity’s Last Exam, codeforces nerd-sniped researchers and could prevent AI labs from developing genuine AGI capable of performing real-world tasks.
Yes there would be differences in taste and preferences but a horrible game/software can be seen and recognized by the majority. Vague Objectives would be set for ai to complete and most humans would be able to verify if said objectives were achieved fully or partially
good to see hle uselessness being confirmed after ~3 months of this thread
actually, it's not just useless it's harmfull signal that somewhat slowed down the progress imo
nitter.cf/andrewwhite01/status/1…
HLE has recently become the benchmark to beat for frontier agents. We @FutureHouseSF took a closer look at the chem and bio questions and found about 30% of them are likely invalid based on our analysis and third-party PhD evaluations. 1/7
GDM achieved the same score at IMO as OpenAI but it will be accessible to ultra subscribers in a trusted program.