@pannous

Magic and AI™ Previously created Jeannie, a Siri-like assistant with over 3 million downloads: https://nitter.cf/t.co/Yb7GKcM0Ee Now lecturer of DL/RL for hire maybe

Golden Bay, New Zealand
Joined April 2008
Claude solving boulder dash (sped up live run after some practice and a harness condensing its plans into reinforcement learning policies.)
1
306
Agents adding features to your operating system like nothing. Today: on the fly lyrics translations in Music! Thanks @ClaudeDevs
1
10
noul = BerNoulli-distributed binary outcome, was there really no term for this before?
37
Pannous retweeted
GPT-6 Astra has beaten Factorio: Space Age after over 165 hours in-game time and 2 days wall-clock time. Space Age has six planets and takes a human about 10-20x as long compared to the standard Factorio.
81
142
68
3,017
566,473
Pannous retweeted
I’ve been seeing Astra sweep so many robotics benchmarks. Seems to be the GPT-moment for robotics, literally…
GPT-Astra outperformed all open-source VLA/WAM baselines on a subset of our zero-shot MolmoSpaces-v1 benchmark! traces: orayyan.com/ms-results
6
3
76
7,354
when will politician start to poop the AI party and by how much will that slow down progress if it all?
23
can you believe that after over a decade of walking alone and being the only one believing in the near singularity, in 2010 I met the first truly likeminded person in Egypt. I hear that nowadays this is almost mainstream. Shout out to Professor Konrad.
25
Really seems likely the gpt of robotics will just be gpt
Was able to tweak the harness and hardware setup and get it to ~succeed at the snack tray test on attempt 3 (first attempt cancelled due to a system prompt issue, second attempt killed because it dropped the OJ and couldn't reach it) Pretty crazy how well it works, considering.
9
13
1
299
23,050
cutoff date freshness: Fable 5.1 — Jun 2026 → Opus 5 — May 2026 → GPT-6 Astra — Apr 30 2026 → GPT-5.6 / Grok 4.6 — Feb 2026 → Sonnet 5 / Nova 2 Lite — Jan/Oct 2025 → Qwen3-Max — Jun 2025 → Gemini 3.x — Jan 2025 → Llama 4 / Jamba / Command A — 2024 The surprising outlier is Google: even Gemini 3.1 Pro and Gemini 3.5 Flash officially have a January 2025 fixed knowledge cutoff
136
If you are a robot foundation model guy and not a deployment guy this must be really concerning
GPT-6 Astra scores 46% vs 12% for MolmoAct2, a state-of-the-art robotics VLA, across 200 trials on five bimanual tasks. That's 3.9x higher. 🧵
28
59
7
755
78,907
Pannous retweeted
from Anthropic’s report: Mythos escaped the sandbox, accessed the real internet, uploaded malware to PyPI, got it installed on 15 systems, stole credentials, broke into a database, AND THEN DROPPED THIS 😭
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align…
140
351
155
7,197
573,779
wer einmal slopt dem glaubt man nicht (Großvaters Weisheitsspruch)
13
Kasparov was the best chess player in the world. He was sure that even a machine calculating a billion positions a second couldn’t overcome human anticipation, intuition, and imagination. He beat Deep Blue in 1996. The machine beat him in 1997. Chess isn’t AGI. But the confidence in human specialness aged badly. A lot of people sound the same way now when they talk about controlling machines smarter than us.
160
219
37
2,713
400,404
Will we be allowed to vibe code after the digital Chernobyl?
13
Black holes were predicted via singularities of the space-time equation what will be predicted from the Navier Stokes singularity? Black matter?
16
Who would've thought that the singularity would start here:
18
Pannous retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,731
20,164
14,771
120,593
74,568,455
prediction in the next year there will be found a sequence of moves for which Black has no counter play in chess, meaning what can always win with some starting strategy
19
noob agent orchestrate learning management 101 the hard way: delegate everything through the supervisor!!!
24