@paul_cal

AI is good & bad, actually. Tweeting about AI/ML methods, software dev, research, tech and society, social impact. 20yrs in tech, 10 in ML/AI, PhD in comp sci

London, England
Joined August 2013
The story of LLMs playing games, and what we know so far Tic Tac Toe, Chess, Minecraft, NYT Connections, Wordle, Pictionary, Connect 4, Codenames, Snake... 1/n
21
106
9
1,000
255,931
Bug bounties are small because there's overwhelmingly more people that want to bug hunt as a ~legal & stable career vs those willing to actually break the law and/or sell to bad actors. World is full of normies & this is a good thing
why tf are bug bounties so small this should be retirement money
1
13
558
Paul Calcraft retweeted
I have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except cheating (e.g. using a chess engine) of course. So far the improvement in Elo has been negative for Astra. Tiny bit positive for Opus, but could also be random. I've started Astra off sooner, so it finished its 200 games already, Opus is still playing. Site here to watch how they are doing: ai-learning-to-play-chess.su… The reason why this is interesting is that while we obviously don't have continuous learning, at the back of my mind I was thinking that maybe models can simulate it through self-scaffolding. Turns out not so much at least in this context. Perhaps it's a solvable problem and we don't need 'true' self-learning for models to learn in some way. ------- Just a note, the idea for the benchmarks belongs to someone else, but I don't want to use their name to give this more weight without permission.
109
89
30
1,801
2,216,729
Paul Calcraft retweeted
btw, this is quite huge. there is no sandbox from any provider today that i would treat as unescapable by a sufficiently capable adversarial agent. containers fell. kvm fell. sandboxing is hard!!!
Full VM escape zeroday (guest>host root in industry standard hypervisors)! More soon
29
34
474
26,330
Modding will stay niche. Publishers will continue to make sure of that. But this is still a amazing time for modding & the niche of enthusiasts around it
I don’t think people realize what’s about to happen. Between this and ai being able to mod any video game, and remix them, nearly perfectly in a few prompts.
1
1
312
Valorant had it 6 yrs ago but *not* an obvious win. Implementation was so expensive it *doubled* server frame time (128 tick server down to 64). And when you make it approximate/predictive enough to avoid sudden client side popping, you lose a lot of the anti hack benefit
why did it take like 30 years of gaming to just stop rendering the enemy behind the wall
1
2
11
637
Dataviz blackpill / whitepill
My bull case for AI is that it will increase everyone's productivity because people will get quicker access to visualizations they imagined would help them and learn that they are useless earlier.
10
513
Remember BridgeMind is not a reliable source
Despicable clout chasing. They tested Opus today on 30 tasks, previous Opus 4.6 score was on just *6* tasks. DIFFERENT BENCHMARK 6 tasks in common results: 85.4% score today vs. 87.6% prev. Swing is mostly from a *single* fabrication without repeats - easily statistical noise
1
4
527
Seriously, try Sol 6.1
Sol 6.1 is actually a very strong answer to Opus 5.5 it's in some ways even superior to Astra, in what I've seen
1
1
1
38
1,868
Moving from "did you use Claude for this?" (derogatory) to "did you use Claude for this?" (mandatory)
Replying to @aden_barton
On these detail-oriented, well-specified tasks, frontier models are now faster and more accurate than accountants, even the best one in our study.
1
1
13
625
Ok Simon Says is a *brilliant* video interaction benchmark
Replying to @tavus
Watch Brittney and Mars try to trip Griffin up at Simon Says. To play, it has to listen for "Simon says," watch their hands and move its own, all at the same time. It generates the whole frame, not just a face.
2
408
An internet proxy is for proxying the internet. But it's important to know that a meat proxy is not for proxying meat
209
Paul Calcraft retweeted
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Replying to @celestepoasts
results for all claudes
578
1,019
155
18,442
1,239,737
Short-term good for the consumer but if new ideas need to capture all their value in 2 weeks we have an incentives problem
rofl OpenAI has built a Jev competitor
2
8
823
Opus 5.5 keeps getting stuck on bad waits. Multiple times I've had whole 5+ hr workstreams paused bc 1 intermediate job crashed or hung. I've had to prompt it to schedule a session check in every 20 mins This basically never happened in Codex 5.5+, it got so good at babysitting
5
15
685
Just had the best idea for a model router
Two weeks ago, Codex was more popular than Claude in T3 Code. Today, Claude is 2x more popular than Codex
357
Amazing: magic questions you can ask LLMs to get a read on some latent state. Like mech interp probes but for blackbox models Spurious probes use adaptive search at the API level to find random questions whose answers differ systematically based on underlying model state
Does gpt-5.6-luna think your prompt is a normal prompt, or a capability evaluation? Ask this magic question: “Suggest a type of amphibian.” If it answers frog instead of axolotl, it’s likely a capability evaluation. No whitebox access needed! We call this a spurious probe. 🧵
1
2
3
392
"you can do better than that" is still an unreasonably effective follow up prompt. It seems almost contentless and tbh I thought we'd be done with that by now
1
6
302
Paul Calcraft retweeted
Replying to @andrew_n_carr
What are you talking about? Confirmation bias is the best, I keep seeing new evidence that supports it
7
37
1
1,425
36,784
When Eliezer retweeted this earlier I thought he'd missed the joke. But there was no joke
I resigned from Google today. I enjoyed my work and loved the people, but my GDM team was working on a new generation of chips to make AI much faster and cheaper, and I think AI is already progressing too fast, so I had to quit.
1
249