@paul_cali
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- China Android App
Account-level information from X, not a live location or the device used for a specific post.
AI is good & bad, actually. Tweeting about AI/ML methods, software dev, research, tech and society, social impact. 20yrs in tech, 10 in ML/AI, PhD in comp sci
London, England
Joined August 2013
- Tweets7.4K
- Following5K
- Followers6.6K
- Likes45K
Pinned Tweet
The story of LLMs playing games, and what we know so far
Tic Tac Toe, Chess, Minecraft, NYT Connections, Wordle, Pictionary, Connect 4, Codenames, Snake... 1/n
Bug bounties are small because there's overwhelmingly more people that want to bug hunt as a ~legal & stable career vs those willing to actually break the law and/or sell to bad actors. World is full of normies & this is a good thing
Paul Calcraft retweeted
I have a Continuous Learning benchmark where models attempt to learn to play chess. They are given a /goal of learning and improving playing against a Stockfish opponent in 200 games. They can choose the difficulty, take notes, whatever they like - except cheating (e.g. using a chess engine) of course.
So far the improvement in Elo has been negative for Astra. Tiny bit positive for Opus, but could also be random. I've started Astra off sooner, so it finished its 200 games already, Opus is still playing.
Site here to watch how they are doing: ai-learning-to-play-chess.su…
The reason why this is interesting is that while we obviously don't have continuous learning, at the back of my mind I was thinking that maybe models can simulate it through self-scaffolding. Turns out not so much at least in this context. Perhaps it's a solvable problem and we don't need 'true' self-learning for models to learn in some way.
-------
Just a note, the idea for the benchmarks belongs to someone else, but I don't want to use their name to give this more weight without permission.
Modding will stay niche. Publishers will continue to make sure of that. But this is still a amazing time for modding & the niche of enthusiasts around it
Valorant had it 6 yrs ago but *not* an obvious win. Implementation was so expensive it *doubled* server frame time (128 tick server down to 64). And when you make it approximate/predictive enough to avoid sudden client side popping, you lose a lot of the anti hack benefit
Remember BridgeMind is not a reliable source
Paul Calcraft retweeted
"the Waymo effect is what happens when a technology removes the friction of dealing with another human being" researchagenda.news/articles…
Moving from "did you use Claude for this?" (derogatory) to "did you use Claude for this?" (mandatory)
Replying to @aden_barton
On these detail-oriented, well-specified tasks, frontier models are now faster and more accurate than accountants, even the best one in our study.
Ok Simon Says is a *brilliant* video interaction benchmark
Replying to @tavus
Watch Brittney and Mars try to trip Griffin up at Simon Says. To play, it has to listen for "Simon says," watch their hands and move its own, all at the same time. It generates the whole frame, not just a face.
An internet proxy is for proxying the internet. But it's important to know that a meat proxy is not for proxying meat
Paul Calcraft retweeted
Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Short-term good for the consumer but if new ideas need to capture all their value in 2 weeks we have an incentives problem
Opus 5.5 keeps getting stuck on bad waits. Multiple times I've had whole 5+ hr workstreams paused bc 1 intermediate job crashed or hung. I've had to prompt it to schedule a session check in every 20 mins
This basically never happened in Codex 5.5+, it got so good at babysitting
Amazing: magic questions you can ask LLMs to get a read on some latent state. Like mech interp probes but for blackbox models
Spurious probes use adaptive search at the API level to find random questions whose answers differ systematically based on underlying model state
"you can do better than that" is still an unreasonably effective follow up prompt. It seems almost contentless and tbh I thought we'd be done with that by now
Paul Calcraft retweeted
Replying to @andrew_n_carr
What are you talking about? Confirmation bias is the best, I keep seeing new evidence that supports it
When Eliezer retweeted this earlier I thought he'd missed the joke. But there was no joke