Pinned Tweet
New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data.
We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
Daphne Cornelisse retweeted
Worst thing about this is we have no idea if GPT-6 was specifically post trained on Nethack or not. Zero transparency about the training data and environments. You can't independently evaluate the significance of anything without knowing this.
In all seriousness, this is a startling achievement for GPT-6 Astra. kenforthewin.github.io/blog/…
(This is GPT-6 Astra beating Nethack on its 3rd try. Nethack is the original roguelike and one of the most famously hard games of all time. I have played a lot, and I've never ascended)
Daphne Cornelisse retweeted
I'm super bored of LLM-stylized papers/blogs. I think we all are?
So I asked Claude to explain its tells: web.mit.edu/phillipi/www/cla…
Let's avoid these! I'd much rather read *your* words, see *your* rickety figures, than face down yet another of Claude's anodized corporate veneers.
Daphne Cornelisse retweeted
i'd like to keep it low key and say this is no big deal, but in reality this has to be the coolest project i've worked on maybe ever, and there's a lot more to come!
Daphne Cornelisse retweeted
Joseph and team have been pursuing a very different approach to AI compared to the mainstream. They build really fast simulators + optimizers — so fast that you can do RL from scratch, even on hard domains. It’s worth checking out! Very cool results!
Daphne Cornelisse retweeted
5 Exciting new results brought to you by our team + contributors: @spenccheng @finlay_sanders (fintern!) @ValtteriValo @daphnesolves @elliotarledge. Play with the agents at puffer.ai! All experimental data available is available in Constellation online + local.
Daphne Cornelisse retweeted
terrytao.wordpress.com/2026/…
Terry’s essay is interesting because it highlights a misalignment between AI companies and many, if not most, social groups in humanity
Many of our institutions are built around improving humanity’s understanding of how the world works and how to create things within it
AI companies like OpenAI are more concerned with producing artifacts that have historically been challenging for people to create
They don’t really care about helping people understand the world
On the one hand, its a fair point that if we can solve a key problem, like finding a cure for cancer, it shouldn’t matter if we understand the solution
on the other hand, this is a deeply antisocial approach to discovery
and one that is increasingly leaving the world — artists, writers, now mathematicians — with animosity towards AI
Rather than build AI as a tool to solve arbitrary problems, we could build it as a tool to empower people as they learn and perform tasks in the world
Some companies, eg @percepta, are taking this approach. I hope more will follow suit
A little late, but happy to share that spiced self-play was accepted to the Conference on Robot Learning (CoRL)! 🎉
Sweden, here we come! 🇸🇪
I’m heading to ECCV, where we’ll host our workshop this Tuesday: emerging-ad.github.io/
We have a fantastic lineup of speakers; see you there!
Daphne Cornelisse retweeted
world models offer new affordances for safety — like simulated counterfactuals to diagnose safety failures.
In this work, @Mingxuan0422 discovered we can use divide and conquer to scale counterfactual debugging to 1M steps
Daphne Cornelisse retweeted
Excited to share that our paper on explainable AI for autonomous driving is out today in @Nature !
Super proud to have been part of this amazing collaboration between @motionaldrive and @MIT_CSAIL!
Details in the 🧵
An interesting read on the virtual cell.
Daphne Cornelisse retweeted
Reward hacking has been in the news a lot lately, but AI researchers have seen surprising examples of it since long before LLMs. We're excited to share “AI Finds a Way,” led by @_aadharna , which brings many of these stories together in one place.
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖
AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential.
We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw
Four favorites:
1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function!
2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI!
3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely.
4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory!
See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it.
A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna
Paper: arxiv.org/abs/2608.23875
Daphne Cornelisse retweeted
A good book, a warm cup
of coffee, and a little silence
- sometimes that's all the soul
needs to feel at home.