@tomssilver

Assistant Professor @Princeton. Developing robots that plan and learn to help people.

Princeton, NJ
Joined October 2011
I like these videos. An LLM agent as an end2end robot policy is very slow! The approach that excites/concerns me more is one where the LLM writes a programmatic policy. Latency is not necessarily an issue then, and programs can be reused. Ofc, Moravec's paradox may still apply.
GPT-6 Astra vs Opus 5.5 vs Grok 4.7 vs MolmoAct2. Real time. No sped-up demos. Zero-shot. No cherry-picking. LLM robot control is progressing, but we’re definitely not there yet. Low-level control for robotics still matters according to Moravec’s paradox.
3
28
2,560
This week's #PaperILike is "Artificial Intelligence: An Empirical Science" (Herbert Simon, AIJ 1995). Timely read as recent results have us revisiting foundational questions. Plus hot takes, eg: engineering is "science for people who are impatient." PDF: ic.unicamp.br/~wainer/cursos…
1
1
38
2,522
Tom Silver retweeted
We put a lot of effort into making the experimental setup clean and relatively strict, and hope the sandbox can be useful for future agent evaluation. Also don’t miss the qualitative examples, where the agents came up with some pretty interesting strategies 👀
Hungry for more Astra robot videos? 🤖 How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way! New preprint 🧵👇 (1/9)
1
1
10
1,198
A number of these problems were unsolvable (or at least very very hard to write down a solution to) ~6 months ago. Wild times!
I’ve seen robots do plenty of unexpected things. That usually means it’s time to fix a bug. This new project is one of only a few times in my career where I’ve seen robots do things that are unexpectedly clever—that none of us on the project anticipated. A few examples...
1
2
16
2,542
Tom Silver retweeted
Agents like Astra has been popular for a while, but “rigorously” what can they actually “solve” at all? Excited to share our latest work on strict, red-teamed sandbox 🔐evaluation on coding agents! (They are indeed surprisingly smart 🤯🤯)
1
6
686
Tom Silver retweeted
This is my first paper with @tomssilver's group, and I'm super happy with the results. Thank you @Bw_Li1024 @thosehippos @yichao_liang @WangRobin58334 @YixuanHuang13 for all the help! 🙏 Excited to keep working together! 🚀 🔗 Paper + videos: agenticgentamp.github.io/ (9/9)
1
5
1
17
1,635
I’ve seen robots do plenty of unexpected things. That usually means it’s time to fix a bug. This new project is one of only a few times in my career where I’ve seen robots do things that are unexpectedly clever—that none of us on the project anticipated. A few examples...
Hungry for more Astra robot videos? 🤖 How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way! New preprint 🧵👇 (1/9)
4
11
2
100
11,789
🏀 🗑️
2
2
1
12
673
The early deadline has passed, but there's still time to submit to LEAP @corl_conf! We're also excited to announce that Dieter Fox has joined our program 🙂 See you in Austin!
9
32
2,506
This week's #PaperILike is "Principles of Animal Cognition for LLM Evaluations: A Case Study on Transitive Inference" (Rane et al., 2025). Still hunting for ways to understand LLMs/agents. This week: animal cognition! PDF: amandaroyka.github.io/ICMLPO… Also: arxiv.org/abs/2503.02882
1
23
1,579
Very excited to ramp up our work working with Tom! We're looking for an exceptional postdoctoral researcher (details below) to work jointly with Tom and Basis on MARA project--our effort to build robotic agents that actively learn models of the physical world.
Replying to @BasisOrg
@basisorg and my group at Princeton are recruiting a postdoc to work at the intersection of robotics and code-based world models. We’re looking for someone excited about abstractions & planning, and who knows their way around a real robot. Link 👇 Thanks for boosting!
1
4
24
1,909
Replying to @BasisOrg
@basisorg and my group at Princeton are recruiting a postdoc to work at the intersection of robotics and code-based world models. We’re looking for someone excited about abstractions & planning, and who knows their way around a real robot. Link 👇 Thanks for boosting!
2
12
1
41
4,351
This week's #PaperILike is "Latent Programming Horizons in Coding Agents" (Silva et al., 2026). Now seems like a good time to better understand what coding agents are doing. Here's a nice example of the kind of analysis one can do. PDF: arxiv.org/abs/2607.05188
4
1
38
3,316
This week's #PaperILike is "Safe Model-based Reinforcement Learning with Stability Guarantees" (Berkenkamp et al., NeurIPS 2017). Some safe RL guarantees the learned policy is safe; this guarantees safety *during* learning. Important for RL in real! PDF: arxiv.org/abs/1705.08551
1
2
44
3,164
Excited to announce the 4th LEAP Workshop at CoRL #CoRL2026! 🚀 We’re looking forward to bringing together the community to discuss learning, reasoning, and planning for robots. Hope to see you there! Learn more: leap-workshop.github.io/
We're excited to announce the 4th Workshop on Learning Effective Abstractions for Planning (LEAP) at #CoRL2026! Previous LEAP papers have gone on to win awards at main conferences (SymSkill, Universal Visual Decomposer). Yes, we'll take all the credit! Workshop link 👇
1
6
722
Proud of being part of this organizing team! The topic is more timely than ever! Make sure you submit your contributions and join us!
We're excited to announce the 4th Workshop on Learning Effective Abstractions for Planning (LEAP) at #CoRL2026! Previous LEAP papers have gone on to win awards at main conferences (SymSkill, Universal Visual Decomposer). Yes, we'll take all the credit! Workshop link 👇
1
14
1,592