@eric_srchen

PhD in Stanford CS, Prev Undergrad at HKU. Interested in robotics

Stanford, CA
Joined September 2023
Humanoids excel in free space but struggle with real-world contact. Meet SceneBot 🤖 the first unified RL framework for ALL free-space locomotion, terrain traversal and object interaction! By conditioning on per-link contact labels, it masters complex, interaction-rich tasks like carrying a box upstairs. Code & data open-sourcing soon! 📦 🪜 Paper: arxiv.org/abs/2606.27581 Website: ericcsr.github.io/scenebot/
6
36
2
159
35,912
Sirui Chen retweeted
Can a robot design a tool from scratch? Introducing HOT: Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization. Check out our project page for more details, videos, and code! 👇 hot.yinghanchen.com
6
1
16
457
What can GPT-6 Astra do on a humanoid when the contact gets hard? Give it the right interaction kernel. Led by @yk20wang, our team presents KPI: A Promptable Kernel for Physical Interaction on Humanoids. With our agentic system, Astra completes long-horizon, contact-rich tasks zero-shot—no task demos or task-specific training. 1/6
9
20
4
159
15,160
We need a powerful system 2
What can Astra do when given a humanoid embodiment? We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning. Here's how we did it 👀: tml.stanford.edu/homebody/
1
28
6,582
Incredibly excited to return to ETH Zurich as Assistant Professor of Robotics and Artificial Intelligence next summer. I'll be recruiting PhD students to build a new lab on general-purpose embodied intelligence for humanoid robots and dexterous manipulation. Details coming soon!
88
55
7
1,135
47,479
Sirui Chen retweeted
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies! Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal (1/6)
6
22
5
202
48,407
Super impressive, congratulations!
Introducing OM-1, our first robot foundation model, zero-shot generalizing to any robot: table-top arms, industrial arms and humanoids. - learned directly from human manipulation data - no teleop/robot data - close to human-level dexterity and efficiency - multi-robot collab
1
4
23
2,458
CHIP has been accepted to CoRL 2026, see you in Austin Texas!
What missing in RL based humanoid controller from industrial robots are precision and force control. CHIP can do both. We propose a simple recipe to build humanoid impedance controller, which can be used for wiping, carrying large objects and multi-robot collaboration.
1
1
23
1,416
Introducing TacThru: Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation, accepted to IEEE Robotics and Automation Letters (RA-L)! Check out our project page for more details, videos, code, and hardware! 👇 tacthru.yuyang.li/
4
7
449
A large-scale motion tracking framework can convert teleoperation, video, text, and music cues into humanoid control signals for running, jumping, or throwing a cup in the trash. Learn more in Science #Robotics: scim.ag/3TYYKlp
1
11
45
3,761
natural language-> CAD design great proj @eric_srchen
As a CS student who suck at CAD design, I use an alternative approach: build a lightweight CAD design harness for an AI agent! Check out: github.com/Ericcsr/AgentCAD It can design complex CAD projects from natural language; both you and the agent can preview the model in real time; it also supports version control, per-part editing, and URDF export. You can also switch between different model backbones
1
6
1,228
As a CS student who suck at CAD design, I use an alternative approach: build a lightweight CAD design harness for an AI agent! Check out: github.com/Ericcsr/AgentCAD It can design complex CAD projects from natural language; both you and the agent can preview the model in real time; it also supports version control, per-part editing, and URDF export. You can also switch between different model backbones
3
1
22
3,140
Sirui Chen retweeted
SONIC is officially published in Science Robotics today and it made the Science front page. We show the promise of scaling motion tracking toward natural, robust whole-body control for humanoid robots. Huge thanks to the team, and more exciting work is on the way. Paper: science.org/doi/10.1126/scir… Code: nvlabs.github.io/GEAR-SONIC
4
21
4
196
19,296
Sirui Chen retweeted
🤖 A robot that transforms between humanoid and dexterous hand? Introducing Handroid, a reconfigurable robot with a shared 27-DoF body 🖐️→🚶→🖐️: • The same joints, different roles • Expanded robot task space • Shared sensing and control interfaces handroid.org
6
33
18
144
33,242
Sirui Chen retweeted
A new piece of research today by @wangyenjen, @jiaman01, @pabbeel et al. Called VLK, for Vision-Language-Kinematics, it is a new model that shows it is possible to learn humanoid loco-manipulation from synthetic scenes. Here is what it means in practice. The team teaches a Unitree G1 humanoid to walk to, pick up, carry, and place objects from a first-person camera plus a spoken instruction All this trained entirely on synthetic data, with zero real-robot data collection (no teleoperation etc.). Here is what the pipeline looks like: - scan a room with an iPhone (Polycam) into a metric-scale 3D Gaussian-Splatting replica (overly simplified: a fast low compute 3D model) - synthesize whole-body G1 motions with conditional diffusion - replay this synthetic data them in Isaac Sim to render egocentric frames with randomized lighting/appearance (i.e. domain randomization in simulator) - fine-tune a VLA to predict a 1-second future kinematic trajectory at 30 fps A separate contact-aware whole-body tracker (SceneBot) then turns those trajectories into joint commands at 50 Hz on the real robot. Total data: 48,000 synthetic trajectories across 8 scanned rooms, evaluated on real hardware with no fine-tuning. Really impressive. What I find really interesting about this approach, is that it could replace teleoperation entirely. No human ever moves the robot or wears a mocap suit. If you compare it to China's playbook (Xiaomi's 100k UMI hours, JD paying collectors ~$3/hr), VLK's entire corpus is super cost effective. Now remains the question: can it scale efficiently and generalize to more complex tasks (dexterous, for example)? Also worth mentioning: this is a strong win for domain randomization. Randomizing lighting beats collecting more data, by far: - no randomization 41% walking success rate - full randomization 90% - lighting-only already gets 87% - camera-jitter-only gets just 48%. Randomized lighting in the renderer costs nothing and is unreasonably effective! Is most of the sim-to-real gap here is a rendering-appearance problem, and not a physics problem? A positive answer to this question would shift the entire robotics data acquisition race happening right now. Last learning from this piece of research: predicting a whole second at once with the VLA/VLK solves the real-time control problem that happens when a single action is predicted instead. Outputing a plan and letting the robot controller take care of the low-level dynamics seems like a much more efficient approach. I hope you enjoy this video as much as I did, where the robots adapts in real time to a changing environment:
1
8
85
5,991
Sirui Chen retweeted
Exciting to see more momentum around real → sim → real for robot learning! We’ve been exploring a related idea in VLK, where we reconstruct real-world scenes to synthesize large-scale vision-language-whole-body kinematics data for training humanoid VLAs, enabling perception-based loco-manipulation directly from RGB and language. Really exciting to see real-world reconstruction becoming an increasingly powerful foundation for scalable robot learning 🚀 VLK: vision-language-kinematics.g…
Historically, RL policies for robots have been trained in synthetic, untextured environments, limiting perceptive policies to depth images where the sim-to-real gap is manageable. RGB has always had more potential, but leveraging it to train policies in simulation remained an open problem. Partnering with @NianticSpatial and @NVIDIARobotics, we built a pipeline that addresses exactly that. We can now scan a real deployment site with off-the-shelf hardware, reconstruct it into a photorealistic Gaussian splat, and run massively parallel RL training. The policies trained in our Gym environment then transfer zero-shot to the real robot and environments they were trained for. This enables faster deployment of more capable and robust policies for the end user. The new resulting capabilities are a big step towards solving sim-to-real and also apply well beyond navigation. Read the full technical breakdown on our blog; link in the comments. #HumanoidRobots #Flexion #NianticSpatial #NVIDIA
5
8
53
5,680
We have a few new SONIC checkpoints that we will be releasing! First one: lower-latency whole-body teleoperation! Already released at: github.com/NVlabs/GR00T-Whol… Next up: Better squatting + wrist tracking!
12
34
5
259
25,371
Task success is not enough for robots. A robot can complete the task and still leave the user feeling dissatisfied. Check out our work E-MPC: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems at #RSS2026. It jointly reasons about task success and human factors.
2
9
1
37
4,146
Sirui Chen retweeted
World models are one of the most exciting frontiers for robotics and embodied AI right now and we're growing the team at Meta FAIR to push on it. Hiring Research Scientists in: 🇺🇸 Menlo Park 🇨🇦 Montreal Links below 👇 or find me at #RSS2026
3
12
1
169
24,852
Sirui Chen retweeted
Excited to present our work "SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation" at RSS 2026! SuperMap is a living spatial memory for embodied AI — it perceives the world, remembers its evolution, and supports reasoning and action. Perceive → Remember → Reason → Act 🌍 Project Page: superodometry.com/supermap #RSS2026 #Robotics #SLAM #EmbodiedAI #SpatialAI #CMURobotics
15
65
4
402
30,340