@eric_srcheni
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
PhD in Stanford CS, Prev Undergrad at HKU. Interested in robotics
Stanford, CA
Joined September 2023
- Tweets177
- Following665
- Followers708
- Likes280
Pinned Tweet
Humanoids excel in free space but struggle with real-world contact. Meet SceneBot 🤖 the first unified RL framework for ALL free-space locomotion, terrain traversal and object interaction!
By conditioning on per-link contact labels, it masters complex, interaction-rich tasks like carrying a box upstairs. Code & data open-sourcing soon! 📦 🪜
Paper: arxiv.org/abs/2606.27581
Website: ericcsr.github.io/scenebot/
Sirui Chen retweeted
Can a robot design a tool from scratch? Introducing HOT: Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization.
Check out our project page for more details, videos, and code! 👇
hot.yinghanchen.com
Sirui Chen retweeted
What can GPT-6 Astra do on a humanoid when the contact gets hard?
Give it the right interaction kernel.
Led by @yk20wang, our team presents KPI: A Promptable Kernel for Physical Interaction on Humanoids.
With our agentic system, Astra completes long-horizon, contact-rich tasks zero-shot—no task demos or task-specific training.
1/6
We need a powerful system 2
What can Astra do when given a humanoid embodiment?
We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning.
Here's how we did it 👀: tml.stanford.edu/homebody/
Sirui Chen retweeted
Incredibly excited to return to ETH Zurich as Assistant Professor of Robotics and Artificial Intelligence next summer. I'll be recruiting PhD students to build a new lab on general-purpose embodied intelligence for humanoid robots and dexterous manipulation. Details coming soon!
Sirui Chen retweeted
Introducing Real-Time EXPO-FT – Fast and Reliable RL for Real-Time VLA Policies!
Real-Time EXPO-FT unlocks π0.5 on challenging dynamic tasks, such as balancing a ball on a plate and striking a ball into the goal
(1/6)
Super impressive, congratulations!
CHIP has been accepted to CoRL 2026, see you in Austin Texas!
Sirui Chen retweeted
Introducing TacThru: Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation, accepted to IEEE Robotics and Automation Letters (RA-L)!
Check out our project page for more details, videos, code, and hardware! 👇
tacthru.yuyang.li/
Sirui Chen retweeted
A large-scale motion tracking framework can convert teleoperation, video, text, and music cues into humanoid control signals for running, jumping, or throwing a cup in the trash.
Learn more in Science #Robotics: scim.ag/3TYYKlp
Sirui Chen retweeted
natural language-> CAD design
great proj @eric_srchen
As a CS student who suck at CAD design, I use an alternative approach: build a lightweight CAD design harness for an AI agent! Check out: github.com/Ericcsr/AgentCAD
It can design complex CAD projects from natural language; both you and the agent can preview the model in real time; it also supports version control, per-part editing, and URDF export. You can also switch between different model backbones
As a CS student who suck at CAD design, I use an alternative approach: build a lightweight CAD design harness for an AI agent! Check out: github.com/Ericcsr/AgentCAD
It can design complex CAD projects from natural language; both you and the agent can preview the model in real time; it also supports version control, per-part editing, and URDF export. You can also switch between different model backbones
SONIC is officially published in Science Robotics today and it made the Science front page.
We show the promise of scaling motion tracking toward natural, robust whole-body control for humanoid robots.
Huge thanks to the team, and more exciting work is on the way.
Paper: science.org/doi/10.1126/scir…
Code: nvlabs.github.io/GEAR-SONIC
Sirui Chen retweeted
🤖 A robot that transforms between humanoid and dexterous hand?
Introducing Handroid, a reconfigurable robot with a shared 27-DoF body 🖐️→🚶→🖐️:
• The same joints, different roles
• Expanded robot task space
• Shared sensing and control interfaces
handroid.org
Sirui Chen retweeted
A new piece of research today by @wangyenjen, @jiaman01, @pabbeel et al.
Called VLK, for Vision-Language-Kinematics, it is a new model that shows it is possible to learn humanoid loco-manipulation from synthetic scenes.
Here is what it means in practice.
The team teaches a Unitree G1 humanoid to walk to, pick up, carry, and place objects from a first-person camera plus a spoken instruction
All this trained entirely on synthetic data, with zero real-robot data collection (no teleoperation etc.).
Here is what the pipeline looks like:
- scan a room with an iPhone (Polycam) into a metric-scale 3D Gaussian-Splatting replica (overly simplified: a fast low compute 3D model)
- synthesize whole-body G1 motions with conditional diffusion
- replay this synthetic data them in Isaac Sim to render egocentric frames with randomized lighting/appearance (i.e. domain randomization in simulator)
- fine-tune a VLA to predict a 1-second future kinematic trajectory at 30 fps
A separate contact-aware whole-body tracker (SceneBot) then turns those trajectories into joint commands at 50 Hz on the real robot.
Total data: 48,000 synthetic trajectories across 8 scanned rooms, evaluated on real hardware with no fine-tuning. Really impressive.
What I find really interesting about this approach, is that it could replace teleoperation entirely.
No human ever moves the robot or wears a mocap suit.
If you compare it to China's playbook (Xiaomi's 100k UMI hours, JD paying collectors ~$3/hr), VLK's entire corpus is super cost effective.
Now remains the question: can it scale efficiently and generalize to more complex tasks (dexterous, for example)?
Also worth mentioning: this is a strong win for domain randomization.
Randomizing lighting beats collecting more data, by far:
- no randomization 41% walking success rate
- full randomization 90%
- lighting-only already gets 87%
- camera-jitter-only gets just 48%.
Randomized lighting in the renderer costs nothing and is unreasonably effective!
Is most of the sim-to-real gap here is a rendering-appearance problem, and not a physics problem?
A positive answer to this question would shift the entire robotics data acquisition race happening right now.
Last learning from this piece of research: predicting a whole second at once with the VLA/VLK solves the real-time control problem that happens when a single action is predicted instead.
Outputing a plan and letting the robot controller take care of the low-level dynamics seems like a much more efficient approach.
I hope you enjoy this video as much as I did, where the robots adapts in real time to a changing environment:
Sirui Chen retweeted
Exciting to see more momentum around real → sim → real for robot learning!
We’ve been exploring a related idea in VLK, where we reconstruct real-world scenes to synthesize large-scale vision-language-whole-body kinematics data for training humanoid VLAs, enabling perception-based loco-manipulation directly from RGB and language.
Really exciting to see real-world reconstruction becoming an increasingly powerful foundation for scalable robot learning 🚀
VLK: vision-language-kinematics.g…
Historically, RL policies for robots have been trained in synthetic, untextured environments, limiting perceptive policies to depth images where the sim-to-real gap is manageable. RGB has always had more potential, but leveraging it to train policies in simulation remained an open problem.
Partnering with @NianticSpatial and @NVIDIARobotics, we built a pipeline that addresses exactly that. We can now scan a real deployment site with off-the-shelf hardware, reconstruct it into a photorealistic Gaussian splat, and run massively parallel RL training. The policies trained in our Gym environment then transfer zero-shot to the real robot and environments they were trained for. This enables faster deployment of more capable and robust policies for the end user.
The new resulting capabilities are a big step towards solving sim-to-real and also apply well beyond navigation. Read the full technical breakdown on our blog; link in the comments.
#HumanoidRobots #Flexion #NianticSpatial #NVIDIA
Sirui Chen retweeted
We have a few new SONIC checkpoints that we will be releasing!
First one: lower-latency whole-body teleoperation!
Already released at: github.com/NVlabs/GR00T-Whol…
Next up: Better squatting + wrist tracking!
Sirui Chen retweeted
Task success is not enough for robots. A robot can complete the task and still leave the user feeling dissatisfied.
Check out our work E-MPC: An Engagement-Aware Human-in-the-loop Framework for Robotic Systems at #RSS2026. It jointly reasons about task success and human factors.
Sirui Chen retweeted
World models are one of the most exciting frontiers for robotics and embodied AI right now and we're growing the team at Meta FAIR to push on it.
Hiring Research Scientists in:
🇺🇸 Menlo Park
🇨🇦 Montreal
Links below 👇 or find me at #RSS2026
Sirui Chen retweeted
Excited to present our work "SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation" at RSS 2026!
SuperMap is a living spatial memory for embodied AI — it perceives the world, remembers its evolution, and supports reasoning and action.
Perceive → Remember → Reason → Act
🌍 Project Page: superodometry.com/supermap
#RSS2026 #Robotics #SLAM #EmbodiedAI #SpatialAI #CMURobotics