@BuiltForRobotsi
iAccount based inSpain
About this account
- Account based in
- Spain
- Connected via
- Spain Android App
Account-level information from X, not a live location or the device used for a specific post.
Sr. Robotics Engineer @PALRobotics. I am a keen Robotics Enthusiast, who gets motivated by robots and their evolution.
Barcelona, Spain
Joined September 2016
- Tweets1.2K
- Following443
- Followers333
- Likes4.4K
Pinned Tweet
Go @OlympusMonsTeam!
Go @PALRobotics!💪🏽💪🏽
2 years of hardwork turned to a dream come true!!
Very proud to be part of the team!
Congratulations to the winners of the @NASAPrize Space Robotics Challenge Phase 2, presented by BHP! These teams have been working since 2019 to help develop code for NASA’s next generation of space robots. Visit the link to learn more! bit.ly/2H4xbyO
🚀 Excited to give a talk with Erik Holum on mujoco_ros2_control: High-Fidelity Physics Simulator for ROS 2!
We’ll show how MuJoCo + ros2_control enables high-fidelity physics, fast simulation, real-time sensors, complex mechanisms, and more.
#ROSCon #MuJoCo @OpenRoboticsOrg
Sai Kishor Kothakota retweeted
This is how Microduck learns to walk. Reinforcement Learning, explained by ducks. 🦆
(impeccable) music arrangement by @antoinepirrone
Sai Kishor Kothakota retweeted
FlashSAC at @holidayrobots just got a Best Paper Award at #RSS2026 🎉
One thing that stood out at #RSS2026 was how widely RSL-RL & PPO are used for sim-to-real RL. So we’re open-sourcing rsl_rl_flashsac, a FlashSAC implementation built on top of RSL-RL.
If you're already on RSL-RL, you can try FlashSAC without changing your existing stack. We're also working on Newton support and sim2sim2real deployment on the Unitree G1. More to come!
Repo: github.com/Holiday-Robot/rsl…
Sai Kishor Kothakota retweeted
New in @ScienceMagazine 🎉
Really excited to share that SONIC is now published in Science Robotics!
SONIC formulates super-scaling motion tracking as the foundational task for more natural and robust whole-body control. Our learned tracker and latent space further support foundation models like VLA/WAM training.
Huge thanks to the team and collaborators who made this happen.
science.org/doi/10.1126/scir…
New checkpoints coming in today!
#ScienceResearch #ScienceRoboticsResearch
Sai Kishor Kothakota retweeted
Ego2Robot might be one of the biggest data-scaling ideas we've seen for robot learning.
Instead of collecting millions of expensive robot demonstrations, Alibaba's Qwen team converts first-person human videos into robot demonstrations.
📹 Human videos → 🤖 Robot data.
18,561 hours.
15 robot morphologies.
That's the largest ego-to-robot dataset so far.
Sai Kishor Kothakota retweeted
69,000+ robot datasets on @huggingface are tagged LeRobot.
Most carry a single task string per episode. No subtasks, no timestamps, nothing a VLA/VAM/WAM can learn properly.
So I started playing with @GoogleDeepMind's Gemini Robotics ER 2 as an auto-annotator: it watches the video and writes it up.
Works pretty well!
Should I open source it?
ft. @alvax64 @leoperzz
Sai Kishor Kothakota retweeted
New framework: Kick down your robot, it will get back up every time 🥋
Chinese startup RoboParty is a Beijing startup founded April 2025 by Huang Yi, originally shipping ROBOTO ORIGIN, the world's first full-stack open-source bipedal humanoid.
They released UFO: Unsupervised Reinforcement Learning Framework for Humanoid Control.
DEFINITIONS -> what differs is where the learning signal comes from:
- SUPERVISED: humans supply the right answers (labels), the model imitates them.
- UNSUPERVISED: no answer key, the model finds structure in raw data on its own.
- REINFORCEMENT LEARNING: no answer key either, the model tries things and a reward scores each attempt.
→ UNSUPERVISED RL: trial and error where the agent invents its own rewards, instead of engineers hand-writing one per task.
REPRESENTATION LEARNING: compress raw states into a useful internal map.
TEMPORAL DISTANCE: distance on that map is "how many steps from A to B."
CONTRASTIVE: trained by pulling together what's close in time, pushing apart what isn't.
-> CONTRASTIVE TEMPORAL-DISTANCE REPRESENTATION LEARNING: the model builds an internal map of body states where distance means how many steps it takes to get from one to another. It is trained by contrast: states that occur close together in a movement get pulled together in the map, randomly paired states get pushed apart.
UFO is an open-source training framework that teaches humanoid robots skills, like getting up, walking, goal-reaching, teleoperation, without reference motions -> no motion-capture or human-video demonstrations to imitate.
Its core is TeCH, a contrastive temporal-distance representation-learning algorithm: the robot explores, builds pseudo-goals by temporal rolling, and learns goal-conditioned policies from a single unified progress reward.
One framework trains five different robots (Unitree G1/H1, RoboParty RP0/RP1, AgiBot X2) with automatic config conversion in ~2–3 hours per robot!
The real novelty here "no demonstrations at all".
No data-collection arms race,the dominant humanoid-locomotion recipe is tracking: imitate mocap/retargeted-human reference trajectories.
The robot self-generates goals from its own exploration and learns from a progress reward, needing zero reference motion data.
Everybody else is fighting over data acquisition, while this team just teleports out of the race entirely (inb4 "competition is for losers 💀 ).
This strategy reminds me of the DeepSeek playbook applied to robots: open-source the whole stack to become the global default and commoditize everyone else.
RoboParty is giving away hardware and now control software (UFO) to be the Android of humanoids.
Yet another reason for the US to ban Chinese open models perhaps 🥶 ?
What I also really like about this approach is the cross-embodiment infrastructure, one framework trains Unitree G1/H1, RoboParty RP0/RP1, and AgiBot X2 with automatic configuration conversion.
Just like Physical Intelligence, RoboParty seems to place itself as a neutral hardware agnostic middle man.
Also woth mentioning: their ability ot perform stable skill injection, e.g. adding a cartwheel without forgetting how to walk.
A common failure of RL humanoid policies is that teaching a new agile skill destabilizes the existing ones (catastrophic forgetting).
UFO claims you can inject rare motions (cartwheel) without collapsing learned behavior.
If it holds, that's a significant incremental/continual skill-learning!
But again, I have to underline it: no arXiv, no external validation, no success-rate numbers.
-> robotics badely needs an independent unbiased evaluator imho.
Still, look at that cool demo: robot is getting kicked and pushed around (serious disturbance) during teleoperation (controlled the person at the back wearing the VR headset), and still managed to always get back up.
This is some serious demonstration of stability and robustness!
Sai Kishor Kothakota retweeted
RoboPlan 0.6.0 is now available!
Highlights include:
📐 RRT planner: Added pose constraints via projection
🐷 Optimal Inverse Kinematics (OInK) solver: Switched backend to ProxQP (ProxSuite)
🪟 Windows support via Pixi (in addition to Linux/macOS)
github.com/open-planning/rob…
Sai Kishor Kothakota retweeted
Today, we’re launching Tau’s humanoid cleaning service in San Francisco at $30 per hour.
Access is initially invite-only as we scale operations. If you don’t have an invite yet, join the waitlist at tau-robotics.com.
All footage is shown at 1× speed. Each humanoid is jointly controlled by a human operator and AI.
Sai Kishor Kothakota retweeted
🚀 New open-source project!
I’ve extended the learning framework from my previous works, GfR (RSS 2026) and HIL (TOG 2026), to two representative tasks on the Unitree G1: box moving and box climbing, built on top of Holosoma.
The same learning framework applies to both locomotion and loco-manipulation tasks. Using only a single reference motion clip for each task, the learned policies generalize across a wide range of task conditions and goals.
The key idea (from GfR and HIL) is simple:
1️⃣ We found that a policy can achieve high-quality motion tracking without receiving the target reference motion as actor input. Instead, task goals are extracted from the reference motion and provided to the actor together with the robot state. The reference itself is used to define the tracking reward.
The policy is conditioned on task goals rather than the reference motion, and the learned behavior can easily transfer to new task conditions.
2️⃣ Based on this, we develop a hybrid multi-task training setup:
• In the tracking task, goals are extracted from the reference motion and the policy is trained with tracking rewards.
• In the general RL task, goals and task conditions are randomly sampled, and the policy is trained with sparse goal-reaching rewards.
• The actor always receives the same observation structure across tasks, while the critic is provided with task-specific signals.
The same policy learns from both tracking and general RL tasks, transferring behaviors acquired through motion tracking to more general goal-driven settings.
This project provides a practical example of how the GfR/HIL methodology can be adapted to new tasks. Note: this is an unofficial research implementation and extension, not the official code release for either GfR or HIL.
🔗 github.com/jiashunwang/Hybri…
Sai Kishor Kothakota retweeted
Humanoid motion planning has a brutal reality check:
A plan can look good in an LLM or search tree and still be physically impossible for the robot.
🎧🎙️That is exactly why I’m excited for tomorrow’s podcast episode with Majid Khadiv.
His group just published FARO, a framework that rapidly checks whether proposed contact sequences are actually executable on humanoid hardware.
If the motion is physically feasible, FARO generates a dynamically consistent trajectory.
An RL-based whole-body controller then tracks that trajectory on the real robot.
That matters because humanoid teams waste too much compute and engineering time on motions that were never possible in the first place.
Foundation models propose. Physics optimization filters. RL executes.
Great work by the entire team of ATARI Lab!!
📌 Paper:
arxiv.org/pdf/2607.18362
Hardware execution:
youtu.be/R6qCHoCormQ
———-
Weekly robotics and AI insights.
Subscribe free: 22astronauts.com
Sai Kishor Kothakota retweeted
Trained on zero real-world data.
Learned to walk, pick up boxes, and follow multi-step instructions...
in the REAL world. ( 📌 Paper below)
Researchers from Amazon FAR, Berkeley, Stanford, and CMU scanned real rooms with an iPhone, rebuilt them as 3D Gaussian Splatting scenes, then generated 48,000 synthetic trajectories of a Unitree G1 walking, grasping, and placing objects inside those virtual replicas. They rendered the robot's first-person camera view from each run and paired it with the matching language instruction and motion data.
That's the dataset every humanoid team needs and nobody has: synced egocentric video + language + kinematics, at scale. Instead of collecting it in the real world, they manufactured it.
They trained a vision-language-kinematics policy on that synthetic data alone, then deployed it on the physical G1 across five task types: navigation to a named object, lifting boxes of three different sizes with no per-size tuning, chained multi-step tasks, robustness to mid-task layout changes and flickering lights, and multi-minute long-horizon runs.
No real-world fine-tuning at any point.
Real-world interaction data has been the hard limit on humanoid learning... slow, expensive, and small. If scanning a room once and synthesizing thousands of labeled interactions holds up as a general recipe, that limit moves. Data stops being the bottleneck robotics teams have to solve for.
📌 Paper: arxiv.org/abs/2606.30645
Project: vision-language-kinematics.g…
——-
Weekly robotics and AI insights.
Subscribe free: 22astronauts.com
One of the toughest challenges for robots in factories is handling large, irregular, flexible parts reliably over long periods.
Here is an uncut video of Xiaomi's humanoid robot continuously sorting center console side covers on the production line.
Sai Kishor Kothakota retweeted
Exciting to see more momentum around real → sim → real for robot learning!
We’ve been exploring a related idea in VLK, where we reconstruct real-world scenes to synthesize large-scale vision-language-whole-body kinematics data for training humanoid VLAs, enabling perception-based loco-manipulation directly from RGB and language.
Really exciting to see real-world reconstruction becoming an increasingly powerful foundation for scalable robot learning 🚀
VLK: vision-language-kinematics.g…
Historically, RL policies for robots have been trained in synthetic, untextured environments, limiting perceptive policies to depth images where the sim-to-real gap is manageable. RGB has always had more potential, but leveraging it to train policies in simulation remained an open problem.
Partnering with @NianticSpatial and @NVIDIARobotics, we built a pipeline that addresses exactly that. We can now scan a real deployment site with off-the-shelf hardware, reconstruct it into a photorealistic Gaussian splat, and run massively parallel RL training. The policies trained in our Gym environment then transfer zero-shot to the real robot and environments they were trained for. This enables faster deployment of more capable and robust policies for the end user.
The new resulting capabilities are a big step towards solving sim-to-real and also apply well beyond navigation. Read the full technical breakdown on our blog; link in the comments.
#HumanoidRobots #Flexion #NianticSpatial #NVIDIA
Sai Kishor Kothakota retweeted
Historically, RL policies for robots have been trained in synthetic, untextured environments, limiting perceptive policies to depth images where the sim-to-real gap is manageable. RGB has always had more potential, but leveraging it to train policies in simulation remained an open problem.
Partnering with @NianticSpatial and @NVIDIARobotics, we built a pipeline that addresses exactly that. We can now scan a real deployment site with off-the-shelf hardware, reconstruct it into a photorealistic Gaussian splat, and run massively parallel RL training. The policies trained in our Gym environment then transfer zero-shot to the real robot and environments they were trained for. This enables faster deployment of more capable and robust policies for the end user.
The new resulting capabilities are a big step towards solving sim-to-real and also apply well beyond navigation. Read the full technical breakdown on our blog; link in the comments.
#HumanoidRobots #Flexion #NianticSpatial #NVIDIA
Sai Kishor Kothakota retweeted
FlashSAC won the Outstanding Paper Award at RSS 2026 🎉
We got off-policy RL fast and stable enough to beat PPO and FastTD3 across 60+ tasks and 10 simulators with minimal tuning!
TL;DR: If you're working on dexterous manipulation, just try FlashSAC!
holiday-robot.github.io/Flas…
anthropic.com/research/claud…
Claude plays robotics
A must-read!!!
Sai Kishor Kothakota retweeted
We're presenting FlashSAC at #RSS2026 tomorrow!
FlashSAC allows humanoid locomotion within only 15-30 minutes using a single RTX GPU.
Meet us at the poster for chat!
📺Presentation: “Control & Dynamics” 09:00-09:15 AM
📜Poster: 12:30-1:30 PM
holiday-robot.github.io/Flas…
We scaled off-policy RL to sim-to-real.
To our knowledge, FlashSAC is the fastest and most performant RL algorithm across IsaacLab, MuJoCo Playground, and many more, all with a single set of hyperparameters.
Project page: holiday-robot.github.io/Flas…
Paper: arxiv.org/pdf/2604.04539