@leoperzz

mathematics @pucp | i love robots | building @muroboticsai

Joined December 2021
Leonardo Perez retweeted
This week, we’re shipping HandUMI to San Francisco, San Jose and London. Excited to see what tasks they collect data for :)
1
3
1
13
422
Yes, I was asking myself this too. One of my favorite attempts: arxiv.org/pdf/2507.16815 We can use the VLM to reason over constraints and condition the actor in latent space, but I think we need either a much higher-frequency approach or a way to decide when reasoning is needed
right now, systems 1 and 2 in robotics communicate with each other mainly through text, one example is task planners prompting VLAs to execute subtasks i think this will be a major bottleneck, it’s hard to express richness of the physical intent with text, if you think about it e.g. when we try to grab a cup, our exact actions depend on whether it was hot/slippery/damaged somewhere/many other things the way our brains “handoff” this information from system 2 to system 1 during actions feels way more integrated
107
Leonardo Perez retweeted
DiscreteRTC has been accepted to CoRL 2026! Also, two other works I contributed to, CLFIT led by @ThomasYuxinChen and Diagnosing Compositional Generalization in Sequential Robot Tasks led by @YixiaoWang777, have also been accepted! See you in Austin!🥳
Real-time Chunking (RTC) is designed to enable smooth asynchronous execution of flow-matching policies. However, it has some critical limitations: its inpainting-based async execution capability comes from inference-time corrections rather than the base policy, yielding little pre-training benefit, specific fine-tuning for better performance (e.g. training-time RTC), heuristic guidance, and extra computation that inflates the latency. In this work, we observe that discrete diffusion policies, which generate actions by iteratively unmasking, are natural asynchronous executors that resolve all limitations at once, being simpler to implement, faster at inference, and better at execution. Paper: arxiv.org/pdf/2604.25050 Code: github.com/outsider86/Discre… Website: outsider86.github.io/Discret…
4
33
3,664
If you wish to accelerate your data collection, some HandUMIs you must buy 😎
HandUMI now also supports I2RT YAM arms :) shop.murobotics.ai/
1
295
Leonardo Perez retweeted
Predictions on where robotics is going in the next few years
54
66
12
486
55,482
Leonardo Perez retweeted
Introducing Deft Robotics' unified deployment platform for physical AI. Frontier models are moving fast. Real world autonomy isn’t. Closing the gap takes more than better models. The missing piece? Infrastructure to deploy, intervene, learn from failures, and continuously improve. For the past year, we’ve been deploying in that gap and building the stack to close it. Now, we’re opening it up. Hardware available now. Software in beta. Here’s how to use it for your use-case 🧵(1/7) ↓
22
12
8
122
15,906
Leonardo Perez retweeted
Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/wri… I think it's an important change in the trajectory of robotics!
44
113
45
749
203,763
Leonardo Perez retweeted
Implementing inpainting/masking with a video model to improve the quality of data collected with HandUMI. Collection, QA, retargeting and inpainting -- all open source: github.com/murobotics-ai/han… DMs are open if you want some curated data for your robots :)
8
5
1
57
3,415
Leonardo Perez retweeted
Robots are not everywhere because the demo to deployment gap needs ridiculous amount of data. Contrary to common belief more SFT data (additional teleop) on the robot foundation model is not as effective in robotics, due to covariate shift. It will need a LOT of in-domain SFT data to fine-tune the model. ICL helps but doest bring us to mastery. Solution: Recursive self-improvement for robot policies via test-time compute without new teleop demos. Instead of collecting more noisy teleop data, Q-Planning uses a learning-based test-time harness to let black-box robot policies iteratively self-improve. It allows any base model to be improved in a manner similar to the "thinking mode" with extra hard thinking. The trick is to learn a evaluator for these possible plans, and turns out we can learn a value function without learning a full dynamics model. Excited to see how thinking mode will scale deployment grade robotics solutions faster. Check out the thread below and more fun info at q-planning.github.io
You fine-tune a robot foundation model on a hard task. It gets 25%. Now what? Introducing Q-Planning, a learning-based harness that lets large black-box robot policies recursively self-improve. On a hard fine-grained task: 25% → 80% in 100 robot attempts (~30 mins), no extra human data. q-planning.github.io/ 1/n 🧵
5
17
1
153
21,620
gemini robotics er-2 from @GoogleDeepMind is really good at annotating data. what other model should we test?
69,000+ robot datasets on @huggingface are tagged LeRobot. Most carry a single task string per episode. No subtasks, no timestamps, nothing a VLA/VAM/WAM can learn properly. So I started playing with @GoogleDeepMind's Gemini Robotics ER 2 as an auto-annotator: it watches the video and writes it up. Works pretty well! Should I open source it? ft. @alvax64 @leoperzz
3
508
🤗
To empower all of humanity, robots must work beyond controlled labs. In recent years, models, hardware, and compute have advanced significantly, but high-quality real-world data remains difficult, expensive, and slow to collect. RoboNet was founded to remove the bottlenecks through research. At @fdotinc Off Season II, we built HandUMI, a device to solve data collection for bimanual robotic arms with parallel-jaw grippers. We'll keep tackling each new bottleneck, one after another, until robots operate widely in the real world. Everything we have built so far is completely open source and available here: github.com/robonet-ai
2
284
This is exciting. I remember one of the sunday robotics founders saying that if you can solve a finite set of household tasks, you have a useful home robot. If we can transfer llm generalization into robot policies, we may not need massive data collection to get there.
Introducing Waddle Labs: Claude Code for robots. Connect our API to your robot and enter a prompt, then our agents write code to achieve the task in 20 minutes. @yiding_song @theWaddleLabs
1
3
222
Next step: Using HandUMI for both robot-free data collection and teleoperation. Collect data faster without the robot, combine it with real robot teleoperation data, and bring everything into the same joint space for more robust policies.
1/ Collect once. Retarget to any bimanual arm with parallel grippers. I firmly believe bimanual robotic arms are the right embodiment to start adding value in the real world, and HandUMI is here to help startups accelerate deployment and researchers run more experiments. Record data with HandUMI, no robot in the loop, replayed straight onto Agilex PiPER, OpenArm, TRLC-DK1, and I2RT YAM. - Robot-free capture. - Modular: swap the gripper on the arm you use and start collecting data - LeRobot v3-compatible dataset format - Built-in calibration + QA before conversion - Sim replay and real teleoperation - Apache-2.0, fully open source ft. @alvax64 @leoperzz @raulb4s @mbrq_13 @robonet_ @0xnonhuman 🧵👇
1
3
375
Leonardo Perez retweeted
Replying to @raulb4s
@raulb4s and @mbrq_13 testing teleop with HandUMI at the @0xnonhuman lab in Peru 🇵🇪, 7,250 km away from SF. This is part of the tests to make sure the IK works correctly and that the data collected with HandUMI feels like teleoperated data, but much cheaper. The complete software will be open sourced in the coming days! Data collection, teleop in sim and in real, data postprocessing, etc. Stay tuned! Want to collaborate on HandUMI or build your own? Join our Discord: discord.gg/V47FuUkFA
2
10
1
16
1,231
Hardware is open source and the software is coming this week. Stay tuned👀
Why do people keep collecting data with teleoperation when a bimanual robot setup costs more than $10k? Isn't there a solution that gives you the same data quality without the robot? At @robonet_, we want to build the Internet of Robotics. As part of that mission, we built HandUMI, a hand-worn data collection device for bimanual arms with parallel-jaw grippers. Specs per unit: - 276.5 grams - $110.68 - Encoder-precision gripper aperture - Integrated wrist camera - Tracking with the VR headset of your choice (Pico/Quest) - More than 5 grippers supported (Piper, Trossen, ARX, Soft gripper, Dream gripper) The best part: all the hardware is open source! Thanks @fdotinc for the hardware lab and the space to make this possible. ft. @alvax64 @leoperzz @raulb4s @mbrq_13 @BryanBRstds @Aryan_Mangla_ , and the rest of the @0xnonhuman team.
2
329
Leonardo Perez retweeted
Who is struggling with data collection for robotics? Stay tuned tomorrow, and you will save up to 8x on data collection costs. Building at Fuera de temporada II @fdotinc
2
10
613
Leonardo Perez retweeted
How can we scale perception-based humanoid learning without collecting massive humanoid teleoperation data? 🚀 Excited to finally share VLK! What excites me most about VLK is that it reframes data collection as a data generation problem. Instead of relying on expensive humanoid teleoperation, we automatically generate synchronized vision, language, and whole-body kinematics from reconstructed real-world scenes. Making this vision a reality required bridging three fundamental challenges: 👀 Perception: Bridging the RGB sim → real gap through visual domain randomization and motion blur mitigation during both training and deployment. 🤖 Embodiment: Bridging the kinematics → dynamics gap with real-time VLA deployment, test-time RTC, and SceneBot, enabling seamless deployment on a real humanoid. 🌍 Environment: Bridging the real-world → synthetic gap to enable scalable Vision-Language-Kinematics data generation through scene reconstruction and interaction synthesis. It has been an amazing journey working with such an incredible team. For a complete walkthrough of the project, check out @jiaman01's thread below 👇 🌐 Project: vision-language-kinematics.g… 📄 Paper: arxiv.org/abs/2606.30645 🎦 Video: youtu.be/ZB6k_iMJP7M Huge thanks to my amazing collaborators @jiaman01 @eric_srchen @TakaraTruong @ Pei Xu, and to our advisors @pabbeel @rocky_duan @KoushilSreenath @akanazawa @carlo_sferrazza @GuanyaShi @ckarenliu.
🤖 How can we scale up humanoid robot learning? Introducing 🌟VLK🌟: generating large-scale synthetic data with paired egocentric observations, text, and full-body G1 kinematics for learning humanoid loco-manipulation. No teleoperation needed! Website: vision-language-kinematics.g…
13
43
10
308
879,120
The next robotic foundation model may not be a model. It may be an autonomous self-improvement learning system. Recent work across @physical_int and @NVIDIARobotics GEAR reveals something more structural than stronger policies. It reveals a shift in what counts as the unit of intelligence in robotics. 1️⃣ Context Conditioning Layer — multimodal context (language, visual subgoals, metadata) that makes policies steerable instead of rigidly task-specified. 2️⃣ Foundation Policy (@physical_int π0.7—future) — a context-conditioned VLA whose steerable prompting produces language coaching and zero-shot cross-embodiment generalization by treating rich context as controllable input. 3️⃣ Skill Acquisition (@nvidia CHORD) — human demonstrations are transformed into object-centric contact-wrench trajectories. The system learns to instantiate the dynamics that move objects rather than imitate human motion. On 1,831 long-horizon tasks this yields 82% success and enables whole-body transfer from hand-only data. The change is not better imitation; it is a change in what is being imitated. 4️⃣ Skill Memory (@nvidia ASPIRE) — evolutionary search over executable control programs, debugged via multimodal traces and distilled into a reusable library of sensorimotor code. Memory now takes the form of programs that can be retrieved and composed, not weights that must be retrained. 5️⃣ Autonomous Research (@nvidia ENPIRE) — coding agents close a physical feedback loop in which they reset scenes, run policies on real robot fleets, analyze results, rewrite code, and iterate. The optimizer is no longer gradient descent on a fixed loss. It is autonomous experimentation that continuously rewrites both the policy and the training process itself. Robotics here is no longer model-centric; it is loop-centric. Our @saturdayrobotic Robotics & World Models Reading Club will delve deep into technical details of NVIDIA #ENPIRE on July 25, stay tuned 👉🏻 RSVP: luma.com/5ltk12w5. 6️⃣ Fleet-Scale Feedback — parallel physical interaction makes robot-hours a first-class scaling dimension. Static datasets are no longer the dominant source of supervision. The physical world becomes a continuously self-generating training signal. The shift is not that these components are becoming better aligned. It is that optimization, memory, representation, and control are collapsing into the same closed loop. The model is no longer the unit of intelligence. The loop is. Foundation models compressed intelligence into weights. Agentic robotics is decompressing it into steerable context, object-centric dynamics, executable programs, autonomous experimentation, and continuous physical feedback. Robotics is shifting from learning policies to building learning systems whose intelligence grows through self-modifying interaction with the world. The first generation of robot foundation models learned from datasets. The next generation may learn from running labs.
💐 Saturday Robotics & World Models Reading Club @saturdayrobotic is 3 months old! (March 28 → June 28) In 3 months, we've been devoted to building the best technical robotics research forum in Silicon Valley. What started as a small weekly reading group has grown into a thriving community where researchers, founders, engineers, and students come together to discuss the latest advances in: 🤖 Robotics 🌍 World Models 🦾 Embodied AI 🧠 Foundation Models for Physical Intelligence Every Saturday, we dive deep into papers, challenge assumptions, and host technical talks from leading researchers and builders across academia and industry. A huge thank you to every speaker, volunteer, and community member who has made this possible. Your curiosity, generosity, and technical depth are what make Saturday Robotics special. We're just getting started. Here's to the next chapter—bringing even more cutting-edge robotics research, world models, and embodied AI discussions to Silicon Valley. See you next Saturday! 🚀 👉🏻 luma.com/saturdayrobotic #SaturdayRobotics #Robotics #EmbodiedAI #WorldModels #PhysicalAI #MachineLearning #ArtificialIntelligence #SiliconValley
1
3
1
15
3,275
really impressive
"A parcel with snacks has been delivered for Flexion. Retrieve it using the stairs and come up using the elevator. Then unpack it and place the items into the empty drawer on the shelf in the snack area." One instruction. No human operator. Everything that follows is autonomous. Today we're introducing Reflect v1.0, our robotics intelligence platform for long-horizon work. From a single natural-language command, the robot understands the task, navigates a multi-floor building, calls elevators, handles doors, uses tools to unpack a box, and puts the items away. The biggest shift in v1.0 is that we use reinforcement learning across every layer, from low-level control to high-level reasoning. Long-horizon autonomy is unforgiving. The robot must recover on its own when things don't go to plan because in the real world, they never do. Combining reasoning, perception, physical execution and runtime robustness into a single mission-capable system is the foundation required to solve humanoid autonomy. Our team is just getting started. #HumanoidRobots #Flexion
142
Leonardo Perez retweeted
Debugging
1
4
9
404