@PokeBot_Global. Final year PhD @Tsinghua University, IIIS.

Joined July 2021
🌶️🤖𝐖𝐞 𝐜𝐨𝐨𝐤𝐞𝐝 𝐭𝐡𝐞 𝐫𝐨𝐛𝐨𝐭𝐢𝐜𝐬 𝐰𝐨𝐫𝐥𝐝’𝐬 𝐟𝐢𝐫𝐬𝐭 𝐩𝐥𝐚𝐭𝐞 𝐨𝐟 𝐌𝐚𝐩𝐨 𝐓𝐨𝐟𝐮.🤖🌶️ This was not just a cooking demo. It was a 30-minute, long-horizon robotics challenge packed with delicate, continuous, high-dexterity actions. With only a small number of demonstrations, our model learned to perform complex manipulation over an extended sequence. At the same time, we pushed hard on motion control and system-level optimization, making the robot move smoothly while keeping the high success rate. For us, this plate of Mapo Tofu is more than a dish. It is a small but meaningful step toward home robots that can handle real everyday tasks with the dexterity, consistency, and reliability people expect.
17
25
9
142
25,739
Zhecheng Yuan retweeted
DINOv2 has never seen a caption, and Qwen3 has never seen an image. We still aligned their embedding spaces without a single image-caption pair. It even works when the images and the captions come from different datasets. Project page: dominik-schnaus.github.io/un… ⬇️
45
172
53
1,248
94,494
Zhecheng Yuan retweeted
In PRH, we argued that representational geometry is converging, but it's still an open question exactly in what sense. I want to share some new evidence that might clarify the picture. The evidence comes from our work here, led by @dominik_schnaus: dominik-schnaus.github.io/un… 1/n
5
25
217
9,740
Zhecheng Yuan retweeted
(1/5) Everyone says dexterous manipulation needs real-world data. These policies were trained with none. Our robots mesh gears, thread nuts, and insert pegs zero-shot from RGB, and the same recipe works for locomotion across terrains much harder than before. The missing piece wasn't a better RL algorithm, but a small tweak to what the robot practices. Today we're releasing Success-Guided Sampling (SGS), a simple change to the task sampler that unlocks new levels of dexterity and agility with large-scale, end-to-end sim2real RL. No demonstrations, no per-task reward tuning. Just PPO at scale. Check out the policies yourself at sgs-rl.github.io! All videos at 1x speed! 🧵 🔈
13
58
34
314
43,665
Zhecheng Yuan retweeted
Can a humanoid learn robust whole-body loco-manipulation directly from a human? No robot, no teleop, no retargeting. Meet Workhorse 🐴: our G1 topples and climbs a suitcase, catches a thrown box, and sorts boxes with hands and a kick. Fully autonomous. More videos at hybridrobotics.github.io/wor…
14
53
18
246
40,314
Zhecheng Yuan retweeted
🤖 Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from? 📼 Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention. 🧠 VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library? 🧵 Video2Skill: From Streaming Experience to Reusable Embodied Skills Why it matters: 📈 It can label much more data for training agents. 🔁 It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new? 📄 Paper: arxiv.org/abs/2609.36691 🌐 Project: andyzworks.github.io/video2s… 🤗 Data: huggingface.co/datasets/Ster…
2
7
1
28
5,165
Zhecheng Yuan retweeted
由于最近读论文的需求激增,加上本人已经受够了沉浸式翻译和DeepL这些翻译软件的垃圾排版和垃圾翻译,受够了Zotero这些old软件的垃圾文献管理,所以忙里偷闲搓了个玩具 EasyRead就是为了让你更舒服地读论文,不用在claude和gpt这些应用里面很难受地读(老艺术家喜欢原汁原味地读全文而不是让大模型直接总结),支持文献管理、自定义模型、原文对照、文献搜索与pdf导入(搜索我只试了arxiv,原谅我,有人用了再说)、能在阅读页直接对话(这不爆杀alphaxiv?),唯一的缺点是烧自己的额度,不过也没多少 项目链接:github.com/Edwardxlai/easyre…,点个star支持一下呀
42
201
12
1,246
103,866
Zhecheng Yuan retweeted
Argus: an open-source robotics data annotation and quality pipeline. With new frontier VLMs like GPT-6 Astra, we can generate rich, high-quality annotations for robotics data for both training and dataset analysis. Argus delivers detailed, timestamp-level annotations while catching issues like mislabeled instructions, sped-up recordings, swapped camera streams, and unflagged operator mistakes. To support better data quality for the robotics community, we’re open-sourcing Argus for anyone to use 🧵
66
68
22
583
84,080
Zhecheng Yuan retweeted
What will be the “RLHF” moment for robotics? What will it take to get robotics to where LLMs are today and beyond? New blog post with @chelseabfinn sharing some thoughts on the state of RL for frontier robotics models and what's missing 👇 Blog: pd-perry.github.io/posts/pos…
19
87
14
518
127,190
Zhecheng Yuan retweeted
CS academia is dead. So where does that leave AI PhDs like me? Did my ICLR reviewer bidding today. Skimmed a few abstracts, and the methods are the exact same recipe I learned when I got into 3DV two years ago. Swap in a newer video gen base model and boom, new paper 🥲 Makes you wonder how many of the 60k submissions were actually thought up by AI. Meanwhile, World Labs' Atlas has basically solved 4D scenes. Academia is so far behind it's not even funny. I still remember the day GPT-6 Astra dropped. My feed was flooded with GPT + Blender doing inverse graphics and GPT driving robot arms through manipulation tasks. The results were so good I literally had to sit down. A year ago, I was dead sure LLMs could never have spatial intelligence. A year later, Astra slapped that belief right out of me 💀 There's no doubt Astra was post-trained on tons of 3D and manipulation data, and it's only going to get bigger and faster. To me, that means any domain that can be represented symbolically, with clean benchmarks for RL, is going to get swallowed by LLMs. Next to real LLM intelligence, most academic papers that add a bit of inductive bias and tune their way to SOTA are just roadkill waiting to happen. Sadly, these papers keep piling up and flooding every conference. It's inertia, plain and simple. We've been chasing SOTA for so long that it's hard to stop overnight. But make no mistake: the paradigms in a lot of areas converged long ago. What used to be research is now pure engineering, and the room for academics to add inductive biases is only going to shrink. So as a PhD student, I see two paths left. One: if you can't beat them, join them. Clean data, build infra, then go to industry and train foundation models. Two: go back to real science. Stop caring about squeezing out another 0.1% on a benchmark, and start asking why this works and that doesn't, with foundation models themselves as the object of study. As a friend put it: AI research might end up looking more and more like biology, except the organisms are silicon-based 🤣 Beyond these two, it's hard to see anything that won't get eaten by LLMs. Still, I'm pretty pessimistic. I can feel the value of human knowledge being eroded bit by bit. Everyone will get their own AlphaGo moment. What will academia even look like after this... 😭
38
124
28
1,187
165,621
Zhecheng Yuan retweeted
Gpt-6 astra can do in-context learning on mobile manipulation! • Different environment • Different camera angle • Different layout No text prompt, it infers from video It even chooses when to use end-effector or joint space Trace recording after ⬇️
35
54
15
522
67,492
Zhecheng Yuan retweeted
“Knock, knock.” Hi! We are PokeBot. We are here to bridge the mind and the physical world—advancing embodied AI toward true Physical AGI. We envision a future where robots step into the messiness of everyday life, taking on the work we dread, so they can join in the work we dream of. A future where a home with robots from PokeBot becomes a space of boundless human creativity — freeing our time to create, to explore, and simply to live. In every home. Across every continent. And one day, beyond this planet. The journey starts now. PokeBot—Born to live with us.
1
4
1
8
529
A year ago, I still needed manual work to make real2sim happen. Things are moving fast.
GPT-6 does real2sim end to end from robot demonstrations. I gave it only multi-view RGB and robot actions, and one simple prompt. It calibrated the cameras, built the object assets, performed physics system-ID, ran MuJoCo, and rendered in Blender.
19
2,852
Zhecheng Yuan retweeted
推荐阅读! GPT-6 Astra + Recurrent depth /Looped transformer
A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer". It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit. About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters." Yes, that's it. The looped transformer idea is just reusing layers in the transformer block. In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights. In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers. Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.) While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation. In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak. Also, the statement "the new technique works in a way that obscures some or all of the AI's reasoning, otherwise known as 'chain-of-thought'" is not necessarily true with respect to the looped transformer method. It's possible that The Information journalist refers to some other technique or misunderstood the looped transformer method. Reusing layers does not by itself suppress visible chain of thought. It adds computation in hidden states before the next token is emitted, just as ordinary transformer layers do. But based on the information we have, the only plausible interpretation here is that if a model uses more of these recurrent passes, it may need to generate fewer intermediate reasoning tokens. So then more of its computation happens in latent activations that cannot be read as text. But we would get the same effect if we were scaling up the model size, like GPT 5.6 Luna -> GPT 5.6 Sol.
4
19
125
21,656
Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs. Checkpoints + training/inference code + Technical Blog ⬇️ (1/6)
70
174
54
1,227
442,608
Zhecheng Yuan retweeted
We've hit 100 episodes! Here's a look back on our journey so far (website with some stats): robopapers100.com/ Some highlights in the thread 🧵:
8
13
8
84
24,087
Great work! Can’t wait to try this solver!
Human data is the cheapest source of whole-body mobile manipulation demos. Turning it into robot data is the hard part.🤖 WARP retargets OFFLINE human motion into replayable robot actions in closed form — precise, consistent, faithful to whole-body intent. 🌍 warp-retargeting.github.io/
5
671
Zhecheng Yuan retweeted
Opus 5 can do this with zero demonstrations despite not being trained for robots. On the flip side Opus took 7 minutes while GEN-1.5 took 7 seconds. 🧵
Replying to @GeneralistAI
Or when fine-tuned to place a block into a bowl, it can clear obstacles (like a piece of paper covering the bowl) to complete the task, despite that not being in the demonstrations.
22
23
8
306
81,382
Zhecheng Yuan retweeted
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
321
1,684
862
12,161
3,427,988
Zhecheng Yuan retweeted
Excitement about world-action models and robot learning has never been higher — they promise a way to use human egocentric data to train massive robotics models which can provide the “GPT” moment for robotics and unlock general-purpose embodied intelligence. And yet there’s been little concrete demonstration of scaling in robot learning. @DynaRobotics aims to change that, with an in-depth look at how scaling works as they approach 1 million hours of training data. @JasonMa2020 @tianyurobot @_anhquanpham and @ChetBhateja joined us to tell us more. They show that as the amount of data they use in pretraining increased, they saw predictable, statistically significant gains on accuracy metrics on held-out data (data not seen during training). They go on to talk about what they learned, and show how this can be applied to many different problems. Watch Episode #99 of RoboPapers, with @micoolcho @chris_j_paxton @DJiafei today to learn more!
1
10
5
75
26,210
Zhecheng Yuan retweeted
🚀 ACE-Data-0 — a large-scale synchronized multimodal dataset🚀 👁️ Ego & multi-view exo videos 🧍 Human motion ✋ Hand motion 📦 Object motion 🎙️ Audio 🧤 Tactile sensing They are all aligned on a unified timeline and a shared coordinated system 📖 Blog: ace-data-engine.github.io/AC… 📄 Tech report: arxiv.org/abs/2607.28625 🤗 Hugging Face: huggingface.co/datasets/ACER……
6
30
4
150
35,942