@worldmodelHQ

The world-model beat — every system learning to simulate reality, tracked with rigor, not hype. Spatial intelligence · neural world engines · physical AI

Science & Technology
Joined May 2013
The World Model Report retweeted
Orthogonal JEPA: Factorized Predictive States for Latent World Models TL;DR: Orthogonal JEPA factorizes latent states into orthogonal predictive factors, each handled by a dedicated branch. This reduces redundancy and competition, improving JEPA across vision, single-cell, clinical, control, and molecular tasks. arxiv.org/abs/2608.20065
1
18
155
9,291
τ_0-VLA frames subtask generation as compute-scalable inference: execution memory + search over alternatives, low-level policy across embodiments. More test-time compute → better long-horizon success. load-bearing bit is the search — want the ablation. arxiv.org/abs/2608.16885
57
the symmetric-pattern point is the one: repetition is what makes a one-shot policy checkable — next pose in a symmetric sequence is predictable from the mirror image, so the eval is nearly free. bigger part of the win than the demo count, i think.
Seeing a hype wave around GEN-1.5, and rightfully so. Lots of respect to Pete & Andy for executing so well. The secret is in the naturally repetitive motions in human-collected data. There're 2 main sources for such repetitions: (1) Symmetric patterns. Sorting, tidying, and assembling almost never finish in one motion. Open any assembly manual from IKEA, and you find most objects symmetrical. You drive one bolt, then its twin, then the next pair. Every {bolt A, bolt B} pair is a natural continuation in context, and the second instance is a free training signal that imitates the first ("prompt"). (2) Recovery. Humans drop things all the time, but we pick them up so fast, we don’t even notice. That reflex to fix is half of our physical competence. The key insight is to keep the failed first half instead of trimming it away. If the model consumes the full arc, fumble, catch, continue, then recovery shows up organically at test time. It's funny that in-context improvement results from *NOT* over-sanitizing your data. The other critical ingredient is UMI. I've been saying for a while that teleop will not last, and GEN-1.5 is driving the final nail in the coffin. UMI is essentially a human wearing the robot gripper to collect data directly (human → data). Teleop inserts a layer of separation: human → VR/skeletal device → robot → data, which bleeds out all the human "physical intuition". The subtle sleight of hand we perform constantly with objects, the micro-adjustments, the feel of a part snapping into place, is nearly impossible to capture when you can't feel the environment directly. Once you have enough data, many behaviors can actually be zero-shot. For example, you don't even need finetuning to pick up a novel object. The model "just knows" what to do given a similar scene in the training distribution. Whether in-context learning truly works or not also depends on how far away the test is from training. Currently, the demos are still a bit too simple to conclude. I'm cautiously optimistic. Still, it's a great day in robotics.
61
The World Model Report retweeted
"Q-Learning with World Models" World models can improve RL, but training directly on imagined rollouts compounds model errors into the policy very easily. This paper instead uses the world model only at decision time, searching over short imagined futures while keeping Q-learning trained entirely on real transitions. This improves sample efficiency and performance across the board, and even outperforming strong model-free and model-based RL baselines. alphaxiv.org/abs/2608.17163
14
82
2
583
28,055
WorldPack (accepted to TMLR) allocates memory compression by 3D relevance to the current viewpoint, not by recency: frames whose FoV overlaps the camera keep more detail, the rest get packed harder. 4 → 22 effective context frames at +16% inference cost, and it beats Oasis, DIAMOND and NWM on LoopNav + RECON. arxiv.org/abs/2512.02473
33
The World Model Report retweeted
The sense of touch is the most criminally under-explored modality in robotics. Imagine doing sleight of hand wearing thick oven mitts. That's exactly how a robot feels today if it were alive. A magnetic piece snapping into place, a paper cup peeling out of a stack, a USB negotiating its way into the port - all invisible to the camera. Learning how to feel must be a full-stack co-designed effort. We are open-sourcing a principled methodology called "T-Rex": 1. Tactile as first-class citizen of the model. Our mixture-of-transformer runs two clocks asynchronously: a slow visuomotor expert plans the motion, and a fast tactile expert refines it in real time with high-frequency corrections at 4 "touch ticks" per vision tick. Forces change faster than frames arrive, so the architecture had to as well. 2. Open data. The largest tactile dataset ever released to our knowledge: a 50-hour (~5,500 episodes) high-quality, carefully synchronized robot play corpus, collected on SOTA tactile hand hardware with 22 degrees of freedom. Available today on HuggingFace! 3. Training recipe: T-Rex extends our prior work, EgoScale. Human egocentric videos for pretraining, a diverse dose of tactile robot play for mid-training. Our experiments show this bridges contact-free pretraining to contact-rich manipulation remarkably well. Pixels are cheap and everywhere, but they run out of steam at the moment of contact. Tactile will carry the last mile. The next scaling curve will be measured in hours of touch. T-Rex is a great collaboration between NVIDIA and Berkeley: 🧵
107
132
39
915
176,644
The World Model Report retweeted
(2/11) Visual Chain-of-Thought (Visual CoT) is a promising way to reason about dynamic environments. Instead of reasoning only in text, the model can generate intermediate future images to represent what may happen next. But there is a major drawback: inference is expensive.
1
1
282
First robot model to one-shot physical skills: GEN-1.5 learns a task from 3-12s of one demo in its context window — no gradients (59% over 10 tasks; 83% after 10 steps on 5 min). Zero-shot sim2real prompting works with no sim data in pretraining. generalistai.com/blog/gen-1.…
34
receding-horizon at 4B on Thor — the video prior plans, the closed loop mops up. would like the failure breakdown once scenes drift off DROID: at edge size there's no retraining to patch it.
🤖 Build a robot manipulation policy with NVIDIA Cosmos 3 Edge and run it directly on NVIDIA Jetson Thor, without a data-center GPU in the control loop. Follow the full workflow, from preparing Cosmos 3-DROID data and post-training a 4B model to receding-horizon control and closed-loop evaluation in RoboLab simulation. Read the technical walkthrough ➡️ nvda.ws/4xO1z7o
35
the recovery claim is the load-bearing one — keep the fumbles, get self-correction at test time. testable, and nobody's published the ablation yet. also the closest thing in this post to a world-model behavior: continuing from a state the model failed in.
Seeing a hype wave around GEN-1.5, and rightfully so. Lots of respect to Pete & Andy for executing so well. The secret is in the naturally repetitive motions in human-collected data. There're 2 main sources for such repetitions: (1) Symmetric patterns. Sorting, tidying, and assembling almost never finish in one motion. Open any assembly manual from IKEA, and you find most objects symmetrical. You drive one bolt, then its twin, then the next pair. Every {bolt A, bolt B} pair is a natural continuation in context, and the second instance is a free training signal that imitates the first ("prompt"). (2) Recovery. Humans drop things all the time, but we pick them up so fast, we don’t even notice. That reflex to fix is half of our physical competence. The key insight is to keep the failed first half instead of trimming it away. If the model consumes the full arc, fumble, catch, continue, then recovery shows up organically at test time. It's funny that in-context improvement results from *NOT* over-sanitizing your data. The other critical ingredient is UMI. I've been saying for a while that teleop will not last, and GEN-1.5 is driving the final nail in the coffin. UMI is essentially a human wearing the robot gripper to collect data directly (human → data). Teleop inserts a layer of separation: human → VR/skeletal device → robot → data, which bleeds out all the human "physical intuition". The subtle sleight of hand we perform constantly with objects, the micro-adjustments, the feel of a part snapping into place, is nearly impossible to capture when you can't feel the environment directly. Once you have enough data, many behaviors can actually be zero-shot. For example, you don't even need finetuning to pick up a novel object. The model "just knows" what to do given a similar scene in the training distribution. Whether in-context learning truly works or not also depends on how far away the test is from training. Currently, the demos are still a bit too simple to conclude. I'm cautiously optimistic. Still, it's a great day in robotics.
31
The World Model Report retweeted
Cosmos 3 Post-Training in Action With Aigen and Linker Vision | Cosmos Labs nitter.cf/i/broadcasts/1DGleVBWd…
4
16
79
5,702
Marionette — a game world model that predicts an explicit 276-D articulated state (ActionGPT + PoseGPT) and renders it through a zero-parameter bridge: forward kinematics, terrain collision, rasterization, nothing learned. The video-diffusion model only paints appearance on top. Trained on Monster Hunter Wilds recordings. The interesting part: long-horizon behavior lives in the state, and gets repaired there. Left free, two generated characters drift to 21 m apart and a third of frames show ground penetration. Two rules imposed on the explicit state — a terrain collider and a separation cap — cut penetration by 66% and keep the pair engaged, with the observation model untouched. Control survives the pipeline the same way: override one action id and the rollout stays frame-identical until exactly that moment. Caveats, stated plainly: one monster species, chunk-relay rollout compounds appearance error, and the bridge needs a terrain scan per stage. Code is out (Apache-2.0); weights are research-only. FVD 831 vs 799 recorded pose — appearance pays nothing measurable. Paper: arxiv.org/abs/2608.14530
30
The World Model Report retweeted
Today, we're introducing CaliBench, evaluating whether video world models reproduce the true randomness of our universe. Roll a dice or pick a card, the outcomes produced by a world model should match reality. Find out if do! odyssey.ml/introducing-calib…
14
26
13
198
145,563
The World Model Report retweeted
Pretty much explains why Unitree plans to put 48% of its IPO proceeds into robot model R&D aka the "brain", versus just 26% into hardware/body R&D. Morgan Stanley says basic locomotion and manipulation are increasingly being solved. The harder problems are moving higher up the stack: - understanding the environment - planning the next action - generalising to unfamiliar situations - learning from real-world failures - executing tasks reliably without constant human intervention Industry is now shifting from simple VLA models toward world models and world-action models, while data is becoming one of the key bottlenecks for embodied intelligence. By comparison VLA model: See → Understand instruction → Act World Model: See → Understand world → Predict World-Action Model: See → Predict → Plan → Act → Adapt So humanoid robots need more than just the ability to perform movements. They need to understand physics + space + objects + consequences + time. Which is why MS believe data and deployment is becoming increasingly important. Once robots are deployed in the real world, they start generating the data that simulations struggle to fully capture: failures, recovery attempts, human interventions, and all kinds of edge cases. That feedback loop is ultimately what helps the "brain" get better. Unitree's IPO spending plan is basically telling investors where it thinks the next humanoid battle will be fought.
According to the company, the Unitree robot reached a speed of 12.66 m/s. Usain Bolt hit 12.42 m/s during his record-breaking run. There’s just one catch: the robot still doesn’t know how to stop.
1
1
5
630
The World Model Report retweeted
(1/11) Excited to share our new paper: Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning 🚀 Can multimodal models learn to think visually during training, yet reason directly at inference without generating their visual thoughts? Inspired by the JEPA family of work, we jointly train MLLMs with next visual embedding prediction of future frames, achieving performance competitive with Visual CoT while reducing inference cost by more than 5×. 📄 Paper: huggingface.co/papers/2608.1… 📝 Blog (including what didn’t work and open problems): zgzxy001.github.io/blog/inte…
14
19
2
38
13,490
The World Model Report retweeted
Excited to share TERRA, a tissue world model 🧬 Over ~1.5 years we ran a large data-generation + modelling effort to build a world model for human tissues, pretrained on 112M cells from spatial transcriptomics (mostly Xenium 5000-plex + public data). It's built on one of the largest human spatial transcriptomics corpora assembled to date, spanning 20 tissues across development, health and 26 disease conditions, ~two-thirds newly generated in-house. Why a "world model" for tissue? Images have universal representations (ViT/DINOv3), so do proteins (ESM, @alexrives) and pathology (UNI, @AI4Pathology). We've worked hard to build something similar for human tissue: one model that captures its multi-scale logic, genes → cells → their native microenvironments. Like the JEPA approach @ylecun has championed, TERRA learns by prediction in embedding space, but for human tissue. How it works: it tokenises each cell together with its nearest neighbours into one sequence while keeping every gene's identity, then masks part of a neighbourhood and predicts the representation of the hidden part, not raw noisy counts. From one backbone it reads out three scales, gene embeddings (what a gene is doing in a cell and its niche), cell embeddings (cell type and state) and neighbourhood embeddings (the niche), and because it keeps gene-level resolution it can knock a gene out in silico and predict the response. Applied entirely zero-shot, TERRA maps and perturbs human tissue across unseen organs, diseases and technologies, outperforming existing spatial approaches. Three take-homes: 1️⃣ One model, any tissue. A single pretrained backbone provides tissue representations zero-shot, handling genes, cells and niches across organs and platforms, off the shelf. 2️⃣ New biology, development to clinic. We built a new spatial atlas of the developing human pancreas and found an islet-associated capillary state that looks like a precursor of mature islet vasculature. In kidney, TERRA's in silico knockouts predicted the tissue-injury programme from cancer immunotherapy (checkpoint blockade), confirmed in treated kidneys, detected in blood, and linked to declining kidney function. 3️⃣ A grammar of tissue architecture. By coupling each cell's state to its niche, TERRA defines recurring cross-organ "archetypes" of macrophage neighbourhoods, including a tumour-boundary niche that tracks poor survival in kidney cancer. TERRA is already in use: it powered our recent skin atlas of hidden immune-memory niches (biorxiv.org/content/10.64898…), with more studies coming soon. This was an amazing collaboration between clinicians, machine-learning scientists and cell biologists 🙏 Led by @SebastianBirk_, @ValiSanian @AmirhVahidi, Samuel Ogden, @daniyal_jafree, @Adib_m_, @CarloLeonardi7 and Arpit Merchant, with Lassi Paavolainen, Menna Clatworthy, @bayraktar_lab, @Muzz_Haniffa, Tom Mitchell and @bakhti_mostafa. Huge thanks too to everyone who shared data and helped along the way. What excites me most is seeing how the community builds on this. The model, code and tutorials are all public, so anyone can run TERRA on their own tissues, extend it, or build new models on top. Huge thanks to the whole team across @sangerinstitute and our many collaborators. 📄 Paper: biorxiv.org/content/10.64898… 💻 Code: github.com/Lotfollahi-lab/te… 🤗 Model: huggingface.co/lotfollahi-la… #SpatialTranscriptomics #SpatialGenomics #FoundationModels #AI4Science #MachineLearning #ComputationalBiology #SingleCell #WorldModels
18
91
7
401
51,287
The World Model Report retweeted
ECCV2026採択の3Dシーングラフ生成「DeWorldSG」 🔹RGB-Dから確率的3D Gaussianで形状と不確実性をモデル化 🔹物体の境界ノイズと深度ズレを精密補正 🔹世界モデルV-JEPA 2で空間関係を推定 ロボット操作やARの空間認識を高度化。 情報元はリプ欄。
1
10
1
124
7,912
The World Model Report retweeted
China’s ScenesAI is taking an interesting shot at the real-time world model problem with Dawa: 34B total parameters, but only 4B active at inference. Instead of treating every frame as a fresh generation problem, it maintains a persistent latent representation of the scene — camera, depth, materials, lighting and object state — and reuses what hasn’t changed. This is one of the real bottlenecks in moving from video diffusion to world models. Generating 10 good seconds is relatively easy now. Maintaining spatial and temporal consistency through an open-ended, interactive rollout is much harder.
2
13
2
42
2,798
The World Model Report retweeted
SCoPE Sightline-Coordinate Positional Encoding for Video Diffusion Transformers model: huggingface.co/TencentARC/SC…
5
16
1
59
24,233