@Murms23i
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
california
Joined August 2011
- Tweets12
- Following60
- Followers24
- Likes35
Mimi Elmasry retweeted
Excited to see S1 finally go public! 🚀 This release also marks exactly one year since I transitioned from vision research to robotics. If I had to share the single biggest lesson from that journey, it’s this:
Before asking "How do we train a better model?", we need to ask "How should the model behave at inference time?"
In academia, we tend to attack problems the same way: set up a benchmark, tweak the architecture or training recipe, and watch the score climb. Today, much of robotics research runs this similar loop: pit VLAs against WAMs, introduce a new loss function, test a different action representation. Model performance on benchmarks keeps improving, but inference-time behavior stays simplistic and fixed.
But my past year at @SkildAI taught me that raw model performance isn't the ultimate driver of real-world success. Inference-time behavior is.
The real world guarantees high stochasticity and endless out-of-distribution scenarios. The dominant deployment paradigm (taking the past X seconds of observations to predict the next Y seconds of actions) won't get us to general-purpose robotics, because every robotic action has compounding consequences. A vision model can guess wrong on a static test set with zero impact on the next sample. A robot cannot. It operates in a closed loop: its own mistakes create the very distribution shifts it must then survive. "Collect more data, fine-tune, redeploy" is not a scalable answer. To truly generalize, a model must adapt on the fly.
So the order of operations has to invert. First, define an inference-time behavior scalable and robust enough to handle the real world. Then work backward to shape data collection and model training. Everything in the S1 release stems from this inversion.
Language modeling already taught us this lesson: LLMs didn't crack complex multi-step math just because someone threw more tokens at pre-training. Instead, the breakthrough came from changing the inference-time behavior first (reasoning step by step) and then redesigning training around it. Betting that data scale alone will magically produce emergent physical skills is just as naive. Robotics needs the same shift.
S1 is proof that working backward from real-world inference changes everything. And this is just the starting point.
Check out the release below, and enjoy the rest of your day! 👇
Introducing S1, our new foundation model that learns from one example.
It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning.
Watch S1 operate in real-time via in-context learning:
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Ohh I see, Long-horizon tasks, deployed. Skild's horizon just got longer..
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Mimi Elmasry retweeted
This! Its not just research...its real deployment. @SkildAI
I am posting this as I am deploying S1 on a customer site. Words can't describe how excited I am.
This was a tremendous effort from the team across data, infra, training, inference, and hardware. Personally I think it is a giant step closer to general-purpose robots creating real value in our lives.
For homes, you can easily teach your robot a new task, or tell it your preferred way of doing a task.
For businesses, you can plug in your SOP and start seeing value creation from Day 1.
Now the foundation has been set, time to bring it to the real world. Still a lot more work to do, stay tuned.
Mimi Elmasry retweeted
Skild AI never shipped a robot. It's worth $14B. Hundreds already run on its software inside NVIDIA's Houston factory. thenextweb.com/news/skild-ai…
Robotics is a data problem.
Today, we’re partnering with @ABBRobotics, @Universal_Robot, and @NVIDIARobotics to deploy the Skild Brain across real-world industries from manufacturing to factory lines.
This will help us build the world’s biggest data flywheel for physical AI.
Announcing Series C
We’ve raised $1.4B, valuing the company at over $14B
With this capital, we will accelerate our mission to build omni-bodied intelligence 🚀
skild.ai/blogs/series-c
Mimi Elmasry retweeted
I have been working on scaling robot learning via videos for almost 5-10 years (initially with @xiaolonw and then other efforts such as R3M, NRNS, and recent factorized policy work from @mangahomanga )...Now we are pushing these directions to scale at @SkildAI
I think scaling via videos is going to play a core role in robot foundation models and this is just the beginning. So excited about the future!
Humans learn by watching. Robots should too.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Mimi Elmasry retweeted
This is your quarterly reminder that I have a job, everyone watch the video and read the blog post
Humans learn by watching. Robots should too.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Learning by Watching Human Videos youtu.be/YRmjBdKKLsc?si=bVjz… via @YouTube