@heimmoritz

creating intelligent robots @microagi | bitter-lesson pilled

SF/LDN
Joined April 2015
In case someone wants to work from any of our cool offices (Munich, SF, Zurich, …), my DMs are open :)
I got to see this office and I think I agree
19
876
Most labs today don’t really look at their data or just evaluate by looking at samples. And they don’t have to. Since they only keep purchasing if data leads to model improvements in experiments, this pushes quality control onto the vendor. OS doesn’t have this incentive.
We (@PantheonInc) have been working on a robotics data quality pipeline that uncovered a series of major problems in public robotics datasets, especially for world modeling. To improve the quality of data available to open-source robotics, we're publishing annotations for four of the most popular datasets. Some examples of issues, and our report 🧵
22
2,088
moritz💭 retweeted
Today we release FLUX 3 Action - an open-weight 7B world-action model, achieving SOTA performance and efficiency on various leaderboards such as RoboLab. I am incredible proud of the team - and both backbone and embodiment-specific finetunes are open weight, alongside the training recipes!
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
23
32
3
347
24,687
This is a good reflection of many of the challenges in human data for robotics. Coming out of LLM expert data, I really appreciate how much easier operations were: One good labeler could drive 100.000s of revenue. In robotics, diversity is the most important. Coordinating 100.000s of people collecting millions of tasks in 10.000s of environments with custom hardware you have to develop and source yourself is actually really hard. And even after solving that, you need to reinvent the complete computer vision stack to make the data actually valuable. The only thing I disagree with is the outlook. I believe that data will remain valuable and crucial for much longer - although the formats might chance. Overall, I‘m very happy with what we as a team have built to collect millions of hours enabling researchers to explore scaling laws. Really looking forward to deploying ever more robots with these models in the future :)
After 2+ years in the robotics data space, we are shutting @Eidon_AI down. The thesis was right. But the business is brutally hard. We close this chapter by open-sourcing everything we built and sharing lessons for anyone venturing into the space.
5
4
50
13,978
If you wonder why 20-year-olds run laps around giant corp. Enterprises are late-stage Versailles: People care much more about status than results
13
769
kimi? oh you mean the claude wrapper?
13
2,880
Insanely proud to have this team. We just beat every existing other SLAM out there.
Today we're announcing microSLAM, a monocular SLAM system built by our computer vision team in Zürich with ETH Zurich's Computer Vision and Geometry Lab. It ranks first on the LaMaria benchmark for monocular SLAM, and holds up against systems that carry many more cameras and IMUs. Ours runs on a single RGB stream. Robot learning is bottlenecked on data, and the largest untapped source is people going about their work. A camera on someone's head for an afternoon in a plant is a record of how that plant actually operates, including all the parts nobody writes down. That data is only useful if you can recover the geometry, and geometry from one moving camera in a world that will not sit still is the hard version of the problem. Solving it monocular is what makes the data cheap. Rigs do not scale to every worker. Glasses do. microSLAM turns that footage into interactive 3D environments where behavioral foundation models can be trained. One person walks a factory floor. A robot learns to navigate it.
14
902
Europe is becoming concerned about us working too hard on its prosperous future😅
POV building a startup in Europe: Government knocks on your door on a Wednesday morning and interviews every single employee to make sure they dont work more than 40 hours per week (not kidding). No wonder Europe is hollowed out and we've become clowns compared to the US & China.
2
11
807
moritz💭 retweeted
Replying to @microagi
@microagi london office vibes
3
2
32
1,069
Cool framework, quick thought: Robotics tasks are bounded (there’s limits to how well a bed can be made/an industrial task can be performed) Model research in physical AI is unbounded (especially if and when we see more training moving towards simulation & self exploration)
9
892
More than impressive. Every business will have a robot in the future :)
Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:
7
408
We are building the most talent-dense team here. Reach out if you want to work here with prior YC founders, Oxford dropouts, Debate Champions, and people building 7-fig revenues before finishing High School.
Britain started the first industrial revolution. We think it should lead the next one too, so microagi is launching in London to put AI-powered robots to work on British factory floors. We are hiring aggressively! Reach out via dm's
2
1
11
547
:)
microagi has raised a $55M seed led by Hummingbird with Northzone, LocalGlobe, Village Global and redalpine, ten months after founding. Atlas trains control policies on worker demonstrations and deploys them to production - starting with the world's largest industrial companies.
1
10
1,386
Let me know if you‘re in SF and need a free private chef! :)
Today, we are bringing shift to San Francisco. Private chefs, completely free. We started with cleaning in New York and Europe. Now it's dinner. One by one, we are making every service more affordable. Comment "shift" below and we'll DM you a priority booking link.
10
489