@dimentary

ml research engineer | physical learning, evals prev google, kaggle, tcl @anyactai

SF
Joined January 2023
Astra, can you write a Python script for the Fibonacci sequence i mean physically
i’ve seen GPT-6 Astra draw in Paint, how about drawing with a robot in a physics simulation? asked it to build a MuJoCo setup and write a controller to draw Picasso’s dove using a robot arm and a five-fingered hand
26
54
14
617
70,105
hilarious to watch how Opus 5.5 designs a 3d asset, prints it, and controls a robo arm to take it out of the printer
55
83
32
1,388
107,054
i basically asked it to design something nice based on @anyactai logo
1
17
2,222
making it open the printer lid was the hardest part
1
11
1,549
beware, my robot is taking his first 3d printing class
10
2
112
4,194
Dmytro Hrybov retweeted
The state of the @jointhebridge house on Sat: @dimentary is terminally locked in with his robots
3
1
20
1,607
so i was curious to see how opus 5.5 compares to astra in token usage when acting as a robotic policy TLDR: opus used ~6x more tokens but astra is still more expensive on average. in this particular example, astra got further and completed more tasks before giving up though was fun watching their approaches to failure recovery nobody beat the triangle
2
3
1
34
8,256
trained a smol RL policy to rotate two baoding balls with a Sharpa hand interestingly, astra managed to define most of the RL environment by itself (including verifiers and rewards) but didn't bother by the fact that balls go around the ring finger while inspecting the policy
6
4
1
90
6,161
Dmytro Hrybov retweeted
I asked Opus 5.5 to create an interactive 3D visualisation explaining how robot actuators work. zerotimedrift.github.io/insi… Repo in comments
3
2
3
22
3,959
Dmytro Hrybov retweeted
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
77
235
54
1,804
168,104
now opus 5.5 visualized how @physical_int π0.5 inference works on DROID, built from open code and paper
love this format, visualized a V-JEPA 2.1 training step based on public materials
7
19
201
14,423
wish i could've visualized all papers like that when i was at uni
180
love this format, visualized a V-JEPA 2.1 training step based on public materials
Continúa la saga ahora con el diagrama de un modelo de difusión, a manos de Opus 5.5
1
2
1
33
15,408
beautiful, asked opus 5.5 and astra to show how they see Kyiv in Monet style there’s something peculiar to the fact that this is all coded, like when we had our early text to image models back in 2021 can you tell which is which?
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
4
25
1,865
compared GPT-6 Sol and Astra on my "robot drawing on the board" bench Sol used ~25k output tokens vs ~22k for Astra. interestingly, Sol's controller drew twice as fast to be honest, i expected stronger spatial reasoning from new Sol here, but it was way cheaper to run (almost 5x)
9
6
5
110
58,912
great paper, their model trained on 1k filtered RL tasks outperformed the one trained on 8k vanilla tasks (on the validation set) 8k vanilla tasks are sampled right after the test construction step 2 as i understand it
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself 🤗: hf.co/papers/2609.22068
11
1,241
or these systems will implicitly emerge inside a single model with enough scale and proper data engineering #bitterlesson
right now, systems 1 and 2 in robotics communicate with each other mainly through text, one example is task planners prompting VLAs to execute subtasks i think this will be a major bottleneck, it’s hard to express richness of the physical intent with text, if you think about it e.g. when we try to grab a cup, our exact actions depend on whether it was hot/slippery/damaged somewhere/many other things the way our brains “handoff” this information from system 2 to system 1 during actions feels way more integrated
2
8
979
right now, systems 1 and 2 in robotics communicate with each other mainly through text, one example is task planners prompting VLAs to execute subtasks i think this will be a major bottleneck, it’s hard to express richness of the physical intent with text, if you think about it e.g. when we try to grab a cup, our exact actions depend on whether it was hot/slippery/damaged somewhere/many other things the way our brains “handoff” this information from system 2 to system 1 during actions feels way more integrated
4
1
3
38
4,525