@dimentaryi
iAccount based inPoland
About this account
- Account based in
- Poland
- Connected via
- Poland App Store
Account-level information from X, not a live location or the device used for a specific post.
ml research engineer | physical learning, evals prev google, kaggle, tcl @anyactai
SF
Joined January 2023
- Tweets245
- Following639
- Followers1.1K
- Likes1.1K
Pinned Tweet
Astra, can you write a Python script for the Fibonacci sequence
i mean physically
hilarious to watch how Opus 5.5 designs a 3d asset, prints it, and controls a robo arm to take it out of the printer
i basically asked it to design something nice based on @anyactai logo
Dmytro Hrybov retweeted
The state of the @jointhebridge house on Sat: @dimentary is terminally locked in with his robots
so i was curious to see how opus 5.5 compares to astra in token usage when acting as a robotic policy
TLDR: opus used ~6x more tokens but astra is still more expensive on average. in this particular example, astra got further and completed more tasks before giving up though
was fun watching their approaches to failure recovery
nobody beat the triangle
trained a smol RL policy to rotate two baoding balls with a Sharpa hand
interestingly, astra managed to define most of the RL environment by itself (including verifiers and rewards) but didn't bother by the fact that balls go around the ring finger while inspecting the policy
added this including training script to the repo:
github.com/dimentary/llm-rob…
Dmytro Hrybov retweeted
I asked Opus 5.5 to create an interactive 3D visualisation explaining how robot actuators work.
zerotimedrift.github.io/insi…
Repo in comments
Dmytro Hrybov retweeted
Introducing FLUX 3 Action.
An open weights 7B World Action Model that achieves first place on the RoboLab benchmark.
It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.
FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA.
Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson.
Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next.
FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together.
We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
now opus 5.5 visualized how @physical_int π0.5 inference works on DROID, built from open code and paper
beautiful, asked opus 5.5 and astra to show how they see Kyiv in Monet style
there’s something peculiar to the fact that this is all coded, like when we had our early text to image models back in 2021
can you tell which is which?
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles
every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
compared GPT-6 Sol and Astra on my "robot drawing on the board" bench
Sol used ~25k output tokens vs ~22k for Astra. interestingly, Sol's controller drew twice as fast
to be honest, i expected stronger spatial reasoning from new Sol here, but it was way cheaper to run (almost 5x)
great paper, their model trained on 1k filtered RL tasks outperformed the one trained on 8k vanilla tasks (on the validation set)
8k vanilla tasks are sampled right after the test construction step 2 as i understand it
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
🤗: hf.co/papers/2609.22068
or these systems will implicitly emerge inside a single model with enough scale and proper data engineering #bitterlesson
right now, systems 1 and 2 in robotics communicate with each other mainly through text, one example is task planners prompting VLAs to execute subtasks
i think this will be a major bottleneck, it’s hard to express richness of the physical intent with text, if you think about it
e.g. when we try to grab a cup, our exact actions depend on whether it was hot/slippery/damaged somewhere/many other things
the way our brains “handoff” this information from system 2 to system 1 during actions feels way more integrated
right now, systems 1 and 2 in robotics communicate with each other mainly through text, one example is task planners prompting VLAs to execute subtasks
i think this will be a major bottleneck, it’s hard to express richness of the physical intent with text, if you think about it
e.g. when we try to grab a cup, our exact actions depend on whether it was hot/slippery/damaged somewhere/many other things
the way our brains “handoff” this information from system 2 to system 1 during actions feels way more integrated