@VmaxAIi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States Android App
Account-level information from X, not a live location or the device used for a specific post.
RL research focused on open-ended learning
San Francisco
Joined December 2024
- Tweets45
- Following1
- Followers507
- Likes391
Nethack is a really underrated benchmark
astra playing NetHack:
- score: 48,978
- BALROG progress: 49.4%
- max dungeon depth: 17
- max experience level: 14
important note: this is a cherry-picked episode (best of 10), and Astra has been able to modify its own game-playing harness
no privileged information or unfair game modifications, but it uses tools such as pathfinding to previously visited landmarks and code-written trackers for previously observed entities. these offload a fair amount of low-level reasoning, while Astra remains responsible for choosing what to do at every step
imo, the ability of LLMs to autonomously create tools that amortize low-level control costs (and then interpret and use their outputs in-context) is remarkable, and one of the reasons i'm excited about LLM agents. like, we could run the same tool computations and... concatenate the outputs into a flat vector for an MLP... but it would still have to learn what those outputs mean and how to use them. with LLMs we get free hierarchies and natural separation of responsibilities...
anyway, i'm now running episodes with the default settings from BALROG for a more fair comparison
the auto-research programme continues. will report back!
full run below ⤵️
Vmax retweeted
Interesting details from @deepseek_ai on the benefits of scaling environment/task synthesis.
If you want to scale open-ended environment generation for your own use case, sign up to the @VmaxAI waitlist - link below
Vmax retweeted
Greatest product placement ever. Placing the Vmax name in the latest Spiderman movie. 👏
Congrats @VmaxAI
Vmax retweeted
Legendary discussions at the office. @rronak_ discussed applying and scaling up SDPO for continual learning in the real world to prevent context degradation and more
@kevingu gave a talk on learning from production traces for knowledge-work tasks that have no direct verifier: how to generate reward signal without ground truth and keep organizational context current as the work changes
@matthewjsargent explained how to stabilize asymmetric self play and new research fields in open ended learning
Thanks to all those who came out and asked great questions we’ll be hosting more of these soon!
Most agents are static once deployed. Trained once, frozen, and brittle outside their training distribution.
We’re bringing together researchers and builders working on systems that keep learning, decompose problems into reusable pieces, and improve their own reasoning over time. Talks are from @kevingu CTO of Thirdlayer, @rronak_ CEO of Trajectory, @matthewjsargent CEO of Vmax followed by open Q&A.
Join us! luma.com/l666hxiw
You should try Silico. Not only for mech interp, but for AI Research in general.
Goodfire have put a lot work into building a great interface for long-horizon agent-driven research.
Vmax retweeted
CISPO outperforms GRPO/DAPO by capping a token's IS weight instead of zeroing out its update, so rare "fork" tokens ("Wait", "However") keep contributing gradient across updates.
Using CISPO, @lorenz_wlf reaches a new Sokoban Speedrun record; for the first time under 20min!!
Following the blog post from our collaboration with @GoodfireAI, the arxiv paper for PROPEL is now available.
Exciting new research led by @lorenz_wlf to accelerate RL task generation using mech interp methods in collaboration with @GoodfireAI
Training a model to generate RL tasks not too hard, not too easy costs many solver runs per task.
PROPEL predicts difficulty via a probe on its activations instead, amortizing cost and speeding up generator optimization.
New open-ended RL research from @Vmax + @GoodfireAI.
Exciting work on unix environment generation led by @radbadgeoffbrad!
Our designer @wwwjim has made something really special for this blog post.
Vmax retweeted
Vmax is building an open-ended learning system that generates and optimizes itself on tasks that it creates, avoiding human bias that may corrupt optimal learning curricula.
In PopuLoRA, we instantiate this as co-evolving populations of LLMs performing asymmetric self-play.
Vmax retweeted
We are so excited to have @tensorfi joining @VmaxAI!
Maxwill joins us from @Meta, where he was working RL and LLMs for recommendation. Previously, he has also worked @Tesla on the autopilot team and also in Quant finance at Kronos research. He also holds an MS in CS from Georgia Tech.
Maxwill simultaneously understands pre-LLM RL fundamentals but also how to scale pipelines for RL training for modern recommendation systems.
Maxwill is already levelling up our pipeline for automated environment design, pushing multiple PRs as soon as he joined.
Really excited about the velocity of his contributions and excited to share more soon.
Welcome Geoffrey!
So excited to welcome Geoffrey Bradway as Member of Technical Staff @VmaxAI.
Geoffrey is a rare catch. He was an engineer at @GoogleDeepMind, Google for Youtube and also has experience in early stage companies, having been a previous @ycombinator founder and also VP of engineering at @numerai.
Fitting the Vmax DNA, he has experience with RL before it was cool (doing RL all the way back in 2014).
Outside of work, Geoffrey does some really cool art with robotic drawing machines.
Cannot wait to share more about what he is cooking
Vmax retweeted
PR review is one of the fast growing categories in AI for SWE, now you can benchmark agents on *real* PRs
Introducing Code Review Bench v0: codereview.withmartian.com
The first independent code review benchmark. 200,000+ PRs. Unbiased. Fully OSS. Updated daily.
Tool performance highlights 🧵👇
Featuring: @augmentcode @baz_scm @claudeai @coderabbitai @cursor @GeminiApp @github @graphite @greptile @kilocode @OpenAIDevs @propelcode @QodoAI
Vmax retweeted
So excited to have @lorenz_wlf join @VmaxAI as a research fellow this spring!
At NeurIPS last year, we caught up with Lorenz, realised how aligned he is with our research vision and invited him to join us shortly after.
Lorenz comes from the @FAICDT1 programme at UCL (where I did my PhD also) and is supervised by @mircomusolesi.
Previously he worked on differential privacy and personalized recommender systems at Apple and did his undergrad in mathematics and statistics at Imperial College London.
Lorenz’s research focuses on RL, RLHF and modular continually learning RL agents. He has contributed to papers in ICLR, TMLR and AI STATS.
So excited for him to join us and accelerate our efforts on unsupervised environment design.
You can read Lorenz's research in the replies.
Much more to come.
Vmax retweeted
22/ Reinforcement learning, but make it automated. @MavorParker & @matthewjsargent showed us how they’re generating long-horizon environments at @VmaxAI.
vmax.ai/
welcome Roger!
Replying to @VmaxAI
@VmaxAI is excited to have @creus_roger joining us as a research fellow!
Roger is joining us from @Mila_Quebec where he works with @pcastr and @GlenBerseth.
Roger Creus Castanyer is a brilliant RL researcher working on exploration, credit assignment, and skill discovery.
He is also fresh off of a NeurIPS spotlight and a recently accepted paper to ICLR, you can find more of his research in the comments.
Roger is significantly accelerating our research on automated environment design - looking forward to sharing what he is cooking!
Vmax retweeted
Replying to @MavorParker @VmaxAI
As an initial step in this direction, we have built on top of methods like SWE-smith and BugPilot, adding to the list of repo profiles built by the swe-bench community
Vmax retweeted
This is a preview of many more tasks to come for Ares!
Replying to @joshgreaves_ml
ARES uses the Harbor task format ( @alexgshaw ). It comes with SWE-Bench Verified, TerminalBench2, SWESmith, and everything else in the Harbor ecosystem.
We're also releasing 1k new JavaScript tasks with @VmaxAI ( @MavorParker @matthewjsargent ) to help the ecosystem grow.