Undergraduate Researcher @HebrewU | NLP • LLM Agents • Interactive Learning 🚀
Joined June 2026
- Tweets61
- Following351
- Followers49
- Likes188
Pinned Tweet
📢 New Paper
Can LLM Agents Infer World Models?
We turn a well studied automata learning task into an agent benchmark, where a hidden automaton serves as a world model that an agent must reconstruct using query tools.
Our findings show that SOTA LLM agents can sometimes perform non trivial interactive discovery, but remain far less robust and efficient than classic automata learning algorithms.
Beyond overall performance, our analysis reveals recurring failure modes in planning, reasoning, and using accumulated information over long contexts.
🔹 Think this task is easy? Try it yourself:
reefmenaged.github.io/Agenti…
🔹 Watch an LLM tackle an automaton of your choice:
agentic-automata-learning.on…
🔹 Read the paper:
arxiv.org/pdf/2606.16576
Grateful to my amazing co-authors for making this work possible: @GiliLior @ravfogel @roeeaharoni @GabiStanovsky
New version of our paper is out!
In v2, we evaluate Agentic Automata Learning with more advanced harnesses and scaffolds, including an OpenHands coding harness and a ReAct style state tracking scaffold.
We find that agents can identify an algorithm for the problem and even implement it with coding tools. Yet in the interactive task, without coding tools, they make recurring errors in query planning, evidence integration, and hypothesis construction, reaching only 30% success on the hardest DFAs.
ReAct style state tracking improves performance, but agents still suffer from these recurring failures, leaving substantial room for improvement compared with classic algorithms in both success rate and interaction efficiency.
Check out the new version and read more 👇
arxiv.org/abs/2606.16576v2
📢 New Paper
Can LLM Agents Infer World Models?
We turn a well studied automata learning task into an agent benchmark, where a hidden automaton serves as a world model that an agent must reconstruct using query tools.
Our findings show that SOTA LLM agents can sometimes perform non trivial interactive discovery, but remain far less robust and efficient than classic automata learning algorithms.
Beyond overall performance, our analysis reveals recurring failure modes in planning, reasoning, and using accumulated information over long contexts.
🔹 Think this task is easy? Try it yourself:
reefmenaged.github.io/Agenti…
🔹 Watch an LLM tackle an automaton of your choice:
agentic-automata-learning.on…
🔹 Read the paper:
arxiv.org/pdf/2606.16576
Grateful to my amazing co-authors for making this work possible: @GiliLior @ravfogel @roeeaharoni @GabiStanovsky
Reef Menaged retweeted
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning arxiv.org/abs/2606.16576
Reef Menaged retweeted
now people think there's just no way to create good new benchmarks. but there is, and it'll look obvious once we find them. it's getting harder but it's still possible. and i think that creating good benchmarks is the best way to advance ai towards being more useful for humans
I think a meaningful part of Astra’s recent jump on ARC-AGI-3 comes from a technical change: it was also evaluated with a Provider Adapter harness, a richer agentic runtime that preserves state across steps, including hidden reasoning state and context compaction.
The key point is that users can’t fully recreate this themselves. You can build an agent loop, memory, and context management around an API, but the model itself remains frozen. The model provider, on the other hand, can do RL and post-training with the model running inside the same runtime it will use at inference time, teaching it how to actually use that persistent state.
That’s why I think this is the right direction for the industry: frontier models should be shipped together with their agent and state-management mechanisms as one system, just like with GPT-6 Astra. I don’t want just a stateless input-output API that I have to build an agent around. I want a native multi-turn system where the provider maintains the working state and the model has been trained to use it.
An advanced harness doesn’t just add capabilities to an LLM agent, it multiplies them: there are many problems the model learned how to solve during training, but solving them requires long-horizon strategies and maintaining a coherent world model across many steps. With a basic harness, the model is limited in its ability to turn that problem-solving knowledge into consistent execution.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Read our post on Astra and what these results mean: arcprize.org/blog/astra
Perfect, I can retire now
We’re excited to open-source Reef: infrastructure we’ve been building for agents (model + harness) to continuously evolve from any signals generated at inference time.
Check out the code — issues and PRs are welcome!
github.com/Human-Agent-Socie…
Great blog post covering many questions that come up when building an agent benchmark
A phenomenon I keep noticing across recent papers is that LLM agents terminate interactions prematurely without fulfilling the task, even when explicitly told they can continue interacting to gather more information. We see this in benchmarks like MT-PINGEVAL [1], where dialogue agents rush to guess instead of using extra turns to uncover critical context, and ProgramBench [2], where coding agents propose the hidden program before exhausting the tool that provides observations on the program. This persistent agentic limitation warrants deeper research!
[1] arxiv.org/pdf/2602.24188
[2] arxiv.org/pdf/2605.03546
Reef Menaged retweeted
Replying to @paulg
I would try to figure out why LLMs can write my essays but not clean my bedroom.
Then I would study topics in college and grad school that could help solve that problem.
I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals.
That last item is also what I would if I were 30, 40, 50, or 66 years old 😉
An interesting project we’ve been working on: From Words to Wins: Analyzing Pre-Tournament Interviews to Predict Tennis Tournament Success 🎾
We built a model that combines interview representations learned by an encoder with structured features such as player ranking, and used it to predict whether a player would perform at least as well as their recent average. In our setting, we found that adding the interviews did not improve prediction performance compared with a model based only on structured features.
I think the most interesting question still remains open: with the right architecture and data at scale, should we expect to see an improvement? Or is there simply not enough predictive information encoded in these interviews to help forecast tournament performance?
$100 for a single agent rollout is insane. Agent evals are becoming so expensive that cost itself is starting to limit who can do meaningful research. Model providers should create programs where researchers can apply to run novel evaluation benchmarks at heavily discounted rates. It’s in everyone’s interest: researchers get affordable access to frontier models, while labs get independent researchers testing their models and uncovering limitations they might otherwise miss.
I agree that evaluating agents with the minimal toolset required to solve the task is the right way to measure their actual capabilities, rather than the strength of the harness around them. For example, in code tasks, adding internet access may turn the benchmark into a test of retrieval if the solution already exists online. With only the tools necessary to solve the task, we get a cleaner signal of the model’s coding ability and how well it generalizes to new problems.
As recent as this past NeurIPS review cycle, I consistently get reviews asking "did you use a stronger agent scaffold than mini-SWE-agent"
@KLieret's mini-SWE-agent *is* a strong agent.
More importantly, i think the framing of "stronger harness", where "stronger" means more tools, skills, add-on's is questionable at times.
Increasingly so, i think "simplicity" in a harness (give model small set of tools, otherwise get out of its way) is the right way to go.
If the model truly needs a tool (e.g., repeatedly invokes 2+ actions), it can synthesize one for itself on the fly (see live-SWE-agent from @steven_xia_ @YuxiangWei9 @LingmingZhang)
Simplicity (just bash, no extra tools or fluff) is exactly what makes it strong.
Reef Menaged retweeted
My last question was more personal:
If you were a new grad today, would you choose Google as the start of your career?
I was seriously considering joining Google, so I wanted to know what Jeff would do in my position.
His answer surprised me.
He told me not to think of Google as one company.
There are pockets doing cutting-edge ML research that he finds incredibly exciting.
And if you get into the right setting, something you build this week can roll out in two weeks and impact a billion people.
But there are also plenty of product areas at Google that Jeff personally wouldn't find that interesting.
His point was that the size of Google means your experience depends heavily on where you end up.
You have to find the pocket where the problems are genuinely exciting to you.
Our 30 minutes were over.
What happened after changed how I approached the rest of my internship.
Reef Menaged retweeted
Every AI agent we use has a system prompt behind it. These hidden instructions define the rules, tone, and priorities for every reply. But you never see them. Today, we are very excited to announce SystemPromptIndex, the largest open system prompt library indexing 1000+ system prompts from 400+ products, along with AISPA, the first assurance standard for AI system prompts.
Reviewers sometimes seem to assume unlimited compute, time, and annotation budgets. In reality, you cannot check everything and have to prioritize. If it were their own paper, they would probably make the same tradeoffs.
This quoted post is unavailable.
Reef Menaged retweeted
Anthropic Engineer Andrej Karpathy:
"The biggest mistake in AI right now: people are forcing agents to work instead of mastering the model first.
We made that mistake in 2016 at OpenAI. It cost us 5 years."
What Karpathy actually means:
step 1 → stop forcing your agent to do everything. Understand the model underneath first.
step 2 → demos are easy; products take a decade. Self-driving proved it. If you skip the foundation, everything breaks.
step 3 → the agent is not the product. The foundation is. Build that—and agents emerge on their own.
"You're building agents right now. You're at the forefront. Not OpenAI. Not DeepMind. You."
watch - bookmark