@NegarEmpr

Postdoc @UCBerkeley @BerkeleySky |👩🏻‍💻Prev @google, @MSFTResearch, @SpotifyResearch | 📚@UWaterloo | Interested in Information Retrieval

Berkeley, USA
Joined April 2017
So excited to share that I received the SIGIR Early Career Research Award! 🥹 I still remember attending my first SIGIR, watching someone receive this award, and thinking, “There’s no way that could ever be me.” And now, here I am!😵 This truly means so much to me. Thank you to the @ACMSIGIR community for recognizing my contributions. #SIGIR2026 P.S. Get yourself a PhD supervisor who matches your award-celebration energy 😄 So cool to celebrate this moment together! @claclarke
22
5
1
99
11,722
Negar Arabzadeh retweeted
Could long-context architectures “find all contradictions” in a science literature? Not yet! 🧵 We study a new class of "high-complexity” tasks whose difficulty scales quadratically with corpus size (as opposed to linearly), reversing common LCLM decisions! (block-sparse attention, hybrid models…)
11
55
7
259
34,656
Negar Arabzadeh retweeted
Impressive paper showing how much the first retrieval step matters for deep research agents. It helps to improve GPT-5.5 from 83.1% to 90.5% on BrowseComp-Plus with the same retriever and the same agent loop. It seems that the gain comes from the opening context. The authors propose Question's Gambit which runs once, before the agent starts searching. It splits the question into clues, turns each clue into complementary searches, pools the results, and reranks them. The agent then starts its loop with that ranked set already in context. The same change lifts GPT-5.4-mini from 68.1% to 79.0% and DeepSeek-v4-pro from 71.4% to 76.9%, and roughly halves calibration error for GPT-5.5. It costs between 2.3 and 5.3 extra tool calls per question. In an error analysis, only 3 of the 79 remaining GPT-5.5 errors come from the gold document never being retrieved. The other 76 happen later, when the agent previews, opens or uses the evidence. Paper: arxiv.org/abs/2609.14412 Chat with Paper: academy.dair.ai/papers/quest…
24
20
1
128
13,789
Thanks @omarsar0 for sharing our paper before we even got the chance to! 🙏 Loved your framing so much we borrowed it for our thread "same retriever, same agent loop, just a better opening move. ♛ "
Impressive paper showing how much the first retrieval step matters for deep research agents. It helps to improve GPT-5.5 from 83.1% to 90.5% on BrowseComp-Plus with the same retriever and the same agent loop. It seems that the gain comes from the opening context. The authors propose Question's Gambit which runs once, before the agent starts searching. It splits the question into clues, turns each clue into complementary searches, pools the results, and reranks them. The agent then starts its loop with that ranked set already in context. The same change lifts GPT-5.4-mini from 68.1% to 79.0% and DeepSeek-v4-pro from 71.4% to 76.9%, and roughly halves calibration error for GPT-5.5. It costs between 2.3 and 5.3 extra tool calls per question. In an error analysis, only 3 of the 79 remaining GPT-5.5 errors come from the gold document never being retrieved. The other 76 happen later, when the agent previews, opens or uses the evidence. Paper: arxiv.org/abs/2609.14412 Chat with Paper: academy.dair.ai/papers/quest…
1
5
14
969
1/ We’re excited to share Question’s Gambit module for search agents. ♛ 🥇Question Gambit got first on the BrowseComp-Plus leaderboard for recall with BM25 as our only retriever! 📈 It focuses on the first retrieval move in search agents and got 96.6% recall with about 6 times fewer agent search calls. Same BM25 index. Same agent loop. Just a better opening move. 📄 Paper: arxiv.org/abs/2609.14412
5
8
1
39
5,028
3/ High recall isn’t enough. The agent has to use what it finds. With GPT-5.5, Pi-Serini surfaces 94.4% of gold documents, but opens or cites only 56.1%. Question’s Gambit raises these to 98.1% and 75.9%. A much larger gain in evidence use than in retrieval.
1
1
7
243
4/ Finally, we're proud that Question's Gambit is the top academic submission on BrowseComp-Plus, at 90.5% answer accuracy. And we're honored to share the top of the leaderboard with @allen_ai @ai2_allennlp , @sailresearchco , and @mixedbreadai . 🙌 And of course, thanks to the whole BrowseComp-Plus team for maintaining such an interesting benchmark! 🙏@xueguang_ma @zijian42chen @luyu_gao @haora @ShengyaoZhuang @crystina_z @beirmug @lintool @WenhuChen 📜: arxiv.org/pdf/2609.14412 💻 : github.com/radinhamidi/Quest… From amazing team: @radin_hrad @amin_bigdelii @sadjadeb @claclarke @ebrahim_bagheri @bfung123
1
7
184
Negar Arabzadeh retweeted
jev for semantic operators is a great example of how LLM data processing at scale requires rethinking across the stack, model architecture and learning algs included new work and many more thoughts on this coming soon 👀 in the meantime, check out lotus, star and stay tuned (link below)
Replying to @EricMao06
I borrowed a ton of concepts from the lotus paper. The idea was well ahead of its time. Credit to the core contributors: @lianapatel_ @sid_jha1 @pgasawa @melissapan Harshit Gupta
9
19
1
104
10,701
Negar Arabzadeh retweeted
Two papers you should read from the information retrieval community about scaling up agent swarms:
2
8
1
92
10,220
Negar Arabzadeh retweeted
We pretrained a 2.3B MoE (360M active) Hybrid Mamba-2 that lands within a few points of Llama-3.2-3B using <1% of its pretraining FLOPs. No dedicated cluster. The run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase. Meet Rigel 🧵
22
80
18
631
82,368
Negar Arabzadeh retweeted
Cool blogpost from Melissa and team on how much the harness matters for coding agents. On the two benchmarks tested, switching harnesses often changed costs much more than success rates, and the default harness of model providers do not necessarily give results that are pareto optimal!
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
2
2
13
1,048
So we found out that your Claude model might not neccesarily need Claude Code 👀
Replying to @melissapan
Finding 1: Harness affects cost more than correctness. Fable 5 can cost twice as much for a 1.1-point gain in success rate. On SWE-bench Lite: - Claude Code: 97.8% accuracy, $1.33/rollout - Pi: 96.7% accuracy, $0.67/rollout While Claude Code reaches the highest success rate on the SWE-bench Lite frontier, Pi and Codex often achieve similar success rates at lower cost across the models we test. So you may be paying a hidden “harness tax” if you pick your harness based only on the success rate… 💸 (2/n)
7
621
Spoiler alert: your harness might be a burden on your wallet 👀💸
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
2
1
5
1,446
Negar Arabzadeh retweeted
If you're building production agents and want to contribute to research, take our survey! We'll release aggregate stats like we did in last year's MAP paper.
1/🤖Agents in production are changing fast. Following our #ICML2026 oral work, we’re running the study again this year to understand what’s changing in how agents are built, deployed and evaluated in real systems. 📊 If you’re working on agents in production, we’d love your input. The survey is short, but it can help us build a much clearer picture of what is actually working in practice. 🔗 Survey: berkeley.qualtrics.com/jfe/f… 🎙️ We’re also interviewing practitioners! Reach out if you’d like to share your experience in more depth. 👇 And if you’re curious about what we learned last year, check out our previous study in the thread below
11
45
5,784
Negar Arabzadeh retweeted
I'm back on Twitter/X, because of Measuring Agents in Production and all the amazing projects in our @UCBerkeley @BerkeleySky Lab! Check it out: berkeley.qualtrics.com/jfe/f… #AI #Research #Agents
The agent ecosystem is evolving rapidly and we are trying to capture a picture of where we are and where things are headed. We're now collecting responses for the 2026 survey of agents. berkeley.qualtrics.com/jfe/f… Help us create a more complete picture of how agent systems are evolving, what challenges remain, and where future research may be most impactful.
3
5
492
Negar Arabzadeh retweeted
The agent ecosystem is evolving rapidly and we are trying to capture a picture of where we are and where things are headed. We're now collecting responses for the 2026 survey of agents. berkeley.qualtrics.com/jfe/f… Help us create a more complete picture of how agent systems are evolving, what challenges remain, and where future research may be most impactful.
1
10
2
15
3,551
1/🤖Agents in production are changing fast. Following our #ICML2026 oral work, we’re running the study again this year to understand what’s changing in how agents are built, deployed and evaluated in real systems. 📊 If you’re working on agents in production, we’d love your input. The survey is short, but it can help us build a much clearer picture of what is actually working in practice. 🔗 Survey: berkeley.qualtrics.com/jfe/f… 🎙️ We’re also interviewing practitioners! Reach out if you’d like to share your experience in more depth. 👇 And if you’re curious about what we learned last year, check out our previous study in the thread below
1
4
1
17
6,517
2/ Our last study looked at how agents are actually being used in production, including common design patterns, deployment choices, and the challenges teams are facing. 🎥 ICML talk: icml.cc/virtual/2026/oral/71… 📄 Paper: arxiv.org/abs/2512.04123 Would really appreciate sharing the survey with others building agents in production.
4
252
Tracking 2 years of frontier models on data-agent benchmarks showed us that the general coding agents now beat carefully hand-designed data agents by up to 37 points and with 4× fewer turns. 😵 👇 Checkout our new work arxiv.org/abs/2609.03141
As better models rapidly eat the stack, what system research problems will remain to enable the next frontier of data agents? We studied the evolution of data agent capabilities the past 2 years, finding general coding agents beat hand-designed data agents by up to 37 points, with 4× fewer turns. So what's actually left for system researchers to solve? More than you might think... 📚 Paper: arxiv.org/abs/2609.03141 🧵👇
1
3
18
1,817
1/ 🧵 BrowseComp-Plus (ACL 2026) set the standard for disentangled agentic search eval... 🙌 𝐁𝐫𝐨𝐰𝐬𝐞𝐂𝐨𝐦𝐩-𝐏𝐥𝐮𝐬_𝐂𝐌 makes it even better! 🚀 Same BCP questions grounded on 𝐂𝐥𝐢𝐦𝐛𝐌𝐢𝐱-𝟒𝟎𝟎𝐁 with 𝟓𝟓𝟑𝐌 𝐝𝐨𝐜𝐬. 💻 github.com/castorini/cmass
🚀 Introducing BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent. It is a new Deep-Research evaluation benchmark built on top of BrowseComp. It features - 📚 a fixed, carefully curated corpus of web documents - ✅ human-verified positive documents - ⚔️ web-mined challenging hard negatives. With BrowseComp-Plus, you can thoroughly evaluate and compare the performance of different components in a deep-research system. e.g. GPT-5 + Qwen3-Embedding. Code, dataset, and leaderboard links are provided at the end of this thread.
1
9
34
4,470