AI Group @ucsantabarbara. Profs. @xwang_lk, @YuhengBu, Xifeng Yan, @WilliamWangNLP, @CodeTerminator. Account run by Student Social Committee.

Santa Barbara, CA
Joined July 2021
AI agents aren't biological individuals. So why make them evolve like one? They can directly share experience and learned artifacts. Not bounded by reproduction, lineage, or genes. Meet ๐—š๐—˜๐—”, accepted to ๐—–๐—ข๐—Ÿ๐—  ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฒ ๐ŸŽ‰ ๐Ÿณ๐Ÿญ.๐Ÿฌ% SWE-bench Verified / ๐Ÿด๐Ÿด.๐Ÿฏ% Polyglot ยท zero human intervention ๐Ÿงต๐Ÿ‘‡ Most self-evolving agent systems follow a similar pattern: select a parent, refine it, produce an offspring, repeat. Evolution unfolds as a tree. It's great at generating diversity, but that diversity gets trapped. Agents explore independently, and instead of serving as stepping stones, their discoveries stay stuck in local branches. Most variants are short-lived. ๐—˜๐˜…๐—ฝ๐—น๐—ผ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ต๐—ฎ๐—ฝ๐—ฝ๐—ฒ๐—ป๐˜€; ๐—ฟ๐—ฒ๐˜‚๐˜€๐—ฒ ๐—ฎ๐—ป๐—ฑ ๐—ฎ๐—ฐ๐—ฐ๐˜‚๐—บ๐˜‚๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฟ๐—ฎ๐—ฟ๐—ฒ๐—น๐˜† ๐—ฑ๐—ผ. It's time to rethink the evolution of AI agents: why keep evolving them like biological individuals? AI agents aren't bound by reproduction, lineage, or genes. They can directly share trajectories, tools, workflows, and learned artifacts, aggregating complementary skills instantly. ๐—ช๐—ต๐˜† ๐—ป๐—ผ๐˜ ๐—ฟ๐—ฒ๐—ฑ๐—ฒ๐˜€๐—ถ๐—ด๐—ป ๐—ฒ๐˜ƒ๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป ๐—ฎ๐—ฟ๐—ผ๐˜‚๐—ป๐—ฑ ๐˜„๐—ต๐—ฎ๐˜ ๐˜๐—ต๐—ฒ๐˜† ๐—ฐ๐—ฎ๐—ป ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐—น๐—น๐˜† ๐—ฑ๐—ผ? That's GEA. GEA makes a ๐—ด๐—ฟ๐—ผ๐˜‚๐—ฝ ๐—ผ๐—ณ ๐—ฎ๐—ด๐—ฒ๐—ป๐˜๐˜€ ๐˜๐—ต๐—ฒ ๐—ณ๐˜‚๐—ป๐—ฑ๐—ฎ๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐—น ๐˜‚๐—ป๐—ถ๐˜ ๐—ผ๐—ณ ๐—ฒ๐˜ƒ๐—ผ๐—น๐˜‚๐˜๐—ถ๐—ผ๐—ป. Each round, a parent group is selected under a Performance-Novelty criterion that balances competence with exploratory diversity. All members within the group then pool their experience, model patches, failure modes, eval logs, and solutions, into a shared pool, and the whole group jointly produces the next generation. Exploration is no longer wasted. It gets consolidated. The results show a significant improvement over prior state-of-the-art self-evolving methods, and GEA matches or surpasses top human-designed frameworks with ๐—ป๐—ผ ๐—ต๐˜‚๐—บ๐—ฎ๐—ป ๐—ถ๐—ป ๐˜๐—ต๐—ฒ ๐—น๐—ผ๐—ผ๐—ฝ. The analysis is the more interesting part. GEA's gains come from explicitly reusing diversity, not from lucky outliers: stronger performance under the same number of evolved agents, more robust to framework-level bugs (repaired in ๐Ÿญ.๐Ÿฐ iterations vs ๐Ÿฑ), and improvements that target workflows and tools rather than overfitting to one model, so they transfer consistently across GPT- and Claude-series backbones. The key to open-ended evolution is not only generating enough diversity. What matters more is whether discoveries accumulate and get reused. ๐—š๐—˜๐—” ๐—น๐—ฒ๐˜๐˜€ ๐—ฒ๐˜…๐—ฝ๐—น๐—ผ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐˜€๐˜๐—ฎ๐—ฟ๐˜ ๐—ฐ๐—ผ๐—บ๐—ฝ๐—ผ๐˜‚๐—ป๐—ฑ๐—ถ๐—ป๐—ด. ๐Ÿ“„ arXiv: arxiv.org/abs/2602.04837 Grateful to my advisor @xwang_lk and my wonderful coauthors @anton_iades @deepaknathani11 @zhenzhangzz @XiaoSophiaPu . Interning in Palo Alto this summer. See you at COLM 2026 ๐Ÿ‘‹
5
52
7
254
21,659
UCSB being creative at @icmlconf
1
26
2,300
How well can AI research agents explore new scientific ideas? The latest paper from our lab, Heuresis, combines coding agents with search and quality-diversity algorithms. Algorithms that provide better exploration are the key to unlock new ideas. Check out this thread๐Ÿ‘‡
Key to realizing Auto Research Agents that can make novel discoveries in AI is understanding and improving their exploration capabilities. To this end, we built Heuresis, a composable framework that combines coding agents with arbitrary search algorithms within a flexible loop.
1
5
934
Key to realizing Auto Research Agents that can make novel discoveries in AI is understanding and improving their exploration capabilities. To this end, we built Heuresis, a composable framework that combines coding agents with arbitrary search algorithms within a flexible loop.
2
13
5
38
7,005
UCSB AI retweeted
I will be attending @icmlconf in Seoul ๐Ÿ‡ฐ๐Ÿ‡ท next week presenting SAW-Bench!! Would love to chat about embodied perception, spatial reasoning, world modeling, and general video understanding ๐Ÿ˜ƒ DMs are open ๐Ÿ™Œ โŒš๏ธTue, Jun 7, 10:30 am -- 12:15 pm ๐Ÿ“Hall A #4309
Human perception is inherently situated โ€“ we understand the world relative to our own body, viewpoint, and motion. To deploy multimodal foundation models in embodied settings, we ask: โ€œCan these models reason in the same observer-centric way?โ€ We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness: - 786 real world egocentric videos - 2,071 human-annotated QA pairs Across all tasks, we evaluate 24 state-of-the-art MFMs: ๐Ÿ“‰ Best model: 53.9% ๐Ÿง‘ Humans: 91.6% Models systematically: โŒ Confuse head rotation with physical movement โŒ Collapse under multi-turn trajectories โŒ Fail to maintain persistent world-state memory ๐Ÿ‘‰ We see that maintaining a stable observer-centric representation remains challenging. As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction. We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
1
1
21
5,751
Auto-research is a search problem. ๐’๐ž๐š๐ซ๐œ๐ก ๐ข๐ฌ ๐š ๐Ÿ๐จ๐ซ๐ฆ ๐จ๐Ÿ ๐ก๐จ๐ฅ๐ข๐ฌ๐ญ๐ข๐œ ๐ซ๐ž๐š๐ฌ๐จ๐ง๐ข๐ง๐  ๐จ๐ฏ๐ž๐ซ ๐š๐ง ๐ž๐ฏ๐จ๐ฅ๐ฏ๐ข๐ง๐  ๐ฅ๐š๐ง๐๐ฌ๐œ๐š๐ฉ๐ž ๐จ๐Ÿ ๐ฉ๐จ๐ฌ๐ฌ๐ข๐›๐ข๐ฅ๐ข๐ญ๐ข๐ž๐ฌ. A scientist does not simply solve a task. They generate hypotheses, design experiments, interpret failures, update beliefs, and decide where to explore next. The central challenge is navigating the search space efficiently. This perspective is largely missing from todayโ€™s autonomous research systems, where reasoning is often confined to a single trajectory. In ARTS, we introduce a reasoning-guided tree search framework with test-time learning, enabling agents to reason about the search process itself. We show that a fine-tuned 4B model can achieve performance comparable to frontier closed-model research agents on MLGym and MLEBench, while operating at substantially lower inference cost. More broadly, we believe that progress in autonomous research will increasingly come from better reasoning-driven search, not just larger models.
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts. Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context
9
32
5
325
47,656
UCSB AI retweeted
While AI scientists are getting better, two core challenges remain: (i) deciding what to explore next and (ii) handling the growing context. ARTS addresses both with reasoning-guided exploration and test-time RL to learn from long search histories.
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts. Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context
8
35
5,163
A good researcher has intuition about when to drop a line of work and when to push. Current auto-discovery systems delegate this decision to scalar scores, which cannot represent this judgment. ARTS (Agentic Reasoning for Tree Search) reasons over the search state to decide.
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts. Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context
1
6
47
6,646
Excited to be at @icmlconf in Seoul๐Ÿ‡ฐ๐Ÿ‡ท๐ŸŽ‰! Presenting my work on Reasoning, RL, Co-Evolution & automated discovery. Would love to chat about RL, Agents, Reasoning, Automated Discovery, Co-Evolution & anything else. Come say hi by my posters July 8 and 9 in Hall A! ๐Ÿ™Œ
3
29
1,341
Interesting observation: even after TTT, diversity doesn't collapse. We test-time-train the ARTS "scientist" on a simple, performance-only reward. Intuition says this should collapse solution diversity. But it doesn't, this is attributed to verbalized sampling which still surfaces the low-probability solution.
Replying to @GurushaJuneja
When the search history outgrows the context window, ARTS* test time trains the scientist on its own search using reinforcement learning, applying GRPO with a percentile reward to instill the search into the weights. A 4B scientist then matches o3 at much lower inference cost.
1
3
11
1,386
Several of our students will be presenting at @icmlconf in Seoul this week! ๐Ÿ‡ฐ๐Ÿ‡ท Work spanning RL, agents, automated discovery, spatial intelligence & robotics. Come say hi by the posters ๐Ÿ‘‹
16
736
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts. Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context
6
28
4
274
69,808
Had so much fun working on this! ABC provides robot data and base models for all!
Introducing ABC: open data, training, and infrastructure for robotics. We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques. @arthurallshire @Cinnabar233 @adamrasb @redstone_hong @davidrmcall
2
3
1
25
3,144
UCSB AI retweeted
Excited to share that our paper SAW-Bench: Learning Situated Awareness in the Real World received the Best Paper Award Runner-Up at the #CVPR2026 WMAS workshop! Congratulations to all co-authors, and thanks to the organizers and reviewers for the recognition. Special thanks to @jieneng_chen for hosting me at WMAS!
Human perception is inherently situated โ€“ we understand the world relative to our own body, viewpoint, and motion. To deploy multimodal foundation models in embodied settings, we ask: โ€œCan these models reason in the same observer-centric way?โ€ We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness: - 786 real world egocentric videos - 2,071 human-annotated QA pairs Across all tasks, we evaluate 24 state-of-the-art MFMs: ๐Ÿ“‰ Best model: 53.9% ๐Ÿง‘ Humans: 91.6% Models systematically: โŒ Confuse head rotation with physical movement โŒ Collapse under multi-turn trajectories โŒ Fail to maintain persistent world-state memory ๐Ÿ‘‰ We see that maintaining a stable observer-centric representation remains challenging. As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction. We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
1
3
17
3,978
Congratulations to @_Chuhan_Li , @xwang_lk, and their collaborators for their SAW-Bench paper on receiving the Best Paper Award Runner-Up from the CVPR 2026 WMAS workshop! ๐Ÿ†
Human perception is inherently situated โ€“ we understand the world relative to our own body, viewpoint, and motion. To deploy multimodal foundation models in embodied settings, we ask: โ€œCan these models reason in the same observer-centric way?โ€ We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness: - 786 real world egocentric videos - 2,071 human-annotated QA pairs Across all tasks, we evaluate 24 state-of-the-art MFMs: ๐Ÿ“‰ Best model: 53.9% ๐Ÿง‘ Humans: 91.6% Models systematically: โŒ Confuse head rotation with physical movement โŒ Collapse under multi-turn trajectories โŒ Fail to maintain persistent world-state memory ๐Ÿ‘‰ We see that maintaining a stable observer-centric representation remains challenging. As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction. We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
3
4
2,151
UCSB AI @ CVPR 2026!
1
2
31
2,686
๐Ÿ“œSAW-BENC evaluates situated awareness in real-world egocentric videos: can a model track where it is, where it came from, and what actions are possible from its current viewpoint? Results reveal a large gap between current multimodal models and humans. nitter.cf/_Chuhan_Li/status/2024โ€ฆ
Human perception is inherently situated โ€“ we understand the world relative to our own body, viewpoint, and motion. To deploy multimodal foundation models in embodied settings, we ask: โ€œCan these models reason in the same observer-centric way?โ€ We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness: - 786 real world egocentric videos - 2,071 human-annotated QA pairs Across all tasks, we evaluate 24 state-of-the-art MFMs: ๐Ÿ“‰ Best model: 53.9% ๐Ÿง‘ Humans: 91.6% Models systematically: โŒ Confuse head rotation with physical movement โŒ Collapse under multi-turn trajectories โŒ Fail to maintain persistent world-state memory ๐Ÿ‘‰ We see that maintaining a stable observer-centric representation remains challenging. As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction. We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
1
1
401
๐Ÿ“œReasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space: A training-free multimodal reasoning framework that interleaves latent reasoning and visual evidence, improving both reasoning accuracy and visual grounding nitter.cf/liuchen02938149/statusโ€ฆ
๐Ÿง  Can Multimodal Models Think Like Humans? Most multimodal models reason in rigid pipelines: either see once, then overthink in text or constantly use external visual tools to re-check. Can reasoning and perception be dynamically interleaved like in the human mind? ๐Ÿ‘ค Human: ๐Ÿ‘€ Look โ†’ ๐Ÿง  Think in mind โ†’ ๐Ÿ” Re-look when confidence is low ๐Ÿค– DMLR (ours): ๐Ÿ‘€ Perception โ†’ ๐Ÿง ย Think in latent space โ†’ ๐Ÿ‘€ Selective re-percept to maximize token confidence 1๏ธโƒฃ ๐Ÿซฃย Seeing at Every Step Is Unnecessary. Only a small subset of reasoning steps require visual input. 2๏ธโƒฃย ๐Ÿงญย Confidence as the Compass. Confidence captures the modelโ€™s intrinsic state, reflecting accuracy, reasoning quality, and visual grounding. 3๏ธโƒฃ ๐Ÿง ย Drafting in the "Mind". DMLR directly optimizes think token in latent space, enabling deeper reasoning without additional generation cost. 4๏ธโƒฃ ๐Ÿ’‰ย Dynamic Visual Injection Strategy. DMLR selects and injects only the most relevant visual patches, dynamically updated across iterations. ๐Ÿš€ Read on to explore more analysis and insights! ๐ŸŽ“
1
265
Going to Denver ๐Ÿ”๏ธ for CVPR? UCSB is having a great presence! Make sure to checkout these works from our lab! ๐Ÿงต๐Ÿ‘‡
1
2
525
๐Ÿ“œSelf-Evolving 3D Scene Generation from a Single Image: A self-evolving framework for single-image 3D scene generation that alternates between reconstruction and novel-view synthesis, progressively improving geometry, coverage, and texture quality nitter.cf/KaizhiZheng/status/199โ€ฆ
๐Ÿš€ Introducing EvoScene: Self-Evolving 3D Scene Generation from a Single Image! Generating complete, textured 3D scenes from a single photo is challenging due to limited coverage and inconsistent textures. EvoScene solves this with a novel, training-free, self-evolving framework that progressively reconstructs high-quality, ready-to-use 3D meshes. ๐Ÿง  What's new: We establish a virtuous cycle where geometry and appearance mutually refine each other by synergistically combining geometric reasoning from 3D diffusion models and visual knowledge from video generation models. This process expands spatial coverage and completes unseen regions. ๐Ÿ† SOTA Results Confirmed: EvoScene achieves superior geometric stability, layout coherence, and photorealistic appearance compared to strong baselines. Human Preference Win Rate: 78.5%โ€“90.5% across all quality criteria. Semantic Fidelity (CLIP): 0.8643, a 15.9% improvement over Trellis. Read the full paper and see the visualizations below! ๐Ÿงต Project page: eric-ai-lab.github.io/evosceโ€ฆ Paper: arxiv.org/abs/2512.08905 Code: github.com/eric-ai-lab/EvoScโ€ฆ #3DGeneration #ComputerVision #SceneGeneration
1
652