AI Group @ucsantabarbara. Profs. @xwang_lk, @YuhengBu, Xifeng Yan, @WilliamWangNLP, @CodeTerminator. Account run by Student Social Committee.
Santa Barbara, CA
Joined July 2021
- Tweets363
- Following637
- Followers2.5K
- Likes640
UCSB AI retweeted
AI agents aren't biological individuals. So why make them evolve like one?
They can directly share experience and learned artifacts. Not bounded by reproduction, lineage, or genes.
Meet ๐๐๐, accepted to ๐๐ข๐๐ ๐ฎ๐ฌ๐ฎ๐ฒ ๐
๐ณ๐ญ.๐ฌ% SWE-bench Verified / ๐ด๐ด.๐ฏ% Polyglot ยท zero human intervention
๐งต๐
Most self-evolving agent systems follow a similar pattern: select a parent, refine it, produce an offspring, repeat. Evolution unfolds as a tree.
It's great at generating diversity, but that diversity gets trapped. Agents explore independently, and instead of serving as stepping stones, their discoveries stay stuck in local branches. Most variants are short-lived.
๐๐
๐ฝ๐น๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป ๐ต๐ฎ๐ฝ๐ฝ๐ฒ๐ป๐; ๐ฟ๐ฒ๐๐๐ฒ ๐ฎ๐ป๐ฑ ๐ฎ๐ฐ๐ฐ๐๐บ๐๐น๐ฎ๐๐ถ๐ผ๐ป ๐ฟ๐ฎ๐ฟ๐ฒ๐น๐ ๐ฑ๐ผ.
It's time to rethink the evolution of AI agents: why keep evolving them like biological individuals?
AI agents aren't bound by reproduction, lineage, or genes. They can directly share trajectories, tools, workflows, and learned artifacts, aggregating complementary skills instantly. ๐ช๐ต๐ ๐ป๐ผ๐ ๐ฟ๐ฒ๐ฑ๐ฒ๐๐ถ๐ด๐ป ๐ฒ๐๐ผ๐น๐๐๐ถ๐ผ๐ป ๐ฎ๐ฟ๐ผ๐๐ป๐ฑ ๐๐ต๐ฎ๐ ๐๐ต๐ฒ๐ ๐ฐ๐ฎ๐ป ๐ฎ๐ฐ๐๐๐ฎ๐น๐น๐ ๐ฑ๐ผ?
That's GEA.
GEA makes a ๐ด๐ฟ๐ผ๐๐ฝ ๐ผ๐ณ ๐ฎ๐ด๐ฒ๐ป๐๐ ๐๐ต๐ฒ ๐ณ๐๐ป๐ฑ๐ฎ๐บ๐ฒ๐ป๐๐ฎ๐น ๐๐ป๐ถ๐ ๐ผ๐ณ ๐ฒ๐๐ผ๐น๐๐๐ถ๐ผ๐ป. Each round, a parent group is selected under a Performance-Novelty criterion that balances competence with exploratory diversity. All members within the group then pool their experience, model patches, failure modes, eval logs, and solutions, into a shared pool, and the whole group jointly produces the next generation.
Exploration is no longer wasted. It gets consolidated.
The results show a significant improvement over prior state-of-the-art self-evolving methods, and GEA matches or surpasses top human-designed frameworks with ๐ป๐ผ ๐ต๐๐บ๐ฎ๐ป ๐ถ๐ป ๐๐ต๐ฒ ๐น๐ผ๐ผ๐ฝ.
The analysis is the more interesting part. GEA's gains come from explicitly reusing diversity, not from lucky outliers: stronger performance under the same number of evolved agents, more robust to framework-level bugs (repaired in ๐ญ.๐ฐ iterations vs ๐ฑ), and improvements that target workflows and tools rather than overfitting to one model, so they transfer consistently across GPT- and Claude-series backbones.
The key to open-ended evolution is not only generating enough diversity. What matters more is whether discoveries accumulate and get reused. ๐๐๐ ๐น๐ฒ๐๐ ๐ฒ๐
๐ฝ๐น๐ผ๐ฟ๐ฎ๐๐ถ๐ผ๐ป ๐๐๐ฎ๐ฟ๐ ๐ฐ๐ผ๐บ๐ฝ๐ผ๐๐ป๐ฑ๐ถ๐ป๐ด.
๐ arXiv: arxiv.org/abs/2602.04837
Grateful to my advisor @xwang_lk and my wonderful coauthors @anton_iades @deepaknathani11 @zhenzhangzz @XiaoSophiaPu .
Interning in Palo Alto this summer. See you at COLM 2026 ๐
How well can AI research agents explore new scientific ideas?
The latest paper from our lab, Heuresis, combines coding agents with search and quality-diversity algorithms. Algorithms that provide better exploration are the key to unlock new ideas.
Check out this thread๐
UCSB AI retweeted
Key to realizing Auto Research Agents that can make novel discoveries in AI is understanding and improving their exploration capabilities.
To this end, we built Heuresis, a composable framework that combines coding agents with arbitrary search algorithms within a flexible loop.
UCSB AI retweeted
I will be attending @icmlconf in Seoul ๐ฐ๐ท next week presenting SAW-Bench!!
Would love to chat about embodied perception, spatial reasoning, world modeling, and general video understanding ๐
DMs are open ๐
โ๏ธTue, Jun 7, 10:30 am -- 12:15 pm
๐Hall A #4309
Human perception is inherently situated โ we understand the world relative to our own body, viewpoint, and motion.
To deploy multimodal foundation models in embodied settings, we ask:
โCan these models reason in the same observer-centric way?โ
We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness:
- 786 real world egocentric videos
- 2,071 human-annotated QA pairs
Across all tasks, we evaluate 24 state-of-the-art MFMs:
๐ Best model: 53.9%
๐ง Humans: 91.6%
Models systematically:
โ Confuse head rotation with physical movement
โ Collapse under multi-turn trajectories
โ Fail to maintain persistent world-state memory
๐ We see that maintaining a stable observer-centric representation remains challenging.
As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction.
We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
UCSB AI retweeted
Auto-research is a search problem.
๐๐๐๐ซ๐๐ก ๐ข๐ฌ ๐ ๐๐จ๐ซ๐ฆ ๐จ๐ ๐ก๐จ๐ฅ๐ข๐ฌ๐ญ๐ข๐ ๐ซ๐๐๐ฌ๐จ๐ง๐ข๐ง๐ ๐จ๐ฏ๐๐ซ ๐๐ง ๐๐ฏ๐จ๐ฅ๐ฏ๐ข๐ง๐ ๐ฅ๐๐ง๐๐ฌ๐๐๐ฉ๐ ๐จ๐ ๐ฉ๐จ๐ฌ๐ฌ๐ข๐๐ข๐ฅ๐ข๐ญ๐ข๐๐ฌ.
A scientist does not simply solve a task. They generate hypotheses, design experiments, interpret failures, update beliefs, and decide where to explore next. The central challenge is navigating the search space efficiently.
This perspective is largely missing from todayโs autonomous research systems, where reasoning is often confined to a single trajectory.
In ARTS, we introduce a reasoning-guided tree search framework with test-time learning, enabling agents to reason about the search process itself. We show that a fine-tuned 4B model can achieve performance comparable to frontier closed-model research agents on MLGym and MLEBench, while operating at substantially lower inference cost.
More broadly, we believe that progress in autonomous research will increasingly come from better reasoning-driven search, not just larger models.
UCSB AI retweeted
While AI scientists are getting better, two core challenges remain: (i) deciding what to explore next and (ii) handling the growing context.
ARTS addresses both with reasoning-guided exploration and test-time RL to learn from long search histories.
UCSB AI retweeted
A good researcher has intuition about when to drop a line of work and when to push. Current auto-discovery systems delegate this decision to scalar scores, which cannot represent this judgment. ARTS (Agentic Reasoning for Tree Search) reasons over the search state to decide.
UCSB AI retweeted
Excited to be at @icmlconf in Seoul๐ฐ๐ท๐!
Presenting my work on Reasoning, RL, Co-Evolution & automated discovery.
Would love to chat about RL, Agents, Reasoning, Automated Discovery, Co-Evolution & anything else.
Come say hi by my posters July 8 and 9 in Hall A! ๐
UCSB AI retweeted
Interesting observation: even after TTT, diversity doesn't collapse. We test-time-train the ARTS "scientist" on a simple, performance-only reward. Intuition says this should collapse solution diversity. But it doesn't, this is attributed to verbalized sampling which still surfaces the low-probability solution.
Replying to @GurushaJuneja
When the search history outgrows the context window, ARTS* test time trains the scientist on its own search using reinforcement learning, applying GRPO with a percentile reward to instill the search into the weights. A 4B scientist then matches o3 at much lower inference cost.
UCSB AI retweeted
Automated discovery is a search problem over combinatorially large spaces. It requires reasoning to explore promising ideas while handling long contexts.
Introducing ARTS, a search algorithm that (i) reasons to guide exploration and (ii) test-time trains to handle long context
UCSB AI retweeted
Had so much fun working on this! ABC provides robot data and base models for all!
Introducing ABC: open data, training, and infrastructure for robotics.
We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques.
@arthurallshire @Cinnabar233 @adamrasb @redstone_hong @davidrmcall
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
UCSB AI retweeted
Excited to share that our paper SAW-Bench: Learning Situated Awareness in the Real World received the Best Paper Award Runner-Up at the #CVPR2026 WMAS workshop!
Congratulations to all co-authors, and thanks to the organizers and reviewers for the recognition.
Special thanks to @jieneng_chen for hosting me at WMAS!
Human perception is inherently situated โ we understand the world relative to our own body, viewpoint, and motion.
To deploy multimodal foundation models in embodied settings, we ask:
โCan these models reason in the same observer-centric way?โ
We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness:
- 786 real world egocentric videos
- 2,071 human-annotated QA pairs
Across all tasks, we evaluate 24 state-of-the-art MFMs:
๐ Best model: 53.9%
๐ง Humans: 91.6%
Models systematically:
โ Confuse head rotation with physical movement
โ Collapse under multi-turn trajectories
โ Fail to maintain persistent world-state memory
๐ We see that maintaining a stable observer-centric representation remains challenging.
As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction.
We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
Congratulations to @_Chuhan_Li , @xwang_lk, and their collaborators for their SAW-Bench paper on receiving the Best Paper Award Runner-Up from the CVPR 2026 WMAS workshop! ๐
Human perception is inherently situated โ we understand the world relative to our own body, viewpoint, and motion.
To deploy multimodal foundation models in embodied settings, we ask:
โCan these models reason in the same observer-centric way?โ
We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness:
- 786 real world egocentric videos
- 2,071 human-annotated QA pairs
Across all tasks, we evaluate 24 state-of-the-art MFMs:
๐ Best model: 53.9%
๐ง Humans: 91.6%
Models systematically:
โ Confuse head rotation with physical movement
โ Collapse under multi-turn trajectories
โ Fail to maintain persistent world-state memory
๐ We see that maintaining a stable observer-centric representation remains challenging.
As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction.
We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
๐SAW-BENC evaluates situated awareness in real-world egocentric videos: can a model track where it is, where it came from, and what actions are possible from its current viewpoint? Results reveal a large gap between current multimodal models and humans.
nitter.cf/_Chuhan_Li/status/2024โฆ
Human perception is inherently situated โ we understand the world relative to our own body, viewpoint, and motion.
To deploy multimodal foundation models in embodied settings, we ask:
โCan these models reason in the same observer-centric way?โ
We study this through SAW-Bench: a novel benchmark for observer-centric situated awareness:
- 786 real world egocentric videos
- 2,071 human-annotated QA pairs
Across all tasks, we evaluate 24 state-of-the-art MFMs:
๐ Best model: 53.9%
๐ง Humans: 91.6%
Models systematically:
โ Confuse head rotation with physical movement
โ Collapse under multi-turn trajectories
โ Fail to maintain persistent world-state memory
๐ We see that maintaining a stable observer-centric representation remains challenging.
As MFMs are increasingly integrated into embodied agents, situated awareness becomes essential for reliable real-world interaction.
We release SAW-Bench and encourage further research toward improving observer-centric reasoning in multimodal foundation models.
๐Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space: A training-free multimodal reasoning framework that interleaves latent reasoning and visual evidence, improving both reasoning accuracy and visual grounding
nitter.cf/liuchen02938149/statusโฆ
๐ง Can Multimodal Models Think Like Humans?
Most multimodal models reason in rigid pipelines: either see once, then overthink in text or constantly use external visual tools to re-check. Can reasoning and perception be dynamically interleaved like in the human mind?
๐ค Human: ๐ Look โ ๐ง Think in mind โ ๐ Re-look when confidence is low
๐ค DMLR (ours): ๐ Perception โ ๐ง ย Think in latent space โ ๐ Selective re-percept to maximize token confidence
1๏ธโฃ ๐ซฃย Seeing at Every Step Is Unnecessary. Only a small subset of reasoning steps require visual input.
2๏ธโฃย ๐งญย Confidence as the Compass. Confidence captures the modelโs intrinsic state, reflecting accuracy, reasoning quality, and visual grounding.
3๏ธโฃ ๐ง ย Drafting in the "Mind". DMLR directly optimizes think token in latent space, enabling deeper reasoning without additional generation cost.
4๏ธโฃ ๐ย Dynamic Visual Injection Strategy. DMLR selects and injects only the most relevant visual patches, dynamically updated across iterations.
๐ Read on to explore more analysis and insights! ๐
๐Self-Evolving 3D Scene Generation from a Single Image: A self-evolving framework for single-image 3D scene generation that alternates between reconstruction and novel-view synthesis, progressively improving geometry, coverage, and texture quality
nitter.cf/KaizhiZheng/status/199โฆ
๐ Introducing EvoScene: Self-Evolving 3D Scene Generation from a Single Image!
Generating complete, textured 3D scenes from a single photo is challenging due to limited coverage and inconsistent textures. EvoScene solves this with a novel, training-free, self-evolving framework that progressively reconstructs high-quality, ready-to-use 3D meshes.
๐ง What's new: We establish a virtuous cycle where geometry and appearance mutually refine each other by synergistically combining geometric reasoning from 3D diffusion models and visual knowledge from video generation models. This process expands spatial coverage and completes unseen regions.
๐ SOTA Results Confirmed: EvoScene achieves superior geometric stability, layout coherence, and photorealistic appearance compared to strong baselines.
Human Preference Win Rate: 78.5%โ90.5% across all quality criteria.
Semantic Fidelity (CLIP): 0.8643, a 15.9% improvement over Trellis.
Read the full paper and see the visualizations below! ๐งต
Project page: eric-ai-lab.github.io/evosceโฆ
Paper: arxiv.org/abs/2512.08905
Code: github.com/eric-ai-lab/EvoScโฆ
#3DGeneration #ComputerVision #SceneGeneration