Tweeting interesting papers submitted at https://nitter.cf/t.co/rXX8x0HzXV. Submit your own at https://nitter.cf/t.co/QhbJKXBd4Q, and link models/datasets/demos to it!

Anywhere
Joined March 2025
Paint-Anything Unified any-color control for image generation and editing. Specify any object's target color with a 24-bit hex value, and the model paints exactly that shade. On FLUX.2-4B, +85.3% ACBench-T2I.
1
1
1
17
1,356
Grounded Action Models A new robot foundation model paradigm built on 3D grounding. Language, points, or boxes become a shared object-centric representation, mixed with robot state history to predict action chunks.
2
4
16
1,602
Ant Group’s Realtime-Venus is now on Hugging Face A full-duplex interaction system that listens while it speaks, sees audio-visual context, and delegates tasks to run in the background without pausing the conversation.
1
3
1
28
1,977
Alibaba's Wan team released WanPE A 397B-parameter prompt enhancement model that turns a short user prompt into a director-level, shot-by-shot cinematic screenplay, boosting video generator preference by up to 50.9 points at 30 seconds
5
15
1
179
11,698
Your LLM can hold two thoughts at once Averaging embeddings of two texts makes an LLM predict both continuations at once. This linear superposition is built into the Transformer, fades during pretraining, but can be restored with light fine-tuning. The authors show how to decode both streams from a single forward pass.
2
16
1
74
3,563
Agent-Editing World Model Rethinks world modeling for LLM agents by judging and editing noisy reasoning-action continuations before execution. AEWM hits 70.5% macro-F1 on Action Judge, +10.6 over the strongest baseline.
3
2
22
1,517
Training Object Permanence in World Models Can video models learn object permanence? WROP offers 150 Blender-rendered cognitive tasks, 1.5M samples, and a 300-question exam. A 16B world model, PWM-WROP, ranks first among continuation models.
3
3
16
1,686
Tencent researchers just released RewardVerse A rubric-guided video reward framework that inserts a dynamic rubric between the evaluation query and the scorer, mitigating scalar drift for stable, interpretable RL rewards
4
2
16
1,377