Opinions are my own. Ph.D. student at @RiceCompSci in @vislang. Previously @RealityLabs @AdobeResearch, @InariAILab, @AdaVivInc.

Houston, Texas
Joined December 2017
Jefferson Enrique Hernandez Cevallos retweeted
Today, most multimodal LLMs are basically a vision encoder stitched to a language model. The alternative is to drop the vision backbone entirely and project raw image patches straight into the decoder. But which one scales better: a rich visual prior or a simpler architecture? 1/n
6
23
1
190
12,669
Jefferson Enrique Hernandez Cevallos retweeted
Can complex multi-step reasoning emerge purely from cells that only talk to immediate neighbors? Happy to share our paper “Reasoning with Neural Cellular Automata (NCA)”, from our team at Google, Paradigms of Intelligence 🧵👇
44
178
52
1,239
125,774
Jefferson Enrique Hernandez Cevallos retweeted
New research (165 pages!) We thought there were too many benchmarks for video understanding, so we squeezed the benchmarks dry and got video-index (cover inspiration: gradient canopy 😜
9
26
2
233
15,191
Jefferson Enrique Hernandez Cevallos retweeted
We let an agent self-improve by exploring and internalizing ”on this sort of task, keep this sort of thing in mind“ bits of self-feedback. It learned to solve tasks where the original policy failed 128 attempts in a row and GRPO flatlined. Paper: arxiv.org/abs/2609.37633 🧵1/8
19
39
3
445
56,441
Jefferson Enrique Hernandez Cevallos retweeted
Most people think JEPAs are inherently non-contrastive. But recent models (LeJEPA, LeWM, etc.) use the SIGReg regularizer, which approximates a sliced MMD whose expansion reveals pairwise repulsion between samples in the batch.
5
19
219
10,994
Jefferson Enrique Hernandez Cevallos retweeted
With every new release, frontier models seem to do more in computer vision, including tasks we once thought only specialist models could do well. 🚀 𝗛𝗼𝘄 𝗺𝘂𝗰𝗵 𝗼𝗳 𝗰𝗼𝗺𝗽𝘂𝘁𝗲𝗿 𝘃𝗶𝘀𝗶𝗼𝗻 𝗰𝗮𝗻 𝗳𝗿𝗼𝗻𝘁𝗶𝗲𝗿 𝗺𝗼𝗱𝗲𝗹𝘀 𝗵𝗮𝗻𝗱𝗹𝗲 𝘁𝗼𝗱𝗮𝘆? 1/4
5
14
2
113
8,925
Jefferson Enrique Hernandez Cevallos retweeted
Paper thread! I recently read this paper from some FAIR colleagues and NYU folks in depth, so making a thread while it's fresh in my head. The main point is seeing how far on the SMALL side we can go in scaling laws without losing fit or predictability, and what it takes?
13
58
705
44,046
Jefferson Enrique Hernandez Cevallos retweeted
OK so let me recap: RL env makers put strings into the RL env that makes it clear it's an RL env. Like "this is not supported in this RL env". Then, lab safety/mechinterp folks be like OMG EvAL aWaReNeSs. Are you effing kidding me?? Just look at your data... surprised Pikachu.
Replying to @j0wimo
and then there are literal strings within the code LOL
31
54
7
968
80,258
Jefferson Enrique Hernandez Cevallos retweeted
Introducing Flow Reasoning Models. We developed a recurrent flow-based architecture to efficiently solve structured reasoning problems (e.g., Sudoku). FRMs apply continuous flows to discrete data and recurrently refine their past mistakes through self-conditioning.
51
330
18
2,928
128,851
Jefferson Enrique Hernandez Cevallos retweeted
Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean reward while simultaneously maximizing coverage over high reward outputs. 🧵 1/n
10
82
15
507
106,755
Jefferson Enrique Hernandez Cevallos retweeted
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
8
71
7
451
74,792
Jefferson Enrique Hernandez Cevallos retweeted
Introducing Speculative Programmatic Tool Calling (sPTC)! A general class of technique for speculating on tool calls during code generation in a harness and queuing them early to overlap with token generation + REPL execution time. Blog: alexzhang13.github.io/blog/2…
58
205
54
1,789
246,385
Jefferson Enrique Hernandez Cevallos retweeted
This cube visualization is my favorite way of remembering how batch norm, layer norm, instance norm, and group norm relate to each other! This figure is from the original “Group Normalization” paper.
1
3
123
5,924
Jefferson Enrique Hernandez Cevallos retweeted
arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (arxiv.org/abs/2608.08888). Why does this work without training?
32
105
32
1,057
236,091
Jefferson Enrique Hernandez Cevallos retweeted
Post-training with RL increases pass@1 accuracy and reduces solution coverage. When scaling test-time compute, maintaining solution coverage is essential! Can a method improve accuracy while maintaining coverage? Yes, Evolution Strategies can! We show this in our new paper!
2
19
3
141
19,591
Jefferson Enrique Hernandez Cevallos retweeted
I've made extensive changes to OpenCLIP over the last several months, especially wrt to supporting variable resolution/aspect NaFlexViT encoders and matching webdataset pipelines. I also integrated existing audio CLAP models into OpenCLIP, allowing existing @laion_ai weights to be used. While fiddling with CLAP I wondered what NaFlex + CLAP would look like so I tried it out. Using the timm NaFlexViT as a base, I put a mel patch embedder integrated into the data pipeline. Very simple and it works w/o noteworthy timm model arch changes. You can choose different patch geometries to cover the mel time vs freq axis and you naturally get variable time capability from the naflex sequence handling along the time axis. Added to the mix is a new 'modern-text' text encoder, much like ModernBERT it uses more recent ideas for the text transformer, and also supports variable length sequences with appropriate masking, in either a pooled or causal form. I tried training a few preliminary, modest size models with the largest audio-text dataset I had access to (WIP, unreleased) to validate the idea and architecture / training code changes. They're decent but definitely room for refinement and more extensive h-param and architecture search. They can be run from OpenCLIP main branch.
8
5
36
5,649
Jefferson Enrique Hernandez Cevallos retweeted
🌟We introduce a method for revisiting historical photos and bringing them to life. Specifically, our model generates video from motion-blurred images. 📢 Blur2Vid (Transactions on Graphics, SIGGRAPH Asia 2025) Webpage: blur2vid.github.io Paper: arxiv.org/abs/2512.19817
14
48
10
267
79,943
Jefferson Enrique Hernandez Cevallos retweeted
WaiT for the Signal: Simple Frequency-Aware Flow-Matching. We show how to natively incorporate a fundamental property of images directly into diffusion models; setting a new pixel-space SOTA on ImageNet, while reducing compute. 📄 arxiv.org/abs/2607.28760 Full breakdown below👇
3
26
3
99
17,984
Jefferson Enrique Hernandez Cevallos retweeted
Jax added support for opaque types with custom tangents. One cool use is proper handling of batched geometric objects. Here's a short demo of a differentiable rasterization, i.e. min L2 w/ 1500 triangles to a photo. Interesting to move beyond tensors. (docs.jax.dev/en/latest/hijax…)
3
11
2
178
21,199
Jefferson Enrique Hernandez Cevallos retweeted
notes on A Defense of the Quadratic Model link - arxiv.org/abs/2607.21716 ------------------------------------------------ it may be possible to use a simple model to predict how a model improves in late stages of training. perhaps the most interesting paper for LLMs in a while to me. how good is that claim? let us see. when we train a language model we are trying to change millions of knobs with each feedback to minimise our errors in prediction. but the easiest way to visualise this for this paper discussion is that you imagine a traveller who has to traverse a really unpredictable terrain all the way to the lowest point possible. now this paper asks that if we were to replace this uneven landscape with a simple bowl then how successful would we be in creating a copy of the traveller that can improve similar to the real uneven landscape? in the early and middle stages of training, this simplistic mapping fails quickly however near the end when the learning rate is slowly getting reduced it can match how the model would improve for several percentages of the training budget (5-10%). this does not work though when the learning rate is constant and works in the case of cosine decay and the model error prediction mapping fails even in the late stages of training in the constant learning rate case. this may be explainable with the help of mainly two changes. first, the learning-rate schedule makes the model take much smaller steps, so it remains inside the tiny neighbourhood where a simplistic local map should work because we expect no major changes in the terrain. second, learning may be saturating iin the late stages of training and the way it would present itself is where most common patterns are already learned. this means lesser and lesser drastic changes to the model landscape. the remaining weight changes could consequently concentrate into a low-rank, slowly rotating set of directions here. and the complete landscape can remain complicated while the narrow corridor in late stages of training actually used by the optimizer becomes simple and stable enough to resemble a bowl. the landscape itself not getting much drastic updates and the traveller slowing down allowing stable prediction for a budget of tokens where the landscape walk is actually much smaller are my intuitive guesses here but i think we can perhaps get creative with the predictable narrow training budgets part.
1
10
63
3,688