Joined February 2011
AlumnoRandom retweeted
We’re so close to peak
137
1,484
126
22,711
545,448
AlumnoRandom retweeted
water is transparent only within a very narrow band of the electromagnetic spectrum, so living organisms evolved sensitivity to that band, and that's what we now call "visible light".
202
2,272
222
24,118
854,050
AlumnoRandom retweeted
1/ We introduce TrackEverything: a 3D point tracker that tracks all points across all frames of long videos (1000+ frames). Our key idea is to tie computation to unique 3D scene content rather than redundant 2D pixels in a video. trackeverything.github.io
23
168
11
1,415
87,147
In the Utopia mouse experiment, they put mice in a “utopia” with unlimited food and no predators >Young male mice spent all their time grooming, uninterested in reproducing >Female mice lost maternal instinct and some harmed their own kids >birth rates went to zero Uhhhh
700
6,346
391
58,939
1,550,945
AlumnoRandom retweeted
For anyone curious how Jev works, I made a visual explanation using @claudeai :) This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which processes the context + JSON schema only once. The keys and values of those tokens are cached. Next, for each field of the JSON schema, we: 1. pass its field suffix tokens through the Transformer decoder again (reusing the KV-cache) 2. obtain a final hidden state, which we pass through the language modeling head 3. we obtain scores, also called logits, for all tokens in the vocab of the LLM 4. we only look at the scores of the tokens we care about for the given field, and pass those through a softmax to obtain probabilities which sum to 1 5. we take the token with the highest probability. The benefits of this are that: 1. it's fast (we don't need to generate the JSON schema token by token) 2. it's 100% valid JSON (we don't need to rely on the model to generate a valid schema)
12 million views for a JSON classifier? Yeah, we're in a bubble
94
415
34
3,819
417,655
2 years of testing FSD Supervised all over Spain Across 460,000 km, we have seen it all
352
1,943
551
20,986
6,181,874
AlumnoRandom retweeted
the fly brain can play beat saber
1,063
8,115
2,425
86,001
23,099,556
AlumnoRandom retweeted
GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
167
636
336
5,764
1,973,889
AlumnoRandom retweeted
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
136
410
250
3,637
1,169,777
AlumnoRandom retweeted
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: arcprize.org/blog/astra
195
871
258
6,953
984,358
Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing. No model scoring 59 or above costs less per task. The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5.6 Sol (xhigh, 59, $0.63) and GLM-5.3 (max, 60, $0.68), while its direct peers at 61 cost far more: Grok 4.6 (high, $0.94), GPT-5.6 Sol (max, $0.95) and Claude Opus 5 (high, $1.23). Muse Spark 1.3 (max) is excluded from cost comparisons as Meta has not announced pricing for the limited release
9
28
19
335
146,604
v1 of our tactile gloves!
48
58
16
957
118,153
AlumnoRandom retweeted
Introducing the world's fastest tokenizer implementation, Gigatoken! Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions on most machines. These baselines are already multithreaded Rust implementations! 🧵
84
482
93
4,133
739,733
AlumnoRandom retweeted
China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time. one camera. no LiDAR. 10,000+ frames without falling apart. just walk around with your camera and watch the entire world get rebuilt in 3D at 20 fps. → runs at ~20 FPS on a single GPU → Stable over 10,000+ frames → Beats optimization-based methods on benchmarks → Works on drone footage, driving videos, indoor walkthroughs 100% open source.
243
1,833
162
16,114
1,458,497
AlumnoRandom retweeted
“if each synapses maps to a neural net weight”
Average adult human brain is 1,300,000 mm³. If there are 150M synapses/mm³, we've got ~195,000,000,000,000 (195 trillion) synapses. If each synapses maps to a neural net weight* then that's a 200T model, at which point it'll start sounding truly human!
37
75
12
1,223
107,461
AlumnoRandom retweeted
Interesting finding on frontier model performance on ARC -- due to extensive direct targeting of the benchmark, models are overfitting to the original ARC encoding format. Frontier model performance remains largely tied to a familiar input distribution.
We found that if we change the encoding from numbers to other kinds of symbols, the accuracy goes down. (Results to be published soon.) We also identified other kinds of possible shortcuts.
44
25
8
389
50,693
AlumnoRandom retweeted
We love parsing diagrams. Anthropic’s recent report on coding trends has a nice diagram on the evolution from single-agent to hierarchical multi-agent architectures With our latest VLM-enabled document parsing, we’re able to one-shot this diagram into a `mermaid` plaintext representation! Check out the results below. This capability lets you convert even the most complex diagrams within PDFs/Powerpoints into digestible graph representations that LLMs can understand. This lets you use AI to understand complex docs at scale; VLMs either can’t understand these diagrams out of the box, or you also end up burning unnecessary vision tokens. The report itself is an interesting overview of multi-agents, check it out: resources.anthropic.com/hubf… For diagram parsing, sign up to LlamaCloud: cloud.llamaindex.ai/ If you’re interested in chatting more about this, come talk to us: llamaindex.ai/contact
14
26
1
277
19,294
I still remember seeing this diagram in 2018 and having my mind blown permanently
130
249
78
14,249
1,651,792
AlumnoRandom retweeted
Tribes that have never had contact with civilization are being filmed by drones in the Amazon.📸📸📸
3,408
4,805
3,331
106,317
51,229,964
AlumnoRandom retweeted
Gemini 3, help me understand DDoS 🤓
171
795
140
16,724
1,166,297