AlumnoRandom retweeted
1/ We introduce TrackEverything: a 3D point tracker that tracks all points across all frames of long videos (1000+ frames).
Our key idea is to tie computation to unique 3D scene content rather than redundant 2D pixels in a video.
trackeverything.github.io
AlumnoRandom retweeted
In the Utopia mouse experiment, they put mice in a “utopia” with unlimited food and no predators
>Young male mice spent all their time grooming, uninterested in reproducing
>Female mice lost maternal instinct and some harmed their own kids
>birth rates went to zero
Uhhhh
AlumnoRandom retweeted
For anyone curious how Jev works, I made a visual explanation using @claudeai :)
This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface
The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which processes the context + JSON schema only once. The keys and values of those tokens are cached.
Next, for each field of the JSON schema, we:
1. pass its field suffix tokens through the Transformer decoder again (reusing the KV-cache)
2. obtain a final hidden state, which we pass through the language modeling head
3. we obtain scores, also called logits, for all tokens in the vocab of the LLM
4. we only look at the scores of the tokens we care about for the given field, and pass those through a softmax to obtain probabilities which sum to 1
5. we take the token with the highest probability.
The benefits of this are that:
1. it's fast (we don't need to generate the JSON schema token by token)
2. it's 100% valid JSON (we don't need to rely on the model to generate a valid schema)
AlumnoRandom retweeted
2 years of testing FSD Supervised all over Spain
Across 460,000 km, we have seen it all
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
AlumnoRandom retweeted
GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we've seen
Our analysis:
AlumnoRandom retweeted
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Read our post on Astra and what these results mean: arcprize.org/blog/astra
AlumnoRandom retweeted
Muse Spark 1.3 (xhigh) is the most cost-efficient model at its intelligence level: $0.55 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing. No model scoring 59 or above costs less per task. The nearest are Gemini 3.8 Flash (high, 59, $0.58), GPT-5.6 Sol (xhigh, 59, $0.63) and GLM-5.3 (max, 60, $0.68), while its direct peers at 61 cost far more: Grok 4.6 (high, $0.94), GPT-5.6 Sol (max, $0.95) and Claude Opus 5 (high, $1.23). Muse Spark 1.3 (max) is excluded from cost comparisons as Meta has not announced pricing for the limited release
AlumnoRandom retweeted
Introducing the world's fastest tokenizer implementation, Gigatoken!
Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions on most machines.
These baselines are already multithreaded Rust implementations! 🧵
AlumnoRandom retweeted
China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time.
one camera. no LiDAR. 10,000+ frames without falling apart.
just walk around with your camera and watch the entire world get rebuilt in 3D at 20 fps.
→ runs at ~20 FPS on a single GPU
→ Stable over 10,000+ frames
→ Beats optimization-based methods on benchmarks
→ Works on drone footage, driving videos, indoor walkthroughs
100% open source.
AlumnoRandom retweeted
“if each synapses maps to a neural net weight”
AlumnoRandom retweeted
Interesting finding on frontier model performance on ARC -- due to extensive direct targeting of the benchmark, models are overfitting to the original ARC encoding format.
Frontier model performance remains largely tied to a familiar input distribution.
Replying to @MelMitchell1 @mikeknoop
We found that if we change the encoding from numbers to other kinds of symbols, the accuracy goes down. (Results to be published soon.)
We also identified other kinds of possible shortcuts.
AlumnoRandom retweeted
We love parsing diagrams. Anthropic’s recent report on coding trends has a nice diagram on the evolution from single-agent to hierarchical multi-agent architectures
With our latest VLM-enabled document parsing, we’re able to one-shot this diagram into a `mermaid` plaintext representation! Check out the results below.
This capability lets you convert even the most complex diagrams within PDFs/Powerpoints into digestible graph representations that LLMs can understand. This lets you use AI to understand complex docs at scale; VLMs either can’t understand these diagrams out of the box, or you also end up burning unnecessary vision tokens.
The report itself is an interesting overview of multi-agents, check it out: resources.anthropic.com/hubf…
For diagram parsing, sign up to LlamaCloud: cloud.llamaindex.ai/
If you’re interested in chatting more about this, come talk to us: llamaindex.ai/contact
AlumnoRandom retweeted
I still remember seeing this diagram in 2018 and having my mind blown permanently
AlumnoRandom retweeted
Tribes that have never had contact with civilization are being filmed by drones in the Amazon.📸📸📸