@srchvrs

Machine learning scientist and engineer @Oracle speaking πtorch & C++. Past @awscloud, @LTIatCMU. Opinions sampled from MY OWN 100T param LM.

Pittsburgh, PA
Joined November 2009
🧵📢Attention folks working on LONG-document ranking & retrieval! We found evidence of a PROFOUND issue in existing long-document collections, most importantly MS MARCO Documents. It can potentially affect all papers comparing different architectures for long document ranking.⏩
4
13
6
134
39,910
It uses GRAG instead! (GRep-augmented generation) PS: jokes inside, it's still RAG even if retrieval is grep-based.
claude code doesn't use RAG either. i wouldn't have expected grep to be this good, but seems like simplicity won, again
3
28
2,413
🧵I actually find the code written agents to be quite good (at least in Python). However, it often implements what's no needed and doesn't implement what's necessary. The intent drift is very real, which is not surprising: You can't get 1000 bits of code from 1 bit of input without making heavy assumptions. ↩️
One of the things that irk me about using coding agents is how shitty their tests are and how defensive their code can be. These fuckers insist on wrapping everything with a try-exept clause and the exceptions being caught are so broad that there will be silent failures. The worst part is that a review agent will read this and be like "this is magnificent, LGTM".
3
9
1,254
What is important about tests though. The hard part of testing was never a unit test, it was realistic integration testing and testing code in the realistic environment, which is not deterministic and which you don't fully understand. Unit tests made by agents are indeed a bit simplistic. 🟦
1
5
148
We're witnessing the rise of a new profession: slopware engineering.
To me the biggest risk of AI is brain rot. I have more and more situations like this. Some coworker shares a report on something they did. I read, and find something fishy. I ask about and get no answer. I ask again and I'm told: the agent said that. And no follow up. It's clear form the interaction that they did not read the report they shared. Their brain is not used as it should be. I am not against AI at all (fortunate for something working at NVIDIA). I use AI a lot. But I use it as a tool. I own the result and whenever I see something fishy I ask about its output I force a change if I am not convinced. I wish there was a way to force people to read the AI output they share with others.
1
4
1,233
BTW: A related article on the topic titled "“Slop Grenades”: Shopify CEO Who Pushed Staff to Use AI Is Now Horrified by What He’s Wrought" futurism.com/artificial-inte…
🧵So you got Codex or Claude Code and thought you had cheated the gods of software development. Suddenly you have superpowers. You can land 10K-, 20K-, even 50K-line PRs. Code production has exploded. @bcherny , @DarioAmodei , and even @sama are applauding. Software engineering has finally been solved. ↩️
1
1
1,194
🧵So you got Codex or Claude Code and thought you had cheated the gods of software development. Suddenly you have superpowers. You can land 10K-, 20K-, even 50K-line PRs. Code production has exploded. @bcherny , @DarioAmodei , and even @sama are applauding. Software engineering has finally been solved. ↩️
5
2
1
10
4,116
Now the gods are laughing. Software engineering learned decades ago that typing out code is not where most of the money goes. Testing, debugging, verification, integration, and fixing defects can dominate development effort. ↩️
1
1
3
368
You made code production 10× cheaper, while potentially introducing enough extra defects to make everyone debug like there’s no tomorrow. Congratulations on your productivity revolution! PS: All characters and events in this post are fictional. Any resemblance to actual organizations, projects, or engineering practices is purely coincidental. 🟦
3
3
333
Leo Boytsov retweeted
This is one of our initial attempts of using applying diffusion models. The plug-and-play property of Uno make it flexible in many areas, think about using them in post training or other complex scenarios. Further, think about it this way: diffusion enhances both depths (denoising steps) and length (parallel testing scaling), it may has the potential to unlock stronger reasoning.
Today’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality. With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality. Paper: arxiv.org/abs/2609.04010 Model available at: huggingface.co/IFM/K2-Horizo…
8
34
2,366
🧵People who claim that SWEs are now managers who just guide their subordinates: Where have you seen managers splitting managers splitting a big problem into multiple small ones, planning each small problem, giving very detailed instructions, checking understanding (and possibly correcting guidelines afterwards), and then even reading outputs? ↩️
2
5
732
That's not management. That's engineering with a very unusual programming interface. The analogy gets especially misleading when the supposed “manager” has to understand the subordinate’s work well enough to detect subtle bugs. If I tell an agent, “implement X,” and then I have to inspect 2,000 lines of generated code, understand its architecture, notice that it misunderstood an invariant, rewrite the instructions, and verify the fix, I haven't magically become a people manager. I've acquired a new programming interface.🟦
1
293
As the US gov considers messing with OPT in a fit of anti-immigrant nonsense, just want to say to all the immigrant students out there that this is: a) not popular b) not reflective of how people feel about you
55
21
3
453
80,120
Leo Boytsov retweeted
Compute Polynomials Twice as Fast - Here's a fun result, done entirely without AI. That because we proved it years ago, but it ran a hundred pages and we weren't quite sure enough to publish. Now AI verified it in Lean. A fun fact about polynomials is that P(x) = x⁴ + a₃x³ + a₂x² + a₁x + a₀ can be evaluated in just two(!) multiplications! The trick is to write y = (x + b₀)x + b₁ P(x) = (y + x + b₂)y + b₃ where b₀…b₃ are easy to calculate from a₀…a₃. Polynomials are everywhere from computing exp/sin, to cryptographic hashes and codes. Over finite fields multiplications are particularly expensive, so a 2x speedup matters. Donald Knuth and others showed ~n/2 multiplications suffice for any degree n, but their preprocessing needs complex roots: numerically unstable and useless over finite fields. Rabin & Winograd fixed that with rational preprocessing, but 2logn extra multiplications, which hurts at the small n we most care about. Our new method solves this. It uses ⌈(n+1)/2⌉ multiplications, and we prove this is optimal. You can try it out on your own polynomials at thomashale.com/fast-polynomi…
20
95
5
771
38,124
In 2015, @ravisujith gave a colloquium talk a substantial part of which was devoted to a large-scale distributed LSH system at Google. I sent a follow up e-mail with pointers to our NMSLIB library, which already included a pre-HNSW graph-based retrieval algorithm SW-graph. Around the same time, SW-graph won the first ANN-benchmarks. I wasn't the inventor of the algorithm, but I had substantially (by an order of magnitude) sped up the initial implementation of SW-graph (a port from Java). Understandably, Sujith Ravi (and other people too) was skeptical and replied that "the major challenge is to scale to large (& distributed) graphs for high dimensional data while maintaining efficiency & good approximation quality compared to methods like LSH". I am happy to see the progress we have made since then.
I’m happy to share that our team's work on PiPNN—an ultra-scalable algorithm for high-dimensional Nearest Neighbor Search—has won three awards at KDD’26, VecDB’26, and SISAP’26 with up to 78x speedup on index build. Nearest-neighbor search is fundamental to deduping in data curation and embedding-driven retrieval for RAG for LLMs. It is great to see our work recognized across the community: Best Paper at KDD’26: PiPNN: A Framework for Ultra-Scalable Graph-Based Nearest Neighbor Indexing. This paper introduces the PiPNN framework, which delivers an order-of-magnitude speedup for building navigable proximity graphs. (dl.acm.org/doi/10.1145/37708…) Best Paper at VecDB’26: Turbocharging PiPNN for Proximity and 𝑘-NN Graph Building. The paper improved the original PIPNN kernels, resulting in an additional 8x speedup (up to a 78x total speedup!), and introduced a highly optimized GPU version. (openreview.net/forum?id=8Gzb…) 1st Place at the SISAP’26 Challenge: Our PiPNN-based solution ranked 1st for Task 1, which challenged teams to develop the fastest k-NN graph building algorithm. (sisap-challenges.github.io/2…) Nearest-neighbor search is an important part of Data Curation for LLMs, and retrieval for RAG pipelines, and it only becomes more important as we move to more long-horizon agents and handling long-context inference-time scaling. Congratulations to the team behind this work: Tobias Rubel, Richard Wen (UMD), Guy Blelloch, Laxman Dhulipala, Lars Gottesbüren(github.com/larsgottesbueren), Jakub Łącki, and Vahab Mirrokni.
3
649
PS: BTW, although scalability has always been somewhat of an issue, but even a 2015 SW-graph was 1. reasonably scalable to run in production (after all indexing cost is often small compared to retrieval cost). 2. was already faster than a typical LSH approach. It would have been slower than FALCON-LSH due to a bug, but it was fixed circa 2016.
In 2015, @ravisujith gave a colloquium talk a large part of which was devoted to a large-scale distributed LSH system at Google. I sent a follow up e-mail with pointers to our NMSLIB library, which already included a pre-HNSW graph-based retrieval algorithm SW-graph. Around the same time, SW-graph won the first ANN-benchmarks. I wasn't the inventor of the algorithm, but I had substantially (by an order of magnitude) sped up the initial implementation of SW-graph (a port from Java). Understandably, Sujith Ravi (and other people too) was skeptical and replied that "the major challenge is to scale to large (& distributed) graphs for high dimensional data while maintaining efficiency & good approximation quality compared to methods like LSH". I am happy to see the progress we have made since then.
There's a new version of this post
230