@rishdotuk

Long document understanding, Multilingual Evals and efficient models mainly, but other #NLProc applications in free time | vim enthusiast

Joined April 2011
Thank you @kunalb11 for this masterstroke to remove that @WhatsApp spamming of messages from every business without explicit consent.
Whatsapp is now going to charge every message sent from October Costs will be 14 paise for each message, a 10,000 daily user support now costs 2L per month from 0 earlier SMS era rates are back, expect a lot more businesses to move to apps Upside maybe reduced Whatsapp spam
2
29
Rishu Kumar retweeted
So yeah how is nobody in jail for all these AI cyber incidents? Aaron Schwartz was prosecuted to death for like 0.0001% of this.
52
278
21
3,110
55,788
Rishu Kumar retweeted
Who wants more speech training data? YODAS v3 is now available @huggingface It’s 1.1M hours - the biggest audio dataset ever. And the first at this scale with stereo audio at 48kHz, with timestamped transcripts and translations. 100+ langs hf.co/blog/espnet/yodasv3 CC-BY-3.0
17
57
11
375
30,448
Rishu Kumar retweeted
Apparently you can now outsource your research experience, recommendation letters, and academic references for $3,000. I received the proposal today. 😭
49
26
11
911
176,425
Rishu Kumar retweeted
Exactly this, it's not only not healthy but it's questionable science. Even with fast pace it's so important to stop and think about whether what you are doing even make sense (don't do things just because you can)
JEV was released in "limited early access" on September 15, that's two weeks ago. How come we have 29 papers about it on arxiv (and, I assume, ICLR submissions) already. This is not healthy.
2
20
1,190
Rishu Kumar retweeted
11
83
2
1,204
20,028
Gonna go to the samsung service center and have my humiliation ritual of paying someone to replace the battery on my phone when I can do it myself way quicker if they sold me the freaking battery.
5
82
Rishu Kumar retweeted
the gemma 4 12b is a beautiful/underappreciated model man... fully linear / standardized multimodal embeddings... reasonably deep... not so large where being dense vs MoE starts to be questionable (i.e 30b)... say what you want about google's RL, but their pretrains are elite
24
16
509
18,581
Rishu Kumar retweeted
hey guys! look at this extremely useless and buggy thing my claude cooked up in just a few prompts! i haven't looked at the code! its in rust though! it's fast for sure, im going to tweet about this now with a link to my github because nobody else must've done this before
9
13
2
348
7,921
Rishu Kumar retweeted
Finally, Apple has reached feature parity with GNOME
Well, bugOS 27 is unreal.
76
1,073
24
24,491
1,314,251
Rishu Kumar retweeted
What BigToken doesn’t want you to know is that they also read arxiv and X to get new ideas. Models are good at magnitude, direction is still hard. There is some uncertainty around for however long though, but its not today Please continue the amazing works like looping transformers and echo.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
8
18
320
20,718
Yep. This is a toxic (and crazy) outlook. There are plenty of important problems that basic research will solve outside of labs. Much of this pessimism btw is created by people that have never experienced excellence in pure research and thus don't have a frame of reference.
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
8
7
1
154
13,512
One of the worst forms of brainrot AI has cultivated is pessimism about basic research, ie the idea that important work can only happen inside a frontier/neo lab and only with 10k+ GPUs, so the rest should not even bother. What a bleak way to think about science. And it's false.
71
233
31
2,009
179,111
Watching my model do hillclimbing in output is making me hungry. what is this.
1
131
I see people being sarcastic but there is a real demand for it: if you want to embed models in industrial processes you need to ease audit and doing local language (and even more local professional register) remove frictions.
Reasoning models think in English, even on German prompts. We asked what it costs to make one think in German. Answer: a valley. Small doses of German reasoning data hurt, large doses mostly recover.
6
4
96
5,247
It is a tool to review/make things. MFs act as if they are debugging cuda-graphs everytime they are using an LLM, while the best they end up doing was to center the div on a slop website or write a stupid python script which is painfully so verbose that will make Guido screech.
I just saw someone using Antigravity in 2026… Like seriously?? 💀
2
144
Rishu Kumar retweeted
Uh yeah what did you guys think LLM as a judge meant
BANGALORE DISTRICT COURT POSTED THE CHATGPT PROMPT IN THEIR JUDGMENT 😭😭 SON😭😭😭 indiankanoon.org/doc/1955248…
9
25
4
819
64,849
Rishu Kumar retweeted
4
1
11
435
Heck yeaahhhh!
did you miss me?
90