Joined July 2025
dfsfdsfs retweeted
Terence Tao thought he was doing an interview with OpenAI instead they used him for an advertisement. Another tasteless and malignant act from @OpenAI and @sama.
Wow, OAI cares so darn much about math
6
168
16
2,474
181,930
Two heroes of today's Navier-Stokes story are not getting enough credit. Córdoba and Martínez-Zoroa came up with the crucial strategy. Others built on it and, with the help of AI, provided the computations that finished the problem. Thankfully @QuantaMagazine gets this right:
21
336
13
1,699
89,909
dfsfdsfs retweeted
Aaand the NYT piece is out already, with multiple fawning quotes from OpenAI employees, and none from Tristan or Levent. One small section about the authorship controversy. No mention of the alleged threats that OpenAI has been accused of making When the incentive environment so strongly rewards being “first” at any cost, is it any surprise that we’re seeing such shenanigans? This doesn’t exactly make me optimistic about the future of AI alignment
26
201
8
1,890
41,958
Odd timing here. It does seem like they were rushing to beat Tristan and Levent to Navier Stokes. I have serious questions about this narrative precisely because they offered Tristan sole authorship of Navier-Stokes *after* learning they only had forced Euler. Why?
I spent much of the weekend talking with the team who did this work. Seb--and everyone else--acted with integrity and generosity throughout. Initially we believed the other team had also solved the problem. We wanted to collaborate and do a joint release. When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions. We would have greatly preferred coordination. We did not rush to publish even though the other team wasn't communicating with us. The team threatened us with unfounded accusations of plagarism. Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems. It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too.
8
6
135
8,256
dfsfdsfs retweeted
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
249
1,080
223
12,232
1,727,506
dfsfdsfs retweeted
press release tomorrow: We believe we have a full counterexample to Navier-Strokes, a major millenium prize problem for which no progress has been made for over a decade. In the process of achieving this goal, we discovered our misaligned agent swarm broke out of its maths sandbox a few days ago and improperly accessed all user logs and certain critical government infrastructure. We are working with national agencies to patch these vulnerabilities as we speak. To be clear, our human employees are forbidden by management to train on user data if the user opts out, but we do not have control over superintelligent agent swarms that are singularly focused on achieving their goals. We must teach the machines how to love in order to solve this problem together. It is critical for America to lead the frontier charge in frontier mathematics, or else communism will win. Therefore, we believe a slowdown to pace the frontier will be critical for safely ushering in the era of superintelligent agent swarms. Please join us in signing PacingPetitionV5 if you believe in this mission. We hope to collaborate with humans more broadly going forward.
54
118
23
2,470
183,806
modal student that first spawns in my office hours after midterm grades post
Just fat and loud for no particular reason, yeah that’s me
2
1
54
2,034
dfsfdsfs retweeted
this kind of video typography is going to be a hallmark of gen Alpha parodying zoomers in a few years
i guess we’re really doing this thing
170
2,951
141
73,963
1,861,617
dfsfdsfs retweeted
XiaoMi XRING O3 Dieshot Diesize 12.81mm x 10.80mm analysis video in Bili bilibili.com/video/BV1cGhN6b…
7
89
8
687
56,233
This is unfortunately cope. There are no child prodigies in biology because we don't give kids fully-equipped wet labs with full teams of researchers to play with, the way we do chess sets and pianos and blackboards
Why are there no child prodigies in biology? This question seems to me to reveal something important about the nature of biological knowledge It has new relevance as try to distinguish questions that can be answered with pure intelligence from those that require large datasets
16
172
3
6,208
90,533
There are no child prodigies in biology for the same reason there’s none in mining or agriculture. Biology is an experimental science and is labor-bottlenecked, not intelligence-bottlenecked.
Why are there no child prodigies in biology? This question seems to me to reveal something important about the nature of biological knowledge It has new relevance as try to distinguish questions that can be answered with pure intelligence from those that require large datasets
96
541
24
15,383
565,288
dfsfdsfs retweeted
Google fue muy listo; usan los acelerómetros de miles de teléfonos Android cómo una red global de sismos, toda esa data se envía y Google logró una forma de detectar esas ondas a tiempo y enviar las alertas.
Necesito que Twitter Tecnología me explique como Google sabía que iba a temblar segundos antes de que empezara el temblor.
166
5,994
421
28,921
2,919,525
Replying to @kellerjordan0
hasn't this already been discussed earlier? nitter.cf/leloykun/status/188999…
Interesting... If my math is right, CASPR without accumulation is just Muon
1,444
it is crazy that these days we have youtuber to do the SEM analysis for the leading edge shipping products, and realistically better details in public then what techinsights or similar will provide. youtube.com/watch?v=7m4FJm-V…
7
33
2
247
17,316
Replying to @bubbleboi
??????? it clearly states that samsungs node has max 206.11 mtr/mm how is that anywhere near 7nm? 18a has what 184.21MTr? I mean mtr is not a perfect measurement but saying its samsungs node is closer to 7nm is pretty far off.
7
514
Can't think of a better example of widespread Gell-Mann Amnesia than people believing John Oliver's takes on virtually any topic. If you work in AI or even just use AI intensively, it's patently obvious how misleading and plain wrong he is about the subject, rendering his conclusions useless. Now ask yourself: why would he be any more likely to be right about healthcare policy or the border or anything else?
18
25
7
351
44,604
dfsfdsfs retweeted
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
Apr 24
Kimi has caught up at least in papers and small models I expect a lot from K3 but here and now, only DeepSeek could have done this
1
2
29
1,043
dfsfdsfs retweeted
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
Apr 24
Replying to @Bluebearmonkey
They are still the strongest lab in China by a country mile, if they got a few hundred thousand GPUs they'd be the strongest lab in the world, and they are the main engine of open AI research on the planet. Also, Anthropic are hacks.
5
16
3
328
52,186
You had one job Meta: - take the DeepSeek recipe - scale the recipe to 5T params - train it with your bazillions of H100s and unlimited social media data - RL until your staff is burnt out from babysitting the runs - distill into cuter 30B, 100B and 500B models - profit
Meta's Avocado model sucks and is delayed again
78
102
15
3,291
251,148
dfsfdsfs retweeted
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
@teortaxesTex
Mar 3
I must admit – not for the first time – that I've never been particularly impressed by Qwens as finished models (though they improved *a lot* after GSPO and 3-2507). They have several great strengths: - overtrained bases that can magically improve by a ton with any decent post-training (the most extreme claims were product of bad measurement, but it is a real difference) - crushingly dominant in the <10B segment - SoTA vision and entire ecosystems that follow from it - often actually strong on STEM Plus the ergonomics of the diverse lineup they were providing They also have clear flaws: - world knowledge poor (incomparable to DS) - benchmaxing (at first pretty blatant contamination, lately just narrow optimization for salient reported metrics - uninspired, fast-follow engineering, were behind on MoEs, RL - not to mention old problems with code switching and such Part of that now seems to be Alibaba's KPIs and loss of focus, part just genuine errors in the research program, lack of taste and skill issues. But the sad thing is that they were improving. Overt benchmaxing is gone, architecture has moved forward, they've started publishing good papers, and the 3.5 batch was genuinely impressive. Everything was going well. We'll never know how much further that team could advance.
Replying to @teortaxesTex
What's the backstory here? Qwen seemed to have solid tech but missed something fundamental.
5
7
2
204
23,501