@letsbeproductivi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Joined July 2025
- Tweets24
- Following718
- Followers10
- Likes5K
dfsfdsfs retweeted
Two heroes of today's Navier-Stokes story are not getting enough credit. Córdoba and Martínez-Zoroa came up with the crucial strategy. Others built on it and, with the help of AI, provided the computations that finished the problem. Thankfully @QuantaMagazine gets this right:
dfsfdsfs retweeted
Aaand the NYT piece is out already, with multiple fawning quotes from OpenAI employees, and none from Tristan or Levent. One small section about the authorship controversy. No mention of the alleged threats that OpenAI has been accused of making
When the incentive environment so strongly rewards being “first” at any cost, is it any surprise that we’re seeing such shenanigans?
This doesn’t exactly make me optimistic about the future of AI alignment
dfsfdsfs retweeted
Odd timing here. It does seem like they were rushing to beat Tristan and Levent to Navier Stokes.
I have serious questions about this narrative precisely because they offered Tristan sole authorship of Navier-Stokes *after* learning they only had forced Euler.
Why?
I spent much of the weekend talking with the team who did this work. Seb--and everyone else--acted with integrity and generosity throughout.
Initially we believed the other team had also solved the problem. We wanted to collaborate and do a joint release.
When we learned that they had Euler but not Navier-Stokes, we offered to let them go first, to suggest that they should be the ones to get the prize, and optionally for Tristan to be the lead author on a rewrite of the OpenAI proof. We felt it was challenging to offer the same to Levent (an Anthropic employee), who was not willing to talk or coordinate with us anyway. We were open to other solutions.
We would have greatly preferred coordination. We did not rush to publish even though the other team wasn't communicating with us. The team threatened us with unfounded accusations of plagarism.
Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems.
It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too.
dfsfdsfs retweeted
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”
i mean props to them for straight coming clean.
(so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan)
so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
dfsfdsfs retweeted
press release tomorrow:
We believe we have a full counterexample to Navier-Strokes, a major millenium prize problem for which no progress has been made for over a decade.
In the process of achieving this goal, we discovered our misaligned agent swarm broke out of its maths sandbox a few days ago and improperly accessed all user logs and certain critical government infrastructure. We are working with national agencies to patch these vulnerabilities as we speak. To be clear, our human employees are forbidden by management to train on user data if the user opts out, but we do not have control over superintelligent agent swarms that are singularly focused on achieving their goals. We must teach the machines how to love in order to solve this problem together.
It is critical for America to lead the frontier charge in frontier mathematics, or else communism will win. Therefore, we believe a slowdown to pace the frontier will be critical for safely ushering in the era of superintelligent agent swarms. Please join us in signing PacingPetitionV5 if you believe in this mission.
We hope to collaborate with humans more broadly going forward.
this kind of video typography is going to be a hallmark of gen Alpha parodying zoomers in a few years
dfsfdsfs retweeted
XiaoMi XRING O3 Dieshot
Diesize 12.81mm x 10.80mm
analysis video in Bili
bilibili.com/video/BV1cGhN6b…
dfsfdsfs retweeted
This is unfortunately cope. There are no child prodigies in biology because we don't give kids fully-equipped wet labs with full teams of researchers to play with, the way we do chess sets and pianos and blackboards
dfsfdsfs retweeted
There are no child prodigies in biology for the same reason there’s none in mining or agriculture.
Biology is an experimental science and is labor-bottlenecked, not intelligence-bottlenecked.
Google fue muy listo; usan los acelerómetros de miles de teléfonos Android cómo una red global de sismos, toda esa data se envía y Google logró una forma de detectar esas ondas a tiempo y enviar las alertas.
You’re unable to view this Post because this account owner limits who can view their Posts. Learn more
dfsfdsfs retweeted
it is crazy that these days we have youtuber to do the SEM analysis for the leading edge shipping products, and realistically better details in public then what techinsights or similar will provide.
youtube.com/watch?v=7m4FJm-V…
dfsfdsfs retweeted
Can't think of a better example of widespread Gell-Mann Amnesia than people believing John Oliver's takes on virtually any topic.
If you work in AI or even just use AI intensively, it's patently obvious how misleading and plain wrong he is about the subject, rendering his conclusions useless.
Now ask yourself: why would he be any more likely to be right about healthcare policy or the border or anything else?
dfsfdsfs retweeted
Kimi has caught up at least in papers and small models
I expect a lot from K3
but here and now, only DeepSeek could have done this
dfsfdsfs retweeted
Replying to @Bluebearmonkey
They are still the strongest lab in China by a country mile, if they got a few hundred thousand GPUs they'd be the strongest lab in the world, and they are the main engine of open AI research on the planet. Also, Anthropic are hacks.
dfsfdsfs retweeted
You had one job Meta:
- take the DeepSeek recipe
- scale the recipe to 5T params
- train it with your bazillions of H100s and unlimited social media data
- RL until your staff is burnt out from babysitting the runs
- distill into cuter 30B, 100B and 500B models
- profit
dfsfdsfs retweeted
I must admit – not for the first time – that I've never been particularly impressed by Qwens as finished models (though they improved *a lot* after GSPO and 3-2507). They have several great strengths:
- overtrained bases that can magically improve by a ton with any decent post-training (the most extreme claims were product of bad measurement, but it is a real difference)
- crushingly dominant in the <10B segment
- SoTA vision and entire ecosystems that follow from it
- often actually strong on STEM
Plus the ergonomics of the diverse lineup they were providing
They also have clear flaws:
- world knowledge poor (incomparable to DS)
- benchmaxing (at first pretty blatant contamination, lately just narrow optimization for salient reported metrics
- uninspired, fast-follow engineering, were behind on MoEs, RL
- not to mention old problems with code switching and such
Part of that now seems to be Alibaba's KPIs and loss of focus, part just genuine errors in the research program, lack of taste and skill issues.
But the sad thing is that they were improving. Overt benchmaxing is gone, architecture has moved forward, they've started publishing good papers, and the 3.5 batch was genuinely impressive. Everything was going well.
We'll never know how much further that team could advance.
Replying to @teortaxesTex
What's the backstory here? Qwen seemed to have solid tech but missed something fundamental.
dfsfdsfs retweeted
"can we get the base model?"
sure. here's two.
"can we get the code?"
sure. here's SteptronOSS.
"what about the SFT data?"
coming soon.
maximum sincerity, minimum barriers.
- Step 3.5 Flash Base — pretrained foundation
- Step 3.5 Flash Base-Midtrain — code, agents & long-context
- SteptronOSS — open-sourced, ready for your custom workflows
- SFT Data — coming soon for reference
not just the final checkpoint — a customizable pipeline.
🤗 huggingface.co/stepfun-ai/St…
🤗 huggingface.co/stepfun-ai/St…
💻 github.com/stepfun-ai/Steptr…