@jaschasd

Member of the technical staff @ Anthropic. Most (in)famous for inventing diffusion models. AI + physics + neuroscience + dynamics.

San Francisco
Joined August 2009
My first blog post ever! Be harsh, but, you know, constructive. Too much efficiency makes everything worse: overfitting and the strong version of Goodhart's law sohl-dickstein.github.io/202… 🧵
43
193
37
1,036
Jascha Sohl-Dickstein retweeted
Very excited to share our interview with @polynoamial on AI for math — the Erdős unit distance problem, saturating the IMO, the future of math research, and more!
24
79
11
642
194,335
Jascha Sohl-Dickstein retweeted
Excited to share @Standard_Kernel's seed round and some reflections on what we’ve learned about kernel generation and what we believe is next. Grateful to our amazing team, supporters, and the broader community pushing this space forward.
48
45
23
520
139,365
Replying to @tedlieu
They didn’t. Altman left it generic saying they agreed to these two principles, and didn’t outline the specific term phrasing. Altman is a master of double speak. What he means is they’ve got rules on both of those but the rules are more lax than the Anthropic ones. For example Anthropic considers combing through metadata collected via dragnet to count as spying on American’s. It’s possible OpenAI considers that to not be “spying” since it’s not doing the collection. Anthropic requires *all* autonomous weapons decisions to have human sign off and not be the sole choice of the AI. OpenAI could easily have a ton of carve outs or conditions and still claim they have “principles on autonomous weapons systems” Whatever it is, they *are* letting the DoD do something with their systems that Anthropic refused on moral grounds.
45
197
21
2,624
166,466
Jascha Sohl-Dickstein retweeted
I feel like I am going insane and no one has read the articles. It appears that OpenAI has not brought about harmony and still has the "all lawful use" clause in their contract that was the issue in the first place? I think they've negotiated functionally the same contact they've always been offered.
For the avoidance of doubt, the OpenAI - @DeptofWar contract flows from the touchstone of “all lawful use” that DoW has rightfully insisted upon & xAI agreed to. But as Sam explained, it references certain existing legal authorities and includes certain mutually agreed upon safety mechanisms. This, again, is a compromise that Anthropic was offered, and rejected. Even if the substantive issues are the same there is a huge difference between (1) memorializing specific safety concerns by reference to particular legal and policy authorities, which are products of our constitutional and political system, and (2) insisting upon a set of prudential constraints subject to the interpretation of a private company and CEO. As we have been saying, the question is fundamental—who decides these weighty questions? Approach (1), accepted by OAI, references laws and thus appropriately vests those questions in our democratic system. Approach (2) unacceptably vests those questions in a single unaccountable CEO who would usurp sovereign control of our most sensitive systems. It is a great day for both America’s national security and AI leadership that two of our leading labs, OAI and xAI have reached the patriotic and correct answer here 🇺🇸
8
11
1
256
26,481
Jascha Sohl-Dickstein retweeted
For those wondering how mass domestic surveillance could be consistent with "all lawful use" of AI models, I recommend a declassified report from the ODNI on just how much can be done with commercially available data (CAI): "...to identify ever person who attended a protest"
18
176
16
593
93,338
Jascha Sohl-Dickstein retweeted
Outside Anthropic's office in SF... intense moment!
475
4,773
502
32,954
1,888,076
This is the kind of day that makes me feel good about my life choices. An unusual thing about Anthropic is that we are trying hard on the inside to do what we say we're trying to do on the outside. The conflict with the DoW isn't great, but it is nice when our commitment to our values is legible. ... and in case you also want to feel very proud of where you work ... we're hiring ...
20
21
1
666
29,332
Jascha Sohl-Dickstein retweeted
New Anthropic Fellows research: How does misalignment scale with model intelligence and task complexity? When advanced AI fails, will it do so by pursuing the wrong goals? Or will it fail unpredictably and incoherently—like a "hot mess?" Read more: alignment.anthropic.com/2026…
152
219
85
1,914
531,042
When AI fails, will it do so by coherently pursuing the wrong goals? Or will it fail the way humans often fail, and take incoherent actions that don't pursue any consistent goal. In other words, like a “hot mess?” How will this change when AI performing limited tasks transitions to AGI performing tasks of unbounded complexity? How does misalignment scale with model intelligence and task complexity? We measure this using a bias-variance decomposition of AI errors. Bias = consistent, systematic errors (reliably achieving the wrong goal). Variance = inconsistent, unpredictable errors. We define "incoherence" as the fraction of error from variance. I am very excited about this framing, because it characterizes types of misalignment in a way that should be amenable to simple theoretical models and clean scaling laws.
3
8
1
118
10,202
Blog post: alignment.anthropic.com/2026… Full paper: arxiv.org/abs/2601.23045 Alex Hägele @haeggee did a superb job leading this project, which he did as part of the Anthropic Fellows Program. Also thank you to collaborators Aryo Gemma, @sleight_henry, @EthanJPerez We will be presenting the paper at ICLR 2026.
2
2
28
3,339
And cross linking that Anthropic account tweet thread: nitter.cf/AnthropicAI/status/201…
New Anthropic Fellows research: How does misalignment scale with model intelligence and task complexity? When advanced AI fails, will it do so by pursuing the wrong goals? Or will it fail unpredictably and incoherently—like a "hot mess?" Read more: alignment.anthropic.com/2026…
4
2,979
Title: Advice for a young investigator in the first and last days of the Anthropocene Abstract: Within just a few years, it is likely that we will create AI systems that outperform the best humans on all intellectual tasks. This will have implications for your research and career! I will give practical advice, and concrete criteria to consider, when choosing research projects, and making professional decisions, in these last few years before AGI. This is my current go-to academic talk. It's mostly targeted at early career scientists. It gets diverse and strong reactions. Let's try it here. Posting slides with speaker notes... -- The title is a play on a very opinionated and pragmatic book by the nobel prize winner ramon y cajal, who is one of the founders of modern neuroscience. To get you in the right mindset, on the right we have a plot of GDP vs time. That is you, standing precariously on the top of that curve. You are thinking to yourself -- I live in a pretty normal world. Some things are going to change, but the future is going to look mostly like a linear extrapolation of the present. And the plot should suggest that this may not be the right perspective on the future. This plot by the way looks surprisingly similar even if you plot it on a log scale. We didn't stabilize on our current rate of growth until around 1950.
57
272
61
1,802
347,235
Let's end on a high note! Here is a followup to the earlier slide, where I mentioned that cancer rates were exponentially falling. Doesn't this image make you happy? (less happy that we need to filter by rich countries. but if we keep making it easier to beat cancer, the rest of the world will catch up)
2
4
69
14,643