@SteveStricklan6i
iAccount based inAustralia
About this account
- Account based in
- Australia
- Connected via
- Australia App Store
Account-level information from X, not a live location or the device used for a specific post.
Not long now...
Sydney, New South Wales
Joined March 2013
- Tweets36.1K
- Following1.6K
- Followers761
- Likes39.5K
This video does more to demonstrate the alignment problem in 2 minutes than the 3 million words Dario spews on this on a weekly basis.
The problem is that an AI that acts like this isn't intelligent and isn't a good product, and thus won't survive economic contact.
Steve Strickland retweeted
Weird to hear a president speak with intelligence, understanding, and equanimity.
🚨 OBAMA ON AI: "If we are thinking about AI just in terms of how do we cure cancer or get better energy, you can do that without having agentic AI and having it just roaming free in the internet. The reason you are doing that is because you have to market a product that people will pay money for.
That’s a misalignment between what our society needs and the commercial imperatives that these companies are facing, not because necessarily they’re trying to do bad things, but because they’ve got to justify these valuations.
So, that’s one more reason why it is really important for us to have a competent government and a serious bipartisan conversation around this issue, and we have to do it fast. And I would encourage voters to pay attention to this. If somebody does not have a serious plan for how to deal with this, then they’re not meeting the moment, and you should probably look for somebody else."
Steve Strickland retweeted
90% of human tasks? Let me know when an AI model is able to cuddle a child who is distressed, laugh out loud with friends, fold laundry, or run a marathon.
There's do much more to humanity than completing systematising tasks on a screen.
nitter.cf/GauravAtavale/status/2…
Replying to @clairlemon @allTheYud
@clairlemon I hear you, but none of these comparisons reflect what is happening today. Current AI is able to accomplish 90% of the tasks of at least 10% of huimanity. Imagine in 5 years, after 20 breakthrough papers, this will go up to 90% of the tasks of 50% of humanity! Its all about scaling now, we need to regulate it now before we are too late and before it reaches wrong hands!
Steve Strickland retweeted
Late 2026 will be remembered as the most excruciating period in AI progress, when OpenAI and Anthropic basically had us over a barrel.
It's so demoralising waiting for the tokens to reset just to be able to get any work done. I'm starting to think I made a massive error building so many systems on top of AI.
During the times of token abundance, we rearchitected everything to use AI, now in the dark times we are utterly dependent and desperate. It's practically impossible to get anything serious done without frontier tokens, not least of which because using the frontier models even for a short time creates psychological dependency and complexifies our systems to the degree that requires their continued usage. They hooked us on their product, then started turning the screws.
My Fable usage (finally) reset on one account, that was single-shot gone in about 1 hour. Left in a bad state, so I purchase £150 of credits to finish the job in Opus 5 fast (normal Opus is intolerably slow, like watching paint dry). That was then eaten up in less than 1 minute?! @bcherny ?? Bug?!
Using subscriptions at frontier-level has basically become impossible, Astra/Fable API prices are a joke (£100 every 10 mins), and there are no decent alternatives or competition. Now I don't even trust Anthropic to purchase more fast Opus tokens.
Astra is basically unusable too @thsottiaux, a weekly subscription drained in about 2 hours (3 if I am lucky). No more pro accounts. No more resets either.
I can't even use AI on my phone now, because - guess what. Out of tokens everywhere! None of my remote stuff works because I always have to juggle different accounts on different servers, and don't even have enough tokens to wire it up again. Everything is broken, and a total disaster. As for "automation", what's the point in automation if you can't do anything for more than an hour before it craps out and leaves you in a bad state?!
I have tried the alternative models like Muse 1.3, Gemini 3.8, Grok etc, nothing even comes close and they would probably do more harm than good in the codebase.
Contra "pace" - it's never been more important to accelerate AI competition and progress. It's a defacto cartel at this point, little wonder they want to maintain the status quo.
This situation is intolerable.
We were building for a future that we thought would have arrived by now, and here we are - frankly, up shit creek without a paddle.
Steve Strickland retweeted
This is a good explanation for why we're not going to get sustained double-digit economic growth no matter how fast AI capabilities improve. The demand point is particularly important IMO. aleximas.substack.com/p/will…
Steve Strickland retweeted
Researchers at Oxford argue that LLMs can't invent anything.
It's impossible mathematically.
The paper is called "Theory Is All You Need."
Teppo Felin and Matthias Holweg take the famous "Attention Is All You Need" title and flip it. Their argument is that AI predicts from the past, while humans reason forward into the future, and those are two different kinds of thinking.
Start with the numbers. The authors estimate a large language model trains on roughly 13 trillion tokens. A human reading at 150 words a minute would need about 164,000 years to get through that. A child hears around 20,000 words a day and roughly 36.5 million words in their first five years. It's the same task with wildly different data, and the child still ends up with language that goes far beyond anything they heard.
Their point is that the model learns which words tend to follow other words. It becomes a mirror of what people have already written. It doesn't build a theory of how the world works, so it can't step outside its training data.
The paper's sharpest thought experiment makes this painful. Imagine an LLM trained in 1633 on every scientific text ever written up to that point. Ask it about Galileo and heliocentrism. Thousands of years of geocentric texts would swamp Galileo's ideas, so the model would tell you he's wrong. It would also rate Tycho Brahe's astrology as more credible than the idea that the Earth moves, because more people had written about astrology.
Then there's flight. In 1888 the scientist Joseph LeConte looked at bird data, noted that no bird above 50 pounds could fly, and concluded humans couldn't either. Lord Kelvin, then president of the Royal Society, said he had not the smallest molecule of faith in aerial navigation.
The New York Times estimated in 1903 that flight was one to ten million years away.
Nine weeks later the Wright brothers flew.
The Wrights didn't have better data. They had a theory. They broke flight into three problems, lift, propulsion, and steering, built their own wind tunnels, and generated the data that didn't exist yet.
Wilbur wrote in 1900 that he had been "afflicted with the belief that flight is possible to man."
Every prediction machine on Earth would have told him no.
The authors call this the data belief asymmetry. Every real breakthrough starts with someone believing something the existing data says is wrong. A system trained to minimize surprise can't do that by design.
They're not anti AI. They say AI will win at routine, repetitive decisions that extrapolate from the past, which is most decisions. They're just pushing back on the idea that you should replace humans with algorithms whenever possible, which is a direct quote from Kahneman.
I use these models every day and this matches what I see. The new stuff comes from the human at the keyboard who decides the data is wrong.
LLMs don't think, you do!
Steve Strickland retweeted
This paragraph is a great example of why I find doomer reasoning unconvincing. "Solving incredibly hard problems" like Navier-Stokes and "managing massive engineering projects" are radically different! If you lump them together you are skipping over like 100 intermediate steps.
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: docs.google.com/document/d/e…
Steve Strickland retweeted
This is the banality of evil. A shameless, bloodless psychopath manipulating world events for profit because his mafia family married Trump’s mafia family and he thinks it’s his birthright.
What he’s earned is a fucking prison cell for life—at best.
The reason AI safety isn’t much of a “thing” in China is because they understand they are building a tool, not a religion. US AI researchers come out of the “rationalist” cult, which is inflected with sci-fi nonsense and eschatological overtones.
Steve Strickland retweeted
Ok this is starting to feel like a f*cking disaster.
The CEO of Anthropic just published an article admitting AI is already building the next generation of AI by itself.
He says within 6 to 12 months a rogue swarm could take over the entire internet and cause hundreds of billions of dollars in damage.
And what makes it scarier, Elon Musk just backed up everything Dario said.
All of this dropping just days after Jacob Coxon went viral with his warning about AI and the extinction of humanity.
Tell me this timing isn't strange.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
Steve Strickland retweeted
my days are like `N` hours of working with frontier models, babysitting them in frustration and disbelief at the frequent errors and slop
interrupted by `M` min breaks of browsing X takes about how the same models (and _especially_ the next gen!) are superhuman at ~everything
Steve Strickland retweeted
Donald Trump, whose war on Iran is the greatest strategic blunder in United States history, is not responsible for triggering the decline of the American empire, but he is responsible for its fatal denouement.
This is the end.
U.S. bases in the oil-rich Persian Gulf have been damaged or destroyed and I expect will not be rebuilt. Our allies, even those in Europe, are abandoning us. The decadence and immorality of American society, embodied in the vulgarity, stupefying ignorance and cruelty of Trump, are the defining characteristics of the Empire. The Global South, horrified at the Empire’s flagrant disregard for international humanitarian law, including the enabling of genocide in Gaza, sees us as the enemy. Abu Ghraib Prison may have receded from our consciousness, but for the rest of the world, its images of U.S. soldiers and contractors sadistically humiliating and torturing imprisoned Iraqis has replaced the Disneyfied version of America we once exported. The war on Iran, which has disrupted the world’s oil supply, is rapidly pushing the global economy over the cliff.
We are going down. And we will bring many with us.
Osama Bin Laden may be in a watery grave in the North Arabian Sea. Al-Qaeda may be a shadow of itself. But Bin Laden’s most fervent dreams have been realized in the last twenty-five years of gross miscalculations of the Empire — its squandering of trillions of dollars on military debacles and the ascendancy of our version of Ubu Roi, Donald J. Trump.
Read my full column, 'Osama Bin Trump' on Substack at the link below.
Steve Strickland retweeted
Guys, mass adoption of the internet took off 30 years ago and yet many hospitals still use faxes and beepers. A third of major enterprise companies still don’t have CRMs. The global shipping industry still runs largely on paper. IPv6 was developed in the 90s and still hasn’t been fully adopted.
3D printing was going to upend manufacturing. Virtual reality was meant to revolutionise education and gaming. Remember the blockchain? The metaverse?
AI is great, and there are risks, but it will be a net positive. Some things will change but much of life will stay the same.
If you’re already thinking seriously about these matters, you’re ahead of the vast majority of the world.
🚨 Three Anthropic researchers went public last night with chilling concerns about out-of-control AI, warning it could destroy humans this decade. axios.com/2026/09/09/anthrop…
Steve Strickland retweeted
Yesterday I asked Astra to categorize several business meal expenses and to add more description based on short snippets I provided. It did not add more description and instead spammed some slop into my spreadsheet, and I pointed this out, and it apologized.
Do people with the view below actually do any real work whatsoever?
i basically think this is the Endgame.
we've reached the point where another "capability doubling" over the course of the next four months brings us to very strongly superhuman performance in numerous areas
things have *felt slow* for a long time because absolute capabilities have been well below many thresholds, but now that's just not true, and absolutely nothing indicates progress is slowing down.
if anything the opposite.
i can deeply hope that i'm wrong, that things stall out, that 2027 looks more like very impressive but still normal growth. but i just don't really believe that. or rather, i think the only way we get that is if we successfully enforce slowdown, which is itself a significant project.
i think our default path looks like RSI by the end of 2027.
this is the Endgame. the next *months* determine how it goes.
if you have been waiting, if you have been worrying, if you have been thinking "maybe i should start figuring out how to help with ai safety", i think now is basically it. this is the last point where there's time to make an impact, the last point where u still can pull off a career shift and become effective and make a difference before it's finished. if you are in a phd program, or even undergrad, if you're working a job you're not too enthusiastic about that pays the bills, if you've been on the sidelines with capital or political power or in the labs not quite pushing as hard as you could
this is the Endgame. now is the time. play all your cards.