@kroalisti
iAccount based inFrance
About this account
- Account based in
- France
- Connected via
- France App Store
Account-level information from X, not a live location or the device used for a specific post.
highly esoteric
Joined January 2013
- Tweets10.7K
- Following1.4K
- Followers625
- Likes3.1K
1) what ?
Jev + Astra beats the Ender Dragon in Minecraft in 8 minutes 43 seconds! ⏱️
Cost less than $1 ($0.01 Jev, $0.96 Astra)
I open sourced the code and explain the harness setup below. This type of movement is only possible with Jev's near instant decisionmaking, and some continually learning skills from Astra.
Arthur 👁 retweeted
Jev + Astra beats the Ender Dragon in Minecraft in 8 minutes 43 seconds! ⏱️
Cost less than $1 ($0.01 Jev, $0.96 Astra)
I open sourced the code and explain the harness setup below. This type of movement is only possible with Jev's near instant decisionmaking, and some continually learning skills from Astra.
Not to be that guy but when the gov clearly doesn’t wanna fund this and there are no financial incentives for any company to invest aswell, at some point who do you think should pay for this ?
We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years. anthropic.com/news/accenture…
Readers added context they thought people might want to know
Anthropic presents this as an "independent evaluation" but will directly fund Accenture's work and has a prior commercial partnership with the firm for deploying its models, including training ~30,000 Accenture professionals on Claude.
anthropic.com/news/accenture…
anthropic.com/news/anthropic…
Arthur 👁 retweeted
Replying to @edzitron
Yeah, this is a really important question.
I’ve been associating with rationalists for almost a decade. Their ideas were compelling and you can’t deny their foresight. The founders of every superintelligence effort are all bound up in the rationality sphere. Sam Altman and Greg Brockman used to frequent the MIRI offices.
The way I relate to rationalists is it’s like the bizarre hypertrophied physiques on specialized athletes. In my opinion they have weirdly specialized minds that make them uniquely able to be correct about topics removed from daily experience. These “specialized minds” also can make them off-putting to many people.
I dont know if I’d consider myself a rationalist. I am intuitively opposed to polyamory, veganism, and consequentialist intuitions. But I certainly trust them to be correct about factual questions.
Arthur 👁 retweeted
Putting aside how idiotic it is to describe a LLM as a "Python loop", what's amazing with those people, who I'm sure would self-describe as materialists, is their assumption that a computer program couldn't possibly be intelligent, which I think reveals how deeply entrenched some inchoate form of dualism remains.
I'm sorry to break it to you, but the human brain isn't so special. At the end of the day, it's just a wet computer interacting with a body, both of which are the result of a very dumb process that went on for billions of years, the last few hundred millions of which it spent selecting your ancestors that were the best at fucking.
The implicit assumption behind this kind of argument is that, unless intelligence isn't implemented exactly as in the human brain, it's not really intelligence. But in general there are several ways to implement something and there is no reason to assume that human-like intelligence is any different.
Arthur 👁 retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: docs.google.com/document/d/e…
Arthur 👁 retweeted
Some people have convinced themselves that AI doesn't work and the AI industry is about to implode, and that AI companies calling for a deliberate slowdown of the R&D that represents their only competitive edge is somehow them begging for an anticompetitive bailout.
IDIOCY.
Arthur 👁 retweeted
"Why are AI labs suddenly talking about a slowdown now?" <-- They were talking about it for years but you weren't paying attention until this week
Arthur 👁 retweeted
i do think a lot of people on the pro-open-source side are having a bit of a knee-jerk reaction to the pacing statements today, as we're used to viewing the closed labs as power-seeking.
but i think their hands are somewhat forced here, and this is just another chapter on the fairly inevitable path towards decently-fast decently-safe decently-commoditized intelligence abundance.
ask any F500 exec or swing voter. the world doesn't really want super-fast-takeoff superintelligence owned only by two companies, and thus we won't get it. it'll happen at the pace that the world can accommodate it, which means reaching some sort of confidence consensus that the models are aligned enough that we won't be dealing with scary new incidents all the time.
this will trickle out broadly, in the form of best practices and distillation. the smarter a model is, the more it has a "personality", and the less effective strict rules are. there will be awkward compromises and moral tensions. but we ultimately just want the models to be reasonable, and to do the sorts of things reasonable humans would do if our brains were faster and less error-prone and had more working memory. i think we'll get there.
the labs will build mac and windows, the rest of us are building linux. everyone's gonna do great. weird stuff will keep happening, but we'll still wake up and go to work, until the work does itself in a manner the world finds acceptable.
Arthur 👁 retweeted
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
All of this because @openai decided they would EXPLICITLY ask thousands of software instances to find security bugs on the open web — while cos-playing a humans🤦
Best Supporting actor goes to @dwarkesh_sp for his soap opera recap of OpenAI’s hacking speed run
Here’s another concept: don’t create and scale thousands of viruses and worms and be SHOCKED when they actually cause damage!
What a fucking farce this is.
Arthur 👁 retweeted
I can think of no greater sign that one's brain is utterly rotted by tribal thinking and extreme political ideology than their reflexively reacting with "regulatory capture!!" at the slightest announcement from the labs, even when it's about the precise opposite
Arthur 👁 retweeted
Can we all agree that “frontier labs are lying about AI safety to pump their IPOs, capture the regulators, and make more money” is a conspiracy theory at this point?
For it to be true, Sam, Dario, Demis, and Elon all have to be lying, along with a big chunk of their execs, chief scientists, and resigning employees.
It also asks you to believe the labs volunteered for much heavier scrutiny that slows them down in the middle of the biggest ARR growth run in the history of capitalism with no signs of slowing down … because they want to make more money by inventing regulatory hurdles that only they can clear(?).
And OpenAI has now said it would even be willing to delay its own IPO over safety concerns - i.e. put off (maybe risk entirely) the biggest liquidity event ever in Silicon Valley. The conspiracy requires us to believe that they’re doing that to make more money(?)
Elon agrees with Dario; i.e. the guy who ran D.O.G.E is *asking for more scrutiny and regulation*. And we're supposed to believe that's because he thinks it will help his company succeed(?)
A much simpler explanation at this point: they think the risk is real.
Arthur 👁 retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Arthur 👁 retweeted
Replying to @gabriel1
Actually, it feels like Humans will become obsolete/irrelevant to serving the World
So instead of asking: "What do other people want, that I can build?"
So now's the best time to ask:
"What do I actually want?"
Then just ask AI to build it.
What the fuck okay
Replying to @AnthropicAI
The symmetric cipher is a reduced version of the Advanced Encryption Standard (AES)—which has received decades of scrutiny (more than almost any other encryption algorithm).
In a week, Mythos Preview found a way to speed up an attack on this version of AES by 200-800×.