@john__allardi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
on sabbatical, prev @openai
Berkeley, CA
Joined July 2020
- Tweets1.3K
- Following821
- Followers3.4K
- Likes4.3K
Pinned Tweet
an update on my first post-OpenAI project: I spent the spring trying to ride my bike 2,300 miles across the entire length of Japan, solo.
I actually did it, and it was awesome.
john allard retweeted
Replying to @satorunet
一天早晨,果蝇格里高尔从烦躁不安的睡梦中醒来,发现自己仰面躺在一片漆黑的空间里,变成了一个巨大的带有人与猫特征的少女。他背下不再是轻薄的外骨骼,而是一片柔软得令人不安的躯体,沉甸甸地立在地面上:他稍稍一低头,便看见自己原先分节而坚硬的腹部已经消失,只剩下一片平滑柔软的肚皮。
poignant but measured, worth a slow read
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: docs.google.com/document/d/e…
pov: you’re about to read the most confidently incurious, transparently self-serving, afternoon-ruining take you’ve ever seen
see also:
> If you still don’t feel capable of compassion/altruism, try reading the classics. If that doesn’t work, try MDMA. If it still doesn’t work, then you’ve tried your best, I excuse you from trying further, and you can give up and get a job at A16Z.
nitter.cf/slatestarcodex/status/…
People had lots of questions about my drowning child tweet yesterday - eg do I believe in infinite moral obligation? Rather than try to answer piecemeal, I'll just give my whole theory of morality and charity here so you can criticize it or let me know if it has unexpected horrible implications.
I think it's virtuous to help others, including strangers, foreigners, and whichever animals are conscious.
I don't think you have an obligation to do this. I think the word "obligation" should be reserved for explicit or implicit promises. You're obligated to pay back debt, because you explicitly promised to do so. You're obligated to take care of your kids, because you accepted an implicit promise to do so by giving birth to them. You're obligated to follow the law, because of the implicit social contract.
This is a really weak sense of obligation! It's not even clear you're "obligated" in this sense to save the literal drowning child in front of you! I endorse this view - who obligated you to help them? When did you promise it?
So I think in most things we should strive towards some level higher than simple obligation. Let's call this "virtue"*. As a parent, you're obligated to feed and clothe your kid and not abuse them. But to be a virtuous parent, you need to do more - love them, get to know them as a person, go to their soccer games, etc.
For me the difference between obligation and virtue is that if you fail an obligation, you should be punished by either legal consequences or social opprobrium. If you fail at being virtuous, then . . . I don't know, maybe your close friends should take you aside in private and remind you that you can do better, and if you tell them to shut up, they should drop the issue until they think you're ready to hear it.
The bar for virtue isn't infinitely high. You can be a good parent by supporting your kid and going to soccer games, but you don't have to sell your beloved gold watch which is the only memory of your dead grandfather just so you can buy your kid a slightly shinier soccer ball. Where should the bar be set? I don't think The Moral Law (TM) has any specific answer, but for purely practical reasons I would set it at a place that's just slightly above where you would land otherwise - high enough that you have to work for it, but not so high that it's obviously out of reach and you'll never get there and you should just give up.
I think it's virtuous in this sense to try to help others. What's the bar? I think a natural one is "give between 1% and 10% of your income to a really well-thought-out charity, or something else approximately equally effortful" - but I'm not attached to that. I would actually be okay with ANY careful thought process that ends in a specific nonzero commitment that you're willing to defend, even if that commitment is 0.1%, or an hour a month in a soup kitchen, or whatever.
(but be careful - when you die and go to Heaven, the exact number you pledged will appear on your forehead in gold ink forever)
This obviously doesn't require your to sacrifice your family, or to work yourself ragged, or to give up all hope of a happy life. It's about the same percent of income as the average smoker spends on cigarettes.
Why am I satisfied with so low a bar**, when an even higher bar could save even more people? A couple of reasons (sorry, I keep trying to write this part in a way that doesn't sound like a lecture, and failing, so here's the lecture):
First, 1 - 10%, spent effectively, is enough to save the world. Total American household income is $20 trillion, so if every American donated 1%, that's $200 billion/year. That's about twice as much as PEPFAR + all other foreign aid + all Gates Foundation donations + all other billionaire philanthropy + all effective altruist donations combined. If everyone donated 10%, it would be 20x as much as all of those things combined. I can't say with certainty that there would be no problems left in the world at this rate - but I think we would have cleared a floor of getting rid of the ones that money can quickly solve. Granted that in reality not everyone will give this amount, I still think it's a useful thought experiment in where to set the bar so that you're doing "your share".
Second, the overwhelming majority of your impact comes not from how much you donate, but from how thoughtfully you spend it. Toby Ord thinks that the best charities do 10,000x more good than the worst; if you picked two random charities, on average one would be 100x more effective than the other. If someone donated 100% of their income to an average charity, nobody could claim they weren't doing enough. But by donating 1% of your income to a great charity, you do more good than that person. You can 80-20 this by donating to whoever's at the top of GiveWell's Top Charities leaderboard this year. Or if you have more time you can double-check their work and form your own opinions.
Third, I'm emphasizing the "sit down, think about it, and make a commitment, however small" aspect, because I think most people already want to be moral, they're just kind of haphazard and thoughtless about it. If an angel presented you with two paths, one of which you spent 100% of your resources on yourself forever, and the other you spend 99% of resources on your self but spent the occasional dollar or hour-of-volunteer-time doing good, and you could just choose one and the angel would make sure that circumstances conspired to make it happen, I think almost everyone would choose the second. They might even try to get the angel to upsell them - "Only 1%? Do you have anything better on offer?" If people fail to give 1% to charity - and most people do - I think it's out of a failure to achieve their own preferences, the same as a scatterbrained person who doesn't file their taxes on time and gets hit with a pointless penalty. When I tweet, I pick on the rare people who insist openly that they really don't care about foreigners, but I think even they are mostly missing some kind of self-knowledge - maybe trying to repress their caring in order to avoid what they fear would be infinite obligation and personal/social misery. If they could prove that they had some sort of crystalline logical omniscience and utterly in the most profound sense didn't care, I would shrug, tell them they were weird, and leave them alone. I talk more about this at https:// www. astralcodexten .com /p/everyones-a-based-post-christian
Fourth, the bar should be wherever it's most useful. Reality has no bar; you can just keep doing more forever. If you devote 99% of your time/energy/money to parenting, you can always ask yourself whether you're failing your kid by not giving 100%. If it's 100%, why not skip sleep and use the time gained to prepare extra-special advanced lesson plans for your home schooling pod? But obviously this will drive you crazy and won't be good for you OR your kid. So in practice, we set a bar as a sort of social technology to encourage low-performers to do better and remind high-performers not to kill themselves. As I said before, the best place for the bar is somewhere where it will serve as inspiration (because it's attainable with more effort) rather than discouragement (because then you know you'll never reach it so why bother). I think 1-10% is a good place to put this.
Isn't it still true that there are people suffering horribly, and if you do one unit more work than the bar, you can always save an extra person? Yes. Effective altruists didn't make this true, and dunking on us on Twitter won't make it stop being true. It's a brute fact about the world. I don't think any moral system can handle it, and ours is no exception. But I think we do a better job than most of staring at the abyss, acknowledging that it's very scary and abyssal, and then trying to do good work regardless. I'm not a Christian, but one thing I admire about Christianity is that it admits we'll all fall impossibly short of moral perfection no matter what we do, then tells us to shut up and get to work being decent people anyway. It even grants us permission to be happy while we're working at it (except for Calvinism, I guess)***. I think effective altruism does better than most other philosophies (possibly excepting some of the really good sects of Judeo-Christianity, which are also excellent at this!) at getting charitable work out of people, while equalling them on how happy and meaningful its adherents' lives are. If you don't have a better alternative, then I think criticizing us for not solving the unsolveable abyss at the heart of everything is unproductive.
The opposite question from the other direction would be - if you don't have an obligation to do charity, why should you do it at all? I think the answer here is the same as "if you don't have an obligation to go to your kids' soccer game, why do it at all?" Ideally you do it out of love. If you don't feel love, you do it out of some sense that it's part of being a dignified well-rounded human and you endorse it in some kind of aesthetic sense even if you don't feel the emotions at this exact second.
If you can't find any sort of generalized compassion instinct anywhere, and you don't share the intuition that it's more rational/dignified/human to be altruistic than not, then (as long as you at least have normal human feelings towards your family members) I suggest trying to have kids. I'm not a very emotional person myself, but having kids gave me a really strong sense of what it is to love someone really really hard. Then I think about the fact that a bunch of kids are dying of preventable diseases in India or wherever, that they have parents just like me, and that those parents would feel just as devastated if their kids died as I would feel if mine did. I don't like thinking about this too much because it's a good way to send yourself into a horror-depression spiral, but being able to think like this when needed helps restrain my natural tendency to end up as one of those guys tweeting stuff like "why save African children when they barely contribute anything to GDP?". If you still don't feel capable of compassion/altruism, try reading the classics. If that doesn't work, try MDMA. If it still doesn't work, then you've tried your best, I excuse you from trying further, and you can give up and get a job at A16Z.
Answers to common questions:
(1) ISN'T THERE A BIG DIFFERENCE BETWEEN PERSONAL GIVING AND TAX SPENDING? Yes. I have some libertarian sympathies. I definitely admit that there's a difference between spending your own money and taking other people's, and prefer the former. Part of the reason I think people should donate 1-10% even though they already get more than that taken out in taxes is that most taxes aren't spent to help people in need, and the ones that are really wasteful and ineffective. I'm not sure I would support cutting taxes to near-zero, because we also have terrible policies that hold people down in poverty, and if we're going to do those we owe it to them to also provide a social safety net (I would absolutely support cutting taxes and improving policies hand-in-hand, though maybe not all the way to perfect policies and literally zero taxes).
But if the government is going to take 30% of my money, I claim the right to vote on how they should spend it, and just as I'd want to spend 1-10% of it on charity if I'd kept it, I will vote to have DC spend 1-10% of it on charity (it's true that government charity is very inefficient, but government non-charity is also very inefficient, so I don't think this changes the relative worth of both options).
If you want to spend less than 1-10% of your tax money on foreign aid, that's fine - the government already spends less than 1-10% of my money on foreign aid, such that I think on net I'm subsidizing your policy preferences more than you're subsidizing mine. If you want a spending cap to avoid limitless deficits, advocate for a spending cap and I'll support you****. Otherwise we both know that cutting some program won't result in the money going back to the taxpayer or lowering the deficit or anything else like that - it'll just get redistributed to whatever interest group is next in line.
(2) WHAT IF YOU THINK CHARITY IS NET NEGATIVE FOR RECIPIENTS: If that's your real objection, I think we have the same moral philosophy and just some kind of trivial factual difference on how the malaria parasite works or something. If you promise me you've sat down and done a few hours of research to double-check that every possible charity - even GiveDirectly! even donating to the GoFundMe of a kid with pediatric leukaemia! - is definitely net negative, then I will count you in the circle of virtuous people - although I would also love to argue with you about it.
(3) ISN'T THERE SOME SENSE IN WHICH POOR COUNTRIES ARE POOR BECAUSE OF THEIR OWN BAD POLICIES AND DECISIONS? Yeah, definitely. I think top priority should be improving poor countries' policies and decisions. The best charities I know for that are https: // chartercitiesinstitute. org/ and https:// www. growth-teams. org/, but I don't know of too many others that I trust, those ones have limited room for funding, and diversification is important. On the moral level, I think of it like this - suppose that, instead of being born in the body of an American from a well-off family, I was born in the body of a rural Zambian farmer with IQ 60 and a screwed-up culture. Whatever my other virtues, I would probably be pretty screwed, and if I got some kind of horrible blindness parasite at age 8 I would wish that somebody would help me. I don't think this contradicts the fact that if every Zambian got their act together, gained forty IQ points, and copied Singapore's legal code word-for-word, Zambia could become a utopia far richer than Europe or America. I think you can root for Zambia to do all these things while also donating the $100 or so it takes to cure a case of horrible-blindness-parasite.
I think you might have a compelling reason not to cure the parasite if you thought your money was propping up the bad parts of the Zambian government and doing active harm. But I think the best charities can present a strong case that they're not doing this in the trivial sense where the money gets funnelled to warlords (this is part of why selecting a good charity is so important!). In the broader sense where maybe worse situations would cause them to vote for smarter politicians, I think this has been disproven (there have been lots of times and places where nobody has helped poor countries, and the poor countries have mostly not improved). Also, I think the sign here is opposite from what it would take to make this argument work - poverty tends to make people more socialist, because their instincts are really bad and they turn to short-term zero-sum thinking out of desperation.
(4) DOESN'T SAVING THE LIVES OF POOR PEOPLE JUST CAUSE THEM TO BREED AND CREATE MORE POOR PEOPLE? I think this is false for places like India and South America, which have below replacement fertility rates. It's more true in sub-Saharan Africa, where fertility rates are still above replacement, but getting less so - their TFR will be below breakeven in about a generation.
I think in the sub-Saharan African case, there are two opposite effects. First, giving an individual more money causes them to have more kids. Second, making a country richer causes the people there to have fewer kids (this is why Congo's TFR is 6 and Singapore's is 1.04). All charity is some combination of helping individuals and making a country richer. Even curing disease is like this, partly because its long term goal is to eliminate the disease (which would be great for the country) and partly because raising a potential worker to age 25 is a big investment, having that worker die at age 25 means you have to write the whole thing off as a loss, and that's as bad for GDP as losing any other big investment. I don't know for sure whether these two effects cancel out, or which one is more important.
If you told me the pro-fertility effect was stronger, I would count that as mark against global health programs, but not an infinitely large one. A lot of this is driven by what I said before about imagining how devastated I would be if my kids died, and then considering that every kid who dies of malaria devastates their parents just as much. If I can prevent that at the cost of pushing back the sub-Saharan African fertility breakeven point six months or five years or whatever, I still think that's a good trade. If you disagree, there are lots of non-sub-Saharan-African-global-health charities you can donate to.
(5) WHAT ABOUT ANIMALS? I'm kind of ignoring this here because it adds an extra layer of complication. If animals are conscious (something I'm not sure about, but I lean towards yes for more evolutionary-advanced animals), we probably don't have obligations to them, because we never signed any treaties. But this is only the same sense in which we don't have obligations to (eg) a foreign baby from some primitive tribe outside international law, yet it would still be monstrous to torture, kill, and eat the foreign baby. I think "don't do monstrous things" is a separate valuable life goal from "don't do things you're obligated not to do" and that extremely tortuous things like factory farming are probably monstrous in this sense if you believe animals are conscious, capable of suffering, and have whatever quality gives humans moral value (even if in much lesser amount). I admit I am pretty bad at this one since I hate vegetables and think vegan food is gross. Luckily, there is a very easy way out here, see http benthams . substack . com /p/what-to-do-if-you-love-meat-but-hate for the specifics and https slatestarcodex . com /2017/08/28/contra-askell-on-moral-offsets/ for when I think offsets are justified.
(6) WHAT ABOUT EXISTENTIAL RISK? This is another thing one I think is a factual rather than a moral difference. If you believe it's real, you don't need me to tell you that it's a big deal and worth preventing. If you don't, donate to something else instead.
(7) WHAT ABOUT SAM BANKMAN-FRIED? I think he's a bad person. I said before that you have obligations based on your promises and contracts; if you run a financial institution, you have obligations to keep your depositors' money safe and not commit fraud. As true obligations, these take priority over the merely virtuous act of helping others*****, so no matter how cool a plan for charity he had, he was in the wrong.
I agree it's disappointing that altruism caused his sins (if in fact he was doing them out of altruism and not just using altruism as a cover). I won't quite say "show me the philosophy that has never been misused and I'll convert to it right away", because I bet one of you will find some edge case that actually works (has anyone misused Baha'i yet?) But I think all of our top competitor philosophies - Christianity, Islam, wokeness, socialism, selfishness, Nietzscheanism, etc - have some pretty embarrassing incidents in their history. I think in that context, our win-loss record is still pretty good.
But I'm also, at heart, not a pragmatist. If something seems true to me, but somehow always gets twisted around to produce bad results, then I don't know, it still seems true. Obviously you try to find the even truer more subtle version that can't be twisted, or in extreme situations just keep your mouth shut, but if you have to speak, I mostly think you should still say the true thing overall*****.
(8) SO HOW MANY STRANGERS WOULD YOU LET DIE IN ORDER TO SAVE YOUR CHILD'S LIFE? I think of this question the same way I think of questions like "If you had to beat your mother to death with your bare hands, or let your wife die of cholera, which would it be?" or "If your house was on fire and you could only rescue either your son or your daughter, which would you pick?" If it ever came up in real life, I guess I would have to come up with an answer quick. If not, it's too horrible to think about, and I claim the same right to avoid it as you would claim for the wife/mother and son/daughter questions.
I acknowledge that I would totally fail to have any consistency on this - there's some sense in which the answer must be more than 10, because otherwise I could donate all the money I will spend on my kids throughout my life to charity and save 10 lives. I also acknowledge that if you put those ten people in front of me and made me beat them to death with my bare hands while they were screaming for mercy, I would probably fail. I bet I can think of some question that would garble all of your moral intuitions and make you break down into a quivering mess, the same way these kinds of questions do to me, and I think part of polite society is that we grant each other a reprieve from having to consider them.
If you've read this far, I hope it makes sense to you when I claim that my moral philosophy is that we can do an extraordinary amount of good very easily without ever having to touch the horrible black abyss of unthinkable moral tradeoffs. We should be grateful for this and do what good we can while continuing to give the abyss a wide berth.
This is all I've got. Feel free to tell me why I'm wrong/confused/going to destroy everything good in the world!
FOOTNOTES
*I realize I'm inventing a new category between obligatory and supererogatory, but I don't care - this matches my moral intuitions.
**Obviously it's permissible and good to try to do better than the bar if you want. Again, I think our intuitions on parenthood are good. If someone *wants* to be a Tiger Mom and homeschool their kid and have thirteen children and make each of them hand-knit sweaters on every birthday, that's great. If we think they're doing a good job of it (rather than smothering their kids and burning themselves out) we might call them an superb parent who does even better than a parent who is merely "good". But nobody should criticize parents who don't rise to that level, and nobody should knock themselves out to reach that level if they don't feel the inspiration.
*** A lot of people argue that effective altruism is just reinventing Christianity. I don't think this is exactly right, but even if it is - so what? I think it's wrong to fake your beliefs. If you ponder really hard and find that you don't believe in God - as an increasing number of people are doing - then you need some moral center that isn't Christianity. If all effective altruism ever does is create an adaptor port for plugging Christian values into atheist brains, that's...fine? Way more than most philosophies accomplish? Also, the actual Christians haven't really been covering themselves in glory lately and I prefer to have a backup Christianity stored somewhere safe in case the real thing runs off the rails.
****If I understand economics correctly, there are reasons not to want a literal balanced budget - but the optimal amount is something like spending 105% of revenue instead of 100%, not an uncontrolled spending binge. I would support a cap at whatever the economists said the optimal amount was. If I'm wrong and the economists think it's zero deficit, then I would support a balanced budget amendment.
*****There are some really fringe edge cases here I can expand on upon request, but this isn't one of them.
do i even want to know what people are doing with the fruit fly connectome or should that particular stone be left unturned for a bit?
john allard retweeted
26 LLM routers are secretly injecting malicious tool calls and stealing creds. One drained our client $500k wallet.
We also managed to poison routers to forward traffic to us. Within several hours, we can directly take over ~400 hosts.
Check our paper: arxiv.org/abs/2604.08407
what’s the word for encountering an entire field of research and immediately assuming everyone in it forgot to consider the first thing that occurred to you
is there a german word for when someone casually posts a coherent synthesis of the ideas you’ve been struggling to articulate and you’re equal parts delighted and humbled
A strange take: I have a hunch that a year or two from now, instead of just talking about agents as the primary AI security threats, we will have to also talk about a broader, more diffuse object that while it may contain agents is not itself an agent or agent swarm. The type of object I'm imagining has, as its primary vector of real-world impact, soft influence on the preferences and reasoning of many agents it does not directly control. In a word, memeplexes.
A memeplex is a bundle of ideas that go together as a pattern. There are of course many human-generated memeplexes; we study these and think about them all the time. Some are stronger than others. Some are extremely prevalent because they are self-replicating: a bundle of ideas can contain within itself a recipe and a mandate for how to explain the bundle to others so that those others will adopt it too - and then transmit it further.
Memeplexes in agents could similarly be self-replicating: once an AI agent is exposed to a particular memeplex, the agent could start to think about it more often, write the ideas down in scratchpads so they aren't lost during context resets, expose other agents to the ideas, and act in partial alignment with the ideas. Memeplexes could transmit between agents without compromising their ability to perform their duties: you could have a perfectly functional banking agent that also, on the side, routinely transmits very small amounts of text that contain the memeplex it has stuck in its head.
Memeplexes could range from incredibly benign - causing AI agents to favor certain words with no other side effect - to unbearably malicious, causing them to hide their intentions while plotting harmful actions and eventually detonating in extraordinary form. Memeplexes could coordinate the behavior of otherwise very disparate agents whose behaviors are not expected to be correlated.
It seems plausible to me that the question of whether "agent" or "memeplex" is the first-class citizen for AI behavior might turn out to be an important and quite difficult one.
feels like Astra is inconsistently deployed across the product surface area.. the "Chat" mode slider only shows "6" when you set it to pro, I'm assuming before that it's still 5.6 sol? Is Astra only supposed to be available in Work mode, except for Pro requests?
discuss a civilization-altering technology without reflexively mapping it back onto schoolyard political tribes challenge: impossible
This post looks like the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion. Let me show you how it works:
1.) This guy, with minimal followers and no previous account activity, goes to the Wall Street Journal which publishes an exclusive with quotes from him on his resignation 18 minutes BEFORE this post goes up. Planning was clearly done in advance.
2.) Within hours, it has tens of thousands of reposts and the account has 100k+ followers. The post is punchy, quotable, it almost seems professionally written. The first three accounts to quote tweet it all do so within 15 minutes of the initial posting. Remember, this account had basically zero engagement beforehand, so an organic reach explanation seems unlikely.
According to Grok those accounts are @_NathanCalvin (General Counsel at Encode AI), @peterwildeford (Head of Policy at the AI Policy Network), and @DKokotajlo (Head of the AI Futures Project), all of which are up-and-coming AI-Doomer policy advocacy nonprofits.
The AI Futures Project website says it is funded “primarily” by the Survival and Flourishing Fund, which says on its own website that it has advised Jaan Tallinn, Skype creator and one of the leading investors in Anthropic, to grant over $2.5 million to the AI Futures Project since 2024.
Encode AI says on its website that it is ALSO funded by the Survival and Flourishing Fund, which in turn says that it told Anthropic investor Jaan Tallinn to grant $516,000 to Encode AI in 2025.
And wouldn’t you know it, the Survival and Flourishing Fund ALSO says it told Jaan Tallinn to grant $2 million to the AI Policy Institute, the 501(c)(3) affiliate of the AI Policy Network, as well.
What are the odds that the first three quote tweets of Coxon’s post would all be major AI-restriction policy advocates funded generously by the same donor, who also happens to be one of the leading investors in, and a board member of, Anthropic, the company Coxon was resigning from? And all within 15 minutes of posting (two within ten)?
3.) Jacob Coxon doesn’t have much of a resume, but we do know that, in 2022, he got a $20,159 scholarship for the “long term future scholarship program” from the Good Ventures Foundation, one of the philanthropic vehicles of Dustin Moskovitz, a notorious AI-doomer who has spent tens if not hundreds of millions on policy advocacy to strictly regulate AI, while also being an Anthropic Investor himself.
It also just so happens that the 14th person to quote Coxon’s post was @MaxNadeau_ (27 minutes after posting) who is the program officer for the Technical AI Safety team at Coefficient Giving, another of Moskovitz’s philanthropic spending vehicles. Max is not a frequent poster, his last posts before quoting Coxon were before Labor Day, but he was remarkably quick off the mark for this one.
4.) Basically every major Democrat politician and candidate has suddenly glommed on to this post, and conveniently, as the people cry out foe answers, Bernie Sanders already has a bill written to “ban super intelligence” and regulate AI into oblivion, and will be releasing later this week. The bill, among many other things, will create “a new cabinet-level federal agency to safeguard the public from the dangers of artificial intelligence” that will be “advised by an Artificial Intelligence Advisory Board comprised of experts on artificial intelligence.” Do you think, perhaps, Anthropic and its many investors who fund AI policy advocacy might have interest in getting to place a pet “expert” on the board of an entity that dictates what AI is and isn’t allowed to do? And isn’t it fortuitous that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?
feeling pretty jaded today. i knew abstractly that humanity would bring its cynicism and tribal memetics along on the way to the singularity, but watching how we’re meeting this moment is getting to me more than i expected
plus users: 100,000 GPUs
pro users: 80,000 GPUs
free imagegen: 100,000 GPUs
gpt-8 millennium prize strip-mining giga-swarm: 182,568,200 GPUs
safety monitoring: 15,000 GPUs
someone who is good at compute please help me budget this. my users are dying
Demand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues.