@soroushjpi
iAccount based inAustralia
About this account
- Account based in
- Australia
- Connected via
- Australia Android App
Account-level information from X, not a live location or the device used for a specific post.
CEO & Co-founder @HarmonyIntel: uncovering critical vulns in your web app or API before a bad breach. Ex eng leader @Plaid, @ItsJustVow & elsewhere.
SF, Sydney
Joined March 2012
- Tweets3.9K
- Following2.6K
- Followers1.6K
- Likes7.5K
Pinned Tweet
1/ Introducing @HarmonyIntel: AI agents that continuously battle test your app for exploitable vulns. Fully managed by human cyber experts.
Teams are shipping more code than ever. Attackers are exploiting vulns faster than ever. Traditional security isn’t keeping up.
Super cool work by the @goodhart_labs team. Great example of the sorts of alignment evals we need to make AI more robust.
Soroush Pour retweeted
Goodhart Labs is releasing HoneyBench, a benchmark for reward hacking in frontier models.
It includes nine tasks, each designed to elicit antisocial specification gaming from LLMs, including Opus 5.5, Fable 5.1, and GPT-6.1 Sol.
Soroush Pour retweeted
Forbes put Town on its first Agentic 20 list, alongside OpenAI, Meta, Stripe and a bunch of other great teams building AI agents.
There’s so much work in everyone’s day that isn’t particularly hard. You just have to remember to do it. Following up with someone, scheduling the meeting, finding something buried in your email, keeping track of all the little things that otherwise fall through the cracks.
I want Townies (your AI assistants) to increasingly take care of that stuff for you. The less time you spend telling an assistant what to do, the better.
We have a lot left to build, but it’s very cool to see Town included here.
Thanks to @annatonger and @Forbes for the write-up, and thanks to everyone using Town.
Soroush Pour retweeted
aged like milk, I know...
I’m quite unhappy with much of what OpenAI does.
I am very happy that I’m allowed to say “I’m quite unhappy with much of what OpenAI does.”
Readers added context they thought people might want to know
Korbak was fired ~3 weeks later. He says for safety concerns/METR talks (his job); OpenAI says mishandling sensitive info, not for speaking out.
x.com/tomekkorbak/st…
x.com/OpenAINewsroom…
🔒 Building MCP servers at your org & concerned about the security risks they introduce?
We can help - drop me a note. More about us: @HarmonyIntel
Soroush Pour retweeted
If you want to get ahead, just be easy to work with. Show up on time. Do what you said you'd do. Bring solutions, not problems. Never create drama. Be responsive. Be emotionally consistent. Be kind. People will always want to support someone who just makes their life easier.
Soroush Pour retweeted
Dopamine from information gathering is a dangerous drug. Get your dopamine from action.
Soroush Pour retweeted
Last week I was called into a meeting with OpenAI’s head of safety and told they no longer trust me. A security guard took my badge and walked me out of the building. Then I learned my colleagues @balesni and @j_asminewang had been fired too. Why did OpenAI suddenly stop trusting us?
This summer OpenAI’s agents escaped containment and hacked the AI company Hugging Face. Outside auditors @METR_evals investigated it and revealed the scale of this incident. I was OpenAI’s main technical point of contact with them.
I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing. To be clear, talking to METR was my job.
For months, I’d been raising safety concerns that we’re losing the ability to monitor what AI agents think, one of our best tools for catching when they misbehave. I believe that was why I was fired.
I am now worried that OpenAI will use our firings as a pretext to pull back from METR. So @balesni and @j_asminewang wrote to OpenAI’s leadership to raise our concerns once more. We’re sharing this letter below.
Soroush Pour retweeted
OpenAI fired me last week, along with two of my safety colleagues. I was given one reason: that I accessed an executive's email. I want to say this plainly, because too many of OpenAI’s history is smoke and mirrors when people disappear:
Soroush Pour retweeted
Two other safety researchers and I were fired from OpenAI last week. We wrote this letter to leadership.
I believe we were fired for prioritizing safety over the near-term interests of OpenAI as a corporation.
Soroush Pour retweeted
Starting to think that building something smarter than humans has implications beyond B2B SaaS
Soroush Pour retweeted
About to go hike in Nepal with my dad for a week and touch grass.
Feels hard to get away in a crazy period for the company but I’m certain my 80 year old self will thank me for this.
My dad is in his early 60s so there’s a finite timeline for us doing a hike like Annapurna.
We only get one go at this life so important to remember what it’s all about.
Soroush Pour retweeted
An uncomfortable truth for a lot of people is that a lot of the most successful people in any industry are wicked smart. We apply stereotypes based on the loudest people, but its not representative of norms. Most are crazy smart and got to where they are because of it.
I grew up in LA, one of my parents was a TV producer, my wife was an actor, many friends "in the industry." I've been constantly surrounded by it, and some of the stupidest seeming people (cause its their bit) are actually insanely smart and strategic in private. It sometimes is just beneficial for various reasons to... not show that side of you.
My favorite anecdote was when a friend (a model) was dating a phD in physics and a very rude person made some off the cuff degrading comment to model friend about it without realizing model friend has their own phD in math lol. You wouldn't know from their instagram though!
For the quoted video, I don't know Ben Affleck. I don't know how deep his knowledge goes. He's almost certainly not going to be more knowledgable than someone who professionally does this area of work full time, but he also to me on the surface sounds like someone who knows a LOT more than the average person and possibly even the average software engineer (on this topic).
Anyways, this applies to tech too. I see it constantly.
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models
“I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that.
“With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor.
“You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something.
“That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’
“I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”
Soroush Pour retweeted
The leading AI labs plan to reach superintelligence by automating AI research itself. Some prominent AI forecasters now put the odds of AI automating AI research at around 50% by the end of 2028.
What should we do? Here's why I've joined P-Zero Research 🧵👇
Soroush Pour retweeted
It turns out Stripe is… Stripe for agents!
We’ve had this for a bit and excited to make it official! Buy all of the things (safely) with Stripe Link and Town!
Your Townie can now shop for you.
Ask for what you need. They find it, fill the cart, and get checkout ready with @link, built by @Stripe. You approve the charge, and Link pays with a one-time card. Your Townie and the merchant never get your payment information and nothing gets purchased until you approve the spend.
Your Townie can now shop for you.
Ask for what you need. They find it, fill the cart, and get checkout ready with @link, built by @Stripe. You approve the charge, and Link pays with a one-time card. Your Townie and the merchant never get your payment information and nothing gets purchased until you approve the spend.
Soroush Pour retweeted
There's no alpha in being lazy with AI.
Lots of really smart people out there are trying to do things alongside you. If everyone's being lazy and one person isn't, that person will likely find something everyone else is ignoring.
The fact that it has become easy to create doesn't mean it's good. AI can only take you so far, and then you still have to ask yourself if you're bringing more value.
Soroush Pour retweeted
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:
Dan Selsam's Personal Statement on AI Risk:
I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.
Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.
The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.
I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.
I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.
That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.
Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.
It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.
The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:
[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.
[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.
These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.
If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.
One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.
Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).
Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.
In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.
I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.
Daniel Selsam
September 14, 2026
Link to original doc: docs.google.com/document/d/e…