@HatforceSec

Sharps or Squares. Personal views only.

Simpli-city
Joined July 2011
AI for Security has never been more exciting. Let me present MAPTA, our multi-agent framework that found multiple (now confirmed!) Remote Code Executions (RCE's) in flagship web products of Tier-1 companies. Why the secrecy? We're good boys, letting them cook patched through responsible disclosure. What's our secret sauce? 1/n
17
109
15
634
147,907
Arthur Gervais retweeted
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field. I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so. First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world. The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability. Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended. I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent. Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.) Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration. Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building. [Original text (with links): deeplearning.ai/the-batch/is… ]
193
379
84
1,416
151,665
Arthur Gervais retweeted
Warren Buffett: “Gerçek zenginlik, en lüks evde oturup en pahalı arabaya binmek değildir. Asıl zenginlik, kafan rahat yaşayacak kadar kazanırken vaktini nasıl harcayacağına, kiminle çalışacağına ve hayatını nasıl kuracağına kendin karar verebilmek. Sonuçta insanın sahip olabileceği en büyük zenginlik, zaman ve özgürlüktür.”
112
2,105
151
9,585
412,314
"""there’s a massive gap between an interesting research and a useful product""" Many people miss that key distinction
poor guy claim to have built Jev a year ago but no one cared, and now Jev stole all the thunder many people are saying “you gotta tell your story” or “marketing is important”, and they just completely missed what actually made the difference here i just looked into this laya model laya.convaiinnovations.com/ and: - it only supports 512-1k context… a lot of use cases won’t fit at all - evaluating the model directly shows its accuracy is as good as a coin flip. in order to get good results, you need to first fine tune it i’m sorry, but that’s not Jev there’s a massive gap between an interesting research and a useful product you can “tell your story” all you like, but you can’t blame Jev for stealing your thunder when Jev did all the work to make a well-packaged solution anyone can just grab and go Jev is not completely new from an academic sense, just like how ChatGPT was not the first LLM don’t underestimate the effort and value in putting together something that’s actually good enough for adoption - it makes all the difference
2
367
Nah, he should just say: give me real entropy
omg I love the term "slop grenade" it's perfect thx @tobi
200
multi-threaded classification model
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
264
oh the bunny is better than everyone in the US ??
gg. we'll take the number 1 spot on @Hacker0x01's US leaderboard with our autonomous pentester. more than number 2, 3 and 4 combined.
7
4
1
147
19,528
Arthur Gervais retweeted
TESLIN FSD Sistem lahko prevzame krmiljenje, vožnjo skozi križišča, menjavo voznega pasu in številne druge manevre, pri tem pa voznik ves čas ohranja nadzor. Moj vtis? Sistem je varen, učinkovit in predvsem zanimiv pogled v prihodnost mobilnosti. Ne utrudi se, ne zaspi in lahko pomembno prispeva k varnejši vožnji. Slovenija je med prvimi 6 državami v Evropi, ki je odprla vrata tej tehnologiji. Gremo na naslednji korak – robotske taksije. 🇸🇮
117
155
57
1,422
560,760
codex schedule is running my eightsleep
my WHOOP subscription expired so i had codex reverse engineer it and build my own app - telemetry goes to my prometheus server - grafana and the app visualize it - hermes monitors metrics and plans workouts next wanna build my own Zwift and Strava the era of personal software is here
337
not many things i hate, but slow server CPUs
145
Nothing will shutdown open source
Nothing can shut down open source
2
304
Arthur Gervais retweeted
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement. I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible. But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier. Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want. Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well. So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it. If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
3,431
12,357
2,700
72,180
9,557,026
existential threat only i can save you you must follow me give me that power otherwise you'll all die .. ehhh, what about no.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
1
9
528
mythos is a myth
Replying to @Xbow
3/ Cost is becoming just as important as raw capability. Lower-cost models like GLM-5.2 and Muse Spark 1.1 showed that “good enough” offensive capability is becoming increasingly affordable, while Grok 4.5 delivered near-frontier performance in the mid-budget range.
1
336
Arthur Gervais retweeted
I resigned from Apple today. I spent the last three years doing foldable research at both google and apple. Neither company is acting responsibly. They are racing straight to self-improving foldables and gambling with our lives. More thoughts below.
1,215
2,583
457
65,942
4,921,524
Arthur Gervais retweeted
Oh this is so shameful, it actually is Anthropic PR.
This post looks like the start of a VERY sophisticated and well-funded PR operation to get support for Democrats to regulate AI into oblivion. Let me show you how it works: 1.) This guy, with minimal followers and no previous account activity, goes to the Wall Street Journal which publishes an exclusive with quotes from him on his resignation 18 minutes BEFORE this post goes up. Planning was clearly done in advance. 2.) Within hours, it has tens of thousands of reposts and the account has 100k+ followers. The post is punchy, quotable, it almost seems professionally written. The first three accounts to quote tweet it all do so within 15 minutes of the initial posting. Remember, this account had basically zero engagement beforehand, so an organic reach explanation seems unlikely. According to Grok those accounts are @_NathanCalvin (General Counsel at Encode AI), @peterwildeford (Head of Policy at the AI Policy Network), and @DKokotajlo (Head of the AI Futures Project), all of which are up-and-coming AI-Doomer policy advocacy nonprofits. The AI Futures Project website says it is funded “primarily” by the Survival and Flourishing Fund, which says on its own website that it has advised Jaan Tallinn, Skype creator and one of the leading investors in Anthropic, to grant over $2.5 million to the AI Futures Project since 2024. Encode AI says on its website that it is ALSO funded by the Survival and Flourishing Fund, which in turn says that it told Anthropic investor Jaan Tallinn to grant $516,000 to Encode AI in 2025. And wouldn’t you know it, the Survival and Flourishing Fund ALSO says it told Jaan Tallinn to grant $2 million to the AI Policy Institute, the 501(c)(3) affiliate of the AI Policy Network, as well. What are the odds that the first three quote tweets of Coxon’s post would all be major AI-restriction policy advocates funded generously by the same donor, who also happens to be one of the leading investors in, and a board member of, Anthropic, the company Coxon was resigning from? And all within 15 minutes of posting (two within ten)? 3.) Jacob Coxon doesn’t have much of a resume, but we do know that, in 2022, he got a $20,159 scholarship for the “long term future scholarship program” from the Good Ventures Foundation, one of the philanthropic vehicles of Dustin Moskovitz, a notorious AI-doomer who has spent tens if not hundreds of millions on policy advocacy to strictly regulate AI, while also being an Anthropic Investor himself. It also just so happens that the 14th person to quote Coxon’s post was @MaxNadeau_ (27 minutes after posting) who is the program officer for the Technical AI Safety team at Coefficient Giving, another of Moskovitz’s philanthropic spending vehicles. Max is not a frequent poster, his last posts before quoting Coxon were before Labor Day, but he was remarkably quick off the mark for this one. 4.) Basically every major Democrat politician and candidate has suddenly glommed on to this post, and conveniently, as the people cry out foe answers, Bernie Sanders already has a bill written to “ban super intelligence” and regulate AI into oblivion, and will be releasing later this week. The bill, among many other things, will create “a new cabinet-level federal agency to safeguard the public from the dangers of artificial intelligence” that will be “advised by an Artificial Intelligence Advisory Board comprised of experts on artificial intelligence.” Do you think, perhaps, Anthropic and its many investors who fund AI policy advocacy might have interest in getting to place a pet “expert” on the board of an entity that dictates what AI is and isn’t allowed to do? And isn’t it fortuitous that this whistleblower came forward with his oh-so scary stories so close in proximity to the release of the most radical piece of AI legislation ever introduced?
47
156
11
4,069
352,528
Arthur Gervais retweeted
Chamath Warns: Your AI Data Isn’t as Private as You Think Today's statement from @OpenAI on the Navier-Stokes Problem and whether it used recent user data to boost its own capabilities: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Here's what @chamath had to say on AI's fragile privacy and the ZDR myth back in July: “Privacy in AI is very fragile and it's very brittle. There are all kinds of non-obvious data leak vectors lurking in AI. If you think that you're going to flip a ZDR switch, zero data retention, which is the magic term that the industry uses to tell you that everything's going to be okay. I think the answer and the message should be, ‘It's not going to be okay because you can't guarantee any of it.’ So the model companies, when they give you these zero data retention policies, are probably trying their best. But I think the reality is you are leaking information where you don't know it. And they, despite their best efforts, may still have trapdoors that they don't even know about until it's figured out by somebody else. You need an independent third-party layer to interface to these models to manage this exposure, because there are trapdoors everywhere.”
39
55
17
420
69,057
Arthur Gervais retweeted
Maybe start valuing SRs yourself by stopping your ridiculous charge to submit model
Security researchers are the backbone of onchain security today and across the broader internet. Because of them, we are surviving the AI security apocalypse. And that has been forgotten. Too many companies went from valuing SRs one day to abandoning them the next. They weren’t even thrown a life raft to follow along. At Immunefi, we see things differently. SRs were the core of security before, they remain the core of security today, and they will be the core of security tomorrow. We need to figure out, as an industry and as a community, how we’re going to adapt to these changes; how to bring all security researchers with us, along with the recognition and support they deserve for keeping us safe thus far. SR Summer was our way of showing a little recognition, and with Immunefi Studio, we will continue to put as much power in the hands of SRs as possible. Thanks to all those who took part, and congratulations to the winners!
2
2
39
2,756
Responsable Disclosure window will become seconds
1
159
Arthur Gervais retweeted
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.
1,861
4,748
2,806
41,830
12,124,301
soon I'm on vacation
1
3
381