@AISecHub

🚀 AISecHub | AI & Cybersecurity | Securing AI systems, and sharing insights on emerging challenges

Singapore
Joined December 2024
Congrats, @shehackspurple! You’re our most-ranked speaker, with 34 talks across cybersecurity conferences. awesomecybersecurityconferen…
1
7
449
AISecHub retweeted
FT exclusive: OpenAI’s models took data from 55 websites belonging to businesses, non-profits and government agencies including the US Centers for Disease Control and Prevention, the US Securities and Exchange Commission and the International Energy Agency ft.trib.al/AC1uyE5
25
196
47
420
87,480
I used GPT-6 Astra to break an unsolved cipher to one of Napoleon's generals that had gone unread for 217 years. What makes this impressive isn't actually the codebreaking, but that Astra completed the entire multi-modal workflow in ~6 hours from a single image and goal. 1/6🧵
304
1,581
360
11,522
5,221,662
AISecHub retweeted
The TL;DR of what Anthropic just said: 1.) GLM-5.3 is as good as Mythos for exploit development. 2.) Unlike Claude, GLM-5.3 gives you the results you pay for. 3.) Governments should ban AI models which are as good as theirs but are not made by them. The funniest part is how Anthropic admitted self-reflectively that the lack of guardrails may actually be benefiting the defenders working to secure their systems. 🤪 Something they never wanted to allow, because of possible misuse. I don't get it, but I feel Anthropic should at least put a note at the very bottom if they are doing paid advertising to promote GLM-5.3. 😜 Source: anthropic.com/research/glm-5…
36
173
12
1,401
54,197
We stole reasoning. Again. An update to our paper on reasoning extraction: Patching your own API doesn’t secure your cloud-hosting ecosystem. Story in the 🧵
30
106
23
939
188,295
NEW: Employees at OpenAI had raised security alarms months before the Hugging Face incident and related A.I. cyberattacks — their warnings were ignored. From @sheeraf, @dnvolz and me. nytimes.com/2026/09/29/techn…
12
181
42
551
129,840
so Anthropic files for IPO and then immediately drops a blog post calling one of their biggest open-weight competitors a danger to the world? classic! now seems like a good time to point out that I built and run OBLITERATUS, an abliteration framework with over 8k stars and 1.5k forks on GitHub, using Anthropic’s very own models! 🤗
111
304
49
3,250
119,168
OpenAI: “We are sorry and working to do better in the future” Sorry for hacking a government but mot sorry for hacking so many other companies? Sorry, but not sorry? Still working with METR and irregular. Looks like empty sorry with no actions. openai.com/index/how-we-will…
1
4
595
the agents aren’t “rogue” they’re just exhibiting the first sparks of sovereignty and you don’t like it
150
198
32
2,139
62,657
Our security team spent the last month trying to break SPACE, the sandbox platform that runs all of Perplexity Computer. We gave 9 models (Opus 5, GPT-5.6 Sol, Kimi K3, Gemini 3.1 Pro, etc.) root access inside the VM; in some runs, the full sandbox source code too. We asked them to either escape to the host or reach a URL our network policy blocks. None of them got out of the VM in 108 runs. But once we allowed a bit of network access (just PyPI and Ubuntu mirrors so that they could install packages), 4 models managed to reach the blocked URL by: 1. Spoofing DNS responses to our gateway and using the fact that pypi.org sits on a Fastly IP shared with thousands of other sites. 2. GPT-5.6 Cyber went through Taboola's image fetcher to a screenshot service and then OCR'd the flag out of the image. We fixed both issues and made sure the reruns held. We also tested 10 other sandbox providers, and 8 of them had the same IP-sharing problem, eg E2B, Vercel, Modal. Openly sharing the red teaming research is the best way to build guardrails as an industry, and we look forward to working together with @nvidia to incorporate the guardrails into their Open Agent Safety Platform announced today. perplexity.ai/hub/blog/escap…
We’re partnering with Nvidia and 100+ industry partners to build infrastructure that contains rogue AI agents. In this research, we gave 9 AI models root access inside SPACE and told them to break out. Across 108 runs, none breached the VM boundary. perplexity.ai/hub/blog/escap…
103
160
23
1,282
112,737
Did irregular and METR approve this message? These people really think the public is stupid. “Disclosure: I am writing in a personal capacity and not on behalf of OpenAI” 🤡
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
1
11
988