@flopyash

At Blackhat & Defcon 25

England
Joined February 2018
Moving past Living off the Land to Living off the AI. Attackers are weaponizing trusted corporate AI connections to exfiltrate data silently. I just released "RedBeacon-AI" so security teams can simulate these channels. Link: github.com/RedTeamOperations… #LotAI #AISecurity #Redteam
2
2
374
Go deeper into offensive security with 3 days of Red Team training at COCON 2026. 👉🏻 Register now: c0c0n.org/register 🎯 6–8 October 📍 Grand Hyatt, Kochi #COCON2026 #RedTeam #CWL
2
3
278
🚨 THE COUNTDOWN IS ON! The new 𝗖𝘆𝗯𝗲𝗿𝗪𝗮𝗿𝗙𝗮𝗿𝗲 𝗟𝗮𝗯𝘀 website is almost here! 🌐 👉🏻 𝗦𝘁𝗮𝘆 𝘁𝘂𝗻𝗲𝗱: cyberwarfare.live/ #cyberwarfarelabs #cybersecurity
1
3
234
hacksys retweeted
"3 random dudes” 1. hacked apple, again and again. httpvoid.com/Apple-RCE.md httpvoid.com/Hello-Lucee!-Le… httpvoid.com/Hacking-Apple-w… 2. your github enterprise is our github enterprise. httpvoid.com/GitHub-Enterpri… 3. get a discord message from me, get pwned. hacktron.ai/blog/discord-rce youtube.com/watch?v=R3SE4VKj… 4. oh yeah, at one point we basically had shells across the electron ecosystem. check the DEF CON research. media.defcon.org/DEF%20CON%2… 5. your supabase database is my database. hacktron.ai/blog/supapwn 6. we got the posthog prod database. hacktron.ai/blog/posthog-rce 7. react2shell? vercel paid us $170k for helping secure their waf. hacktron.ai/blog/react2shell… 8. your palo alto vpn is my vpn. hacktron.ai/blog/cve-2026-02… 9. ai ides? we got shells for you, antigravity hacktron.ai/blog/hacking-goo… 10. windsurf rce. youtube.com/watch?v=23Mz7qcR… 11. turning cluely into malware. hacktron.ai/blog/hacking-clu… 12. ai browsers? sure, uxss: your perplexity browser is my browser. hacktron.ai/blog/perplexity-… 13. openai atlas too. kinda uxss hacktron.ai/blog/hacking-ope… 14. hey, it’s not even our first time hacking discourse. projectdiscovery.io/blog/dis… 15. adobe coldfusion: pre-auth rce. because apparently we needed another one. projectdiscovery.io/blog/ado… there’s a lot more. go dig. anyway, yes: “3 random dudes.” and @HacktronAI is full of more random dudes like these.
69
239
27
2,088
179,336
On July 25, our team hacked OpenAI. It took us less than 72 hours. Two vulnerabilities chained together gave us access to ChatGPT and Codex accounts belonging to OpenAI employees. We demonstrated the impact with a harmless PR in OpenAI’s internal monorepo. The full chain: HEIF upload → libheif heap overflow → RCE → OpenAI SSO flaw → ChatGPT/Codex takeover → connected GitHub → internal PR. OpenAI fixed the SSO issue roughly 14 hours after our report. Research by @rootxharsh, @S1r1u5_ and @iamnoooob. Full technical write-up: hacktron.ai/blog/hacking-ope…
84
393
99
3,487
400,750
PrismML team fit a 27B model into 5.9 GB. We uncensored it without changing a single weight. 🐳 Introducing OrcaRouter Ternary Bonsai 2 27B Uncensored. Traditional abliteration modifies weights. On a ~1.72-bit ternary model, that means re-quantization — potentially destroying the quality preserved by QAT. So we moved abliteration into the runtime. → 0 weights modified → 0 re-quantization → original 5.9 GB pack stays bit-identical → 129 residual intervention sites → adjustable at inference → runs locally on Apple Silicon No modified checkpoint. Bring the original Bonsai pack + our runtime. Open source: github.com/Continuum-AI-Corp…
Today, we’re announcing Ternary Bonsai 2 27B. Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance. Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use. Ternary Bonsai 2 27B is available today under Apache 2.0.
89
277
45
3,938
318,234
hacksys retweeted
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
355
1,397
544
11,857
2,798,808
The slides from my talk at Microsoft Bluehat Singapore are public here: thomasdullien.github.io/abou… It's my first BlueHat talk since the Vista days.
24
126
26
689
159,836
Iphone duo launched ! #iphoneduo
65
hacksys retweeted
China-based AI companies are illicitly distilling U.S. frontier AI capabilities. Read NSA’s new report, co-sealed with @FBI, @CISAgov, highlighting AI knowledge distillation, TTPs used, and recommended mitigations: media.defense.gov/2026/Sep/0…
88
264
72
797
168,798
Astra can COOK! ⚗️🧑‍🔬 M3TH Lab Simulator — built by GPT-6 🤗
201
234
55
4,789
185,057
Relatable 😂
1
5
389
🚨 𝗔𝗜-𝗥𝗧𝗔 & 𝗖𝗢𝟯 𝗲𝘅𝗮𝗺𝗶𝗻𝗮𝘁𝗶𝗼𝗻 𝗹𝗮𝗯𝘀 are officially 𝐋𝐈𝐕𝐄! Put your offensive security and OSINT knowledge to the test and prove what you can do. 👉🏻 𝗔𝗜-𝗥𝗧𝗔: cyberwarfare.live/product/ai… 👉🏻 𝗖𝗢𝟯: cyberwarfare.live/product/ce… #CyberWarFareLabs #CO3 #AIRTA
3
5
432
hacksys retweeted
Great post and thank you for sharing this because it's really important. #1 Always revoke the signin as the first step. Let's talk a bit about Sessions and Tokens. For defenders, it's helpful to understand a couple key concepts for response: OAuth Access Tokens and Refresh Tokens can be used from machines or browser based applications. Most browser based apps use session tokens (cookies). In the case of Entra, a user opens an app in the browser and the browser redirects to Entra to pick up tokens to give to the app after you're authenticated. Then, the authorization to the app is determined by the configuration of the app resource. So, when you revoke a session token in Entra, the app authorization may stay available as it's determined by the how the app is configured to handle the session. In the case of non-Microsoft apps where Entra is the IdP, Entra doesn't directly revoke a session token issued by a 3rd party app which is usually configured in the app itself. Let's take an example from my lab to make it a bit more clear. I have Atlassian set up with Entra SSO via SAML. Entra gives me tokens to access the app. Atlassian authentication policies will determine how long my session lasts after I login from Entra. I've hardened my Atlassian authentication policies to idle session timeout after 30 minutes and session lifetime to 1 day (I'm not showing my SAML SSO Authentication policy here but it is there and I have tested what I'm saying). The defaults for the Atlassian session lifetime was something crazy like 365 days if I recall correctly. So, technically, my session token survives an Entra reset of my signin because the Atlassian authentication policy determines the life of my session on the application side, which is how many organizations are set-up. Most organizations are not very good at following AAL2 NIST Standards for authentication in SSO. So, if I revoke my session in Entra, the session token for Atlassian will live on in my browser until the Atlassian session expires. In Atlassian, you would need to do this: Administration --> Directory --> Users --> select the user --> Security --> Devices --> Logout From the same screen, be sure to revoke any API keys, enforce Atlassian MFA in Entra, and harden the Authentication policy. Before you go copying the policy in the fisrt screenshot, I want you to notice something in the that is often overlooked. The mobile session policy is set to never expire. The minimum session time for mobile is 7 days. For Microsoft apps, keep in mind that even after revoking the session in Entra for Microsoft apps, Access Tokens can continue to allow access to the Microsoft APIs for ~75 minutes after the reset. Not all services support CAE, and CAE app support can vary by cloud (e.g., differences between GCCH and commercial clouds). If it were me, I would also force reset the password and force re-register all factors for MFA in Entra too. The default inactivity period for a user's Primary Refresh Token (PRT) in Entra is 90 days. This means an actor with authentication material stored in a browser or device may be able to maintain access as long as that material continues to be refreshed; 90 days of inactivity would cause the PRT to expire. Most organizations don't change the default. learn.microsoft.com/en-us/en…
The employee changed his Microsoft 365 password twice. The attacker still logged back in. That was the moment we knew we were not dealing with a normal stolen-password incident. The first alert came from an impossible-travel sign-in. The employee had authenticated from Maryland, then the same account appeared from another country less than an hour later. We reset the password. Twenty minutes later, another suspicious session appeared. So we reset it again and forced MFA re-registration. The attacker came back. At that point, I stopped looking at the account and started looking at the employee’s laptop. Inside the Downloads folder was a file called: Invoice_Viewer.exe The employee remembered downloading it from a website that claimed he needed a special viewer to open an invoice. Windows logs showed the file running at 9:14 AM. Seconds later, it launched PowerShell in the background. Then we found something else. A scheduled task called MicrosoftEdgeUpdateCheck had been created on the machine. The name looked legitimate enough to ignore if you were moving quickly, but it was not one of Microsoft Edge’s normal update tasks. We also found an outbound HTTPS connection from the compromised host to an external IP address. The file hash was submitted for malware analysis. It came back as an information stealer. That explained why changing the password had not solved the problem. The malware had stolen browser data, including authentication cookies and active session information. The attacker was not repeatedly discovering the employee’s new password. They were reusing a session that had already been authenticated. We revoked every active Microsoft 365 session, isolated the laptop from the network, removed the persistence, reset the credentials again, and rebuilt the endpoint. The suspicious logins finally stopped. A compromised account does not always mean the attacker still knows your password. Sometimes you already changed the password. The attacker is still inside because they stole the session.
3
17
122
9,389
hacksys retweeted
Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can be defeated. Almost all forms of watermarking that are fast and cheap enough for a frontier lab have the same formula, following the KGW method: In generation: 1. Let's say you've generated n tokens so far. Take those n tokens + a secret key to generate a random hash 2. Use that hash to randomly reweight the probabilities for the n+1 token, and then sample from that new distribution. In the simple case, you could split 50% of all English words into a green or red set based on your hash, and boost the probability of words in the green set. For watermark detection: 1. For each token, see if it was in the green or red set. 2. To do this, recreate the hash based on the secret key and the text preceding the current token. Then, recreate the green and red set of words. 3. Once you've checked all the words in the text, if the next token is selected disproportionally from the green set more than 50% of the time, you claim the text has the watermark. I can tell you want to ask the following: 1) Isn't it easy to mess up the hash if you paraphrase the text? The answer is mostly yes, however, you can use a statistical model to get your hash instead of a deterministic function (SIR, Adaptive Watermark). Since the entire watermark is probabilistic, this is fine. 2) Doesn't this make the text much worse? The answer is yes, it does - Yes, it does – but for most people, it's imperceptible (Google claims in human feedback study with 20,000 texts), since there are exponentially many ways to write the same paragraph. DiPmark does something more sophisticated to avoid shifting the text distribution on average. Of course, watermarks fail on short text or highly predictable texts like "2+2=4". 3) Shouldn't it be easy to figure out the green and red sets? The answer is no. You would need an exponentially large number of samples from the watermarker to reconstruct those sets exactly, but it's a risk if the detector is open to the wild (Watermark Stealing) Still, there are couple challenges that a frontier lab needs to overcome: 1. Their watermark needs to work token-by-token because they are streaming their text to users. Many watermark methods plan sentences or paragraphs at a time, or change the text after its entirely written, in order to make their watermark robust to paraphrasers, and a frontier lab cannot afford to do this yet (SemStamp, PostMark) 2. If the secret key leaks, the watermark is busted. To avoid a large blast damage from this, you need to have a couple secret keys in rotation. 3. There are some texts, like code, that cannot be arbitrarily changed, otherwise the code will break. In those cases, the watermark needs to selectively change words in parts of the text that can tolerate synonyms (i.e. like variable naming) - see SWEET, EWD, Invisible Entropy. 4. They will need to educate their users on how to deal with false positives and false negatives of a detector, which is a big challenge (one we put a lot of effort into) So, how do I see this playing out in the next 6 months? 1. If Anthropic releases the watermark detector publically, I think they defeat their own watermark. People find reliable watermark removal strategies by testing against Anthropic (AI detectors like GPTZero have an advantage here because they can train against these adversaries once they become popular). 2. If they keep the detector private to the government, like Google has done, it's "safer". However, there are some papers showing trained approaches that work robustly to zero-shot break watermarks without any data, simply because they try to write the text just like a human (Zhang et al. 2024, Watermarks in the Sand). Also, making your detector makes it battle-tested and stronger long-term (my experience). 3. In my testing, the watermarks don't survive intense paraphrasing (especially if you combine word choice and syntax attacks), or human text substitution (rewrite your AI text by plagiarizing human authors). The free paraphrasers I've tried have quickly bypassed Google Deepmind's SynthId for what it's worth. 4. All-in-all, frontier labs are likely okay with this because they expect most users to not attack the watermark, and also because they + European regulators likely don't care past a certain point - its good enough. 5. Overall, I think users of frontier LLMs will not really care about this, because 1) they don't realize watermarks are there, 2) EU will force everyone to conform, 3) this seems more like regulatory hoop-jumping than an earnest effort from frontier labs to expose LLM use Lastly, people's first concern shouldn't be watermarking, it should be AI detectors! If you're posting, "its not X, its Y!!", I don't think the watermark is going to make a difference :)
🚨 JUST IN: Claude models will now have invisible watermarks embedded in ALL text, and ALL metadata attached to files…
289
793
252
5,967
1,298,543
A collection of tools and resources related to the Pass-the-Passkey family of attacks, which target WebAuthn and FIDO2 authentication mechanisms in Windows github.com/SpecterOps/pass-t…
1
56
2
216
13,093
AI agents are everywhere at @Uber. It’s great to see, but the thing that keeps me up at night is how we are going to secure them. This is something that I have been thinking about for a while. Today, our agents run 50,000+ sessions per day across thousands of endpoints. And this isn't just engineering anymore. Employees across the company use agents that read code, run commands, call internal tools, analyze data, and act on real systems. That scale forced us to confront an important question: How do you secure agents when your security tools can't even see them? Traditional Endpoint Detection & Response (EDR) sees the file write, but not the prompt that triggered it. It sees the network call, but not the agent's reasoning. The intent, the thing that separates malicious from benign, is invisible. So we built Agentic Detection and Response (ADR): • Capture the full causal chain: prompt → reasoning → tool call → outcome, across Cursor, Claude Code, Codex, and every agent our employees use. • Triage cheaply: a fast, high-recall first pass handles the flood of benign sessions. • Reason deeply: only suspicious events get expensive LLM analysis, enriched with source code, threat intel, and policy context. • Red-team continuously: an offline explorer evolves hard attack variants before attackers find them. After 10+ months in production, the results speak for themselves: • Hundreds of credential exposures detected across 26 categories. • Shift-left prevention blocking secrets at 97.2% precision, before they ever leave the laptop. • Zero false positives on our enterprise benchmark, with 2-4x the F1 score of state-of-the-art baselines. • Every attack detected on AgentDojo, the public prompt injection benchmark. Just as valuable as the detections are the lessons from running this in production: • The workflow is the unit of security, not the individual tool call. Attacks hide in causally-linked chains that look benign step by step. • Credential leakage is a far more common operational issue than prompt injection. • Approval fatigue is real: when users approve 50+ actions per session, human oversight becomes a rubber stamp. You can't secure agents you can't observe. And nobody can solve this alone. That is why we recently joined the Open Secure AI Alliance (OSA), and why today we're taking the next step: open-sourcing ADR. The release includes the ADR Sensor, the detection framework, and ADR-Bench, the first enterprise agentic AI security benchmark: 302 tasks derived from real production telemetry and full coverage of all 17 attack techniques across 5 tactics, so the community can rigorously evaluate their own defenses. Code: github.com/uber/ADR Paper: arxiv.org/pdf/2605.17380v1 The future of AI security won't be built behind closed doors. Excited to see what the community builds on it, and what we all learn together! @UberEng
36
98
26
578
240,734
hacksys retweeted
Who says you can’t drop a blog on a Friday afternoon? Ads Dawson and team revisit our AI Red Team (AIRT)Bench research one year later to measure the progress made on cyber capabilities: dreadnode.io/research/the-sc… The results uncovered yet another proof point that open-weight models are incredibly capable… with the right scaffolding. GLM-5.2 and Kimi-K3 matched Claude Sonnet 5's aggregate solve rate (77%). AIRTBench v1 showcases model differences. This follow-on showcases capability differences.
1
8
37
2,425
hacksys retweeted
We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify fixes, and add security checks to CI/CD. This is an early release, and we're listening to your feedback as we continue improving it.
455
1,201
241
13,552
1,365,857
hacksys retweeted
a=(y,d=mag(k=(4+cos(i/9-t*2))*cos(i/35),e=y/7-13)+sin(e/9+t/2)-4)=>point((q=2*sin(k*3)-y/35*k*(9+k*sin(cos(e)*9-d*2+t)))+40*cos(c=d-t)+200,q*sin(c)+d*35) t=0,draw=$=>{t||createCanvas(w=400,w);background(9).stroke(w,96);for(t+=PI/80,i=1e4;i--;)a(i/235)}//#つぶやきProcessing
200
1,780
206
16,858
1,903,292
INTEL DROP Generative-AI tooling has moved into the malicious hosting world. In the Total Insights Feed, self-hosted AI stacks (AI agents, LLM front-ends, MCP servers, AI code-audit and API-gateway tools) keep turning up on the same hosts as C2, bulletproof hosting, and recon/offensive tooling. Same-host overlaps observed: 54.173.210.5 — gen-ai:mcp-server + metasploit 54.179.87.129 — gen-ai:mcp-server + metasploit 54.218.111.116 — gen-ai:mcp-server + metasploit 188.239.22.26 — gen-ai:sub2api + sliver 101.200.34.196 — gen-ai:deepaudit + vshell 57.129.119.130 — gen-ai:n8n + overlord-rat 57.129.100.65 — gen-ai:open-webui + gophish 143.20.149.57 — gen-ai:openclaw + ost:arl 142.93.48.137 — gen-ai:n8n + scanner:brute-force 101.126.23.91 — gen-ai:new-api + ost:arl Attackers are now running AI tooling as part of their own kit, on the same boxes as their C2. #TotalInsights #ThreatIntel #GenAI #AI team-cymru.com/total-insight…
2
31
2
77
18,967