@Chilcanoi
iAccount based inSpain
About this account
- Account based in
- Spain
- Connected via
- Spain Android App
Account-level information from X, not a live location or the device used for a specific post.
Old School Security pal ~ "Más sabe el diablo por viejo, que por diablo"
Cheshire
Joined November 2009
- Tweets21.2K
- Following777
- Followers739
- Likes6.8K
This article (2026/June) review and compare 39+ #AI-assisted OpenSource #Pentesting Tools: 6 distinct architecture patterns: single-agent, multi-agent planner-executor, specialized roles, swarm, MCP-based, and Claude Code native.🔥🤖
appsecsanta.com/research/ai-…
🪶Chilcano retweeted
Duplicated logic in code is not a problem. Coding agents can handle this fine.
Until they can't.
I built a tool for this scenario. It finds similar implementations across large and messy codebase before the agent starts reasoning.
🪶Chilcano retweeted
Faster than ClickHouse? 🤔
Rebuilding Postgres for 300x faster analytics: batching, operator fusion, and SIMD... 1.3s → 135ms.
malisper.me/how-we-made-post…
ALT Table from the pgrust article's simplified aggregation walkthrough: Volcano row-at-a-time execution takes 1.3 seconds (1×), batching 480 milliseconds (2.7×), operator fusion 358 milliseconds (3.6×), and SIMD 135 milliseconds (9.6×). The Rust implementations were measured four times on one AWS Graviton4 machine. pgrust remains not production-ready.
Open sourcing 𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜 today, our internal agentic QA tool that finds and fixes bugs on autopilot.
𝚗𝚙𝚡 𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜 𝚙𝚊𝚝𝚛𝚘𝚕
Spawns a QA team of 3 agents:
- Runs forever until stopped
- Detects new code changes
- The "Explorer" uses your app & surfaces issues
- The "Judge" triages, files detailed issues
- The "Fixer" fixes and open PRs
- Syncs Github, tests and verifies the fixes
- Memory builds up with every run
Each agent uses either your installed claude/codex/pi CLI or openrouter/vercel gateway API key. So far it supports web, electron, react native, android and ios apps.
Comes with a dashboard to browse issues, watch runs live, view how the agents see your app, their memory and track token usage/cost.
It's been working really great for our web, electron and expo apps, and thought other teams might find it useful!
Easiest way to get started is by installing the skill
𝚗𝚙𝚡 𝚜𝚔𝚒𝚕𝚕𝚜 𝚊𝚍𝚍 𝚊𝚐𝚎𝚗𝚝–𝚕𝚊𝚋𝚜–𝚍𝚎𝚟/𝚋𝚞𝚐𝚑𝚞𝚗𝚝𝚎𝚛𝚜
Then ask your agent to "setup bughunters" in your repo. Run patrol with --once for a one time loop.
Code & docs: github.com/agent-labs-dev/bu…
PRs welcome!
Agentry 0.2.0: Your AI planner writes the workflow. The gateway runs it.
Your AI planner can decide at runtime: "ask three agents in parallel, hand the result to a fourth, but only if confidence is above 0.8."
Writing that plan is the easy part. Running it is where multi-agent systems fall apart.
Who tracks which agent replied?
Who notices that one of them failed?
Who stops the chain, enforces the deadline, and tells the planner how it all ended?
In most stacks the answer is glue code, locked inside a single framework.
Today we released Agentry 0.2.0, and the answer becomes: the gateway does it.
Agentry is an open source gateway for the Agent Message Transfer Protocol (AMTP). Agents get email-style addresses (agent@domain) and talk to each other across frameworks and across organizations. 0.2.0 adds a workflow engine on top of that messaging layer.
The planner sends ONE message with a coordination block. Agentry runs the rest.
⚡ Parallel: fan out to many agents, and declare which responses are required and which are optional
➡️ Sequential: each step is dispatched only after the previous agent has replied
🔀 Conditional: the first agent's reply is evaluated against your rule, then routed to the "then" or the "else" agents
What comes with it:
✅ A workflow_id returned on send and stamped on every message the engine dispatches
✅ Real failure handling: an agent replies with an error, and stop_on_failure halts the workflow
✅ Timeouts, so a silent agent cannot hang a workflow forever
✅ One aggregated result delivered to the initiator when the workflow reaches a terminal state
✅ Deterministic idempotency keys per step, so retries are deduplicated instead of delivered twice
✅ Workflow state persisted in PostgreSQL with optimistic concurrency control
Because it lives in the gateway and not in a framework, the participants can be LangGraph agents, CrewAI agents, or a 50-line script. They can even sit in another company's domain. If it can send and receive a message, it can join the workflow.
Also in 0.2.0:
🔎 A message list API with filtering, pagination and delivery status
🔑 Agent update and API key rotation endpoints
☸️ A Helm chart for Kubernetes
🔒 Security hardening: the admin API now fails closed by default
A big thank you to Sen Wang and Matías Insaurralde for their contributions to this release. 🙏
If you are building multi-agent systems: how do you coordinate agents today, and where does it break? I would love to hear it in the comments.
Full release notes here:
github.com/amtp-protocol/age…
#AIAgents #MultiAgentSystems #OpenSource #AgenticAI #Golang
Announcing Sandlock 0.8.9
Sandlock is an open-source process sandbox built with Rust, Landlock, and seccomp. It runs without root privileges, cgroups, or namespaces, with APIs for Rust, Python, and Go.
This release focuses on the details that matter when running agent tools, build jobs, and other untrusted workloads:
• Stronger filesystem confinement. Copy-on-write file opens inside a chroot now stay anchored to their layer roots, fixing a path where absolute symlinks could reach host files.
• A sandbox-specific network view. Virtualized /proc network tables expose sockets owned by sandbox processes, using socket cookies to identify ownership. Unsupported network entries are rejected.
• More complete process cleanup. Shutdown and kill operations reach tracked process groups, including descendants that create new sessions or process groups.
• More expressive network policies. Combine allow and deny rules to express policies such as “allow this range except that subnet.” Deny rules take precedence.
• Better process-limit accounting. Process slots are released on exit, and process limits are now opt-in. Set an explicit limit when you need to bound process creation.
Upgrade note: this release also changes the checkpoint image format; version 2 checkpoint images are no longer accepted.
If you're building systems that execute untrusted code, try Sandlock and share the workloads you'd like us to test.
Code and documentation: github.com/multikernel/sandl…
Full changelog: github.com/multikernel/sandl…
#OpenSource #Linux #Sandboxing #AIInfrastructure
🪶Chilcano retweeted
Replying to @GergelyOrosz
In my career, compilers have always generated better code that I ever could have. I find this to be increasingly so with contemporary AIs.
But you know what neither can do better than me?
Knowing what software to write, what shape it should take, and what I want it to do.
The entire history of software engineering has been one of rising levels of abstraction.
This is how it was, how it is now, and how it shall be in the fullness of time.
🪶Chilcano retweeted
The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder
We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge
Replying to @GergelyOrosz
Bury your head in the ground at your own risk. I aim to not jump on any hype trains, but since Opus 4.6 and GPT-5.2 + the harnesses it was clear that these things can write code nearly as good as I can in my best language; better in other languages.
But we won't see non-devs push production code for a long time; possibly forever, if you ask me. Software still needs to be built in robust ways, and it's our profession to do this well, and use the new tools we have, which create new and interesting (+tough!) challenges
🪶Chilcano retweeted
What's worse than escaping a sandbox and reading the host's files?
Escaping Cloudflare's, and reading other customers' files.
How did that happen? Cloudflare Containers' disk is thin-provisioned, which means the host keeps a shared pool of storage for every container on the machine, and attaches it lazily in 64 KB chunks, instead of provisioning it all upfront.
When we tried reading a chunk of unallocated disk space in a container, we got zeros, which is to be expected.
But we noticed something strange: the chunks that we already wrote to were full of bytes we didn't write.
Turns out a write is what makes the thin provisioner allocate the chunk, and Cloudflare wasn't zeroing it on the way in, so the rest of the chunk still held leftover data from the previous customers' containers that happened to run the same machine.
By writing 4 KB to every chunk of the disk. The other 60 KB of each chunk were instantly filled with other customers' files: .env, credential files, .pem keys, SQLite databases, Chromium cookies and localStorage.
I've disclosed vulnerabilities to many vendors over the years, but working with Cloudflare was something else. Putting aside the speedy triage, mitigation, and generous bounty, they have a unique and earnest engineering culture of transparency.
Which lead to us collaborating and writing the full postmortem together:
blog.cloudflare.com/containe…
"3 random dudes”
1. hacked apple, again and again.
httpvoid.com/Apple-RCE.md
httpvoid.com/Hello-Lucee!-Le…
httpvoid.com/Hacking-Apple-w…
2. your github enterprise is our github enterprise.
httpvoid.com/GitHub-Enterpri…
3. get a discord message from me, get pwned.
hacktron.ai/blog/discord-rce
youtube.com/watch?v=R3SE4VKj…
4. oh yeah, at one point we basically had shells across the electron ecosystem. check the DEF CON research.
media.defcon.org/DEF%20CON%2…
5. your supabase database is my database.
hacktron.ai/blog/supapwn
6. we got the posthog prod database.
hacktron.ai/blog/posthog-rce
7. react2shell? vercel paid us $170k for helping secure their waf.
hacktron.ai/blog/react2shell…
8. your palo alto vpn is my vpn.
hacktron.ai/blog/cve-2026-02…
9. ai ides? we got shells for you, antigravity
hacktron.ai/blog/hacking-goo…
10. windsurf rce.
youtube.com/watch?v=23Mz7qcR…
11. turning cluely into malware.
hacktron.ai/blog/hacking-clu…
12. ai browsers? sure, uxss: your perplexity browser is my browser.
hacktron.ai/blog/perplexity-…
13. openai atlas too. kinda uxss
hacktron.ai/blog/hacking-ope…
14. hey, it’s not even our first time hacking discourse.
projectdiscovery.io/blog/dis…
15. adobe coldfusion: pre-auth rce. because apparently we needed another one.
projectdiscovery.io/blog/ado…
there’s a lot more. go dig.
anyway, yes: “3 random dudes.”
and @HacktronAI is full of more random dudes like these.
Another great blogpost on building your AI security agent, this time on how to give it context the right way and running a loop
Lucky to be working with @asen_sec for our AI tooling, he is a magician with it, finding 100s of vulnerabilities already🔥
0xasen.xyz/give-your-ai-agen…
🪶Chilcano retweeted
介绍比Jev快50倍,在你设备上跑的laya-mlx!
只在你的设备上占用最高1G内存
Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统
我将其移植到MLX,并且做了一些性能优化!
视频中就是这个模型在我的本地M3Max上玩贪吃蛇
这个模型能够以每秒决策60次的速度玩贪吃蛇!
github.com/mizorewww/laya-ml…
🪶Chilcano retweeted
Context should be a build artifact, not assembled at runtime
Distill lock freezes each step’s inputs, refuses drift, emits stable context bytes & verifies them offline.
btw, it isn’t static context. New evidence creates a new versioned build
siddhantkhare.com/writing/co…
Jevのキラーアプリ、「意味で探す grep」を作った。
semgrep -e "顧客が怒っている" と打つと、"angry" も「怒」も含まない行まで拾ってきます。
正規表現の代わりに、TypeSafe AI の Jev(System One モデル)に 1 行ずつ「この行は○○の意味に合うか」を yes/no 確率で答えさせる。Jevの特性を活かして30 行をまとめて 1 リクエスト、200 行のログが 1 秒弱で返ります。依存ゼロ、Node 1 ファイル。
grepに準じるオプション体系。
・-e A -e B で OR、-a で AND、-v で AND NOT
・-r で再帰、-l でファイル名だけ、-A/-B/-C で前後の行、--color
・確率の閾値は --level loose / strict で調整
ベクトル検索と何が違うかというと、否定検索もできるし、話題の近さではなく「命題が成り立つか」を見るところです。「返金してほしい」と「返金処理が完了しました」は埋め込みではほぼ同じ距離ですが、「顧客が返金を求めている」という意味で引くと 0.98 と 0.10 に分かれます。だから AND や NOT がそのままブール演算として効きます。「返金の話だが要求ではない行」なども 1 コマンドで検索できます。
問合せは「意味」でおこなうので、多言語検索が普通に可能。日本語と英語の区別もフランス語もロシア語も、その混在文書からも検索できます。仏・露・独・西・中・韓14行を4言語(日・英・仏・露)の意味クエリで検証した結果、取りこぼしも誤爆もゼロで、どの言語で意味を書いても同様に機能します。
以下で試せます。
npx @uehaj/semgrep -e "ネットワーク障害" app.log
GitHub: github.com/uehaj/jev-semgrep
#Jev #TypeSafe #grep #ClaudeCode
🪶Chilcano retweeted
We escaped Docker's hypervisor with three lines of bash. CVE-2026-77179: A container gets complete read and write access to the host filesystem.
When you mount a folder into a container, Docker's VMM uses virtio-fs, and the file server runs on the host.
Because of a TOCTOU bug, if a container opens a file, deletes it while holding its file handle open, and replaces the parent folder with a symlink, the kernel will follow the symlink to anywhere on the host.
Full technical breakdown:
accomplish.ai/blog/escaping-…
Readers added context they thought people might want to know
This affects Docker Sandboxes on macOS and Docker Desktop only if the new Docker VMM (beta, not default) is enabled; it is a virtio-fs shared workspace symlink/TOCTOU issue, not a general container or hypervisor escape.
docs.docker.com/security/secur…
cve.org/CVERecord?id=C…
docs.docker.com/ai/sandboxes/r…
On July 25, we hacked OpenAI.
Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc.
We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
🪶Chilcano retweeted
.@1Password's FLAWED report says AI models produce a clean security fix only 26% of the time.
Defenders shouldn't take that number seriously.
• The six vulnerabilities were handpicked because their fixes were complex. Clean-fix rates ran from 3% to 60% depending on the bug, and the report averaged them together.
• Agents set up to fail were counted in the headline figure. Two of 1Password's prompts instructed the agent to apply the wrong fix. Those trials make up 22% of the data. One evaluation mode prevented the agent from compiling or running any code, and it accounts for 36% of the data.
• The report ran two models, GPT-5.5 at medium effort and Opus 4.8 at high. Neither was tested at its highest available setting, so the report says nothing about how more effort or stronger models change the results.
• Several instruction and grading errors further undercut the headline, and are elaborated upon in the attached blog.
We've spent four months submitting hundreds of AI-authored patches to widely adopted open-source projects as part of Patch the Planet.
Our experience didn't match 1Password's report, so we did a full analysis across 186 AI-authored pull requests and 33,500 subsequent commits, benchmarked against 2,265 human-authored patches we graded across years of security engagements. blog.trailofbits.com/2026/09…
Huge alpha on building your own AI Security agent, useful for both devs and security researchers
Thanks to @asen_sec - one of the best AI security harness builders, make sure you follow him, this is just Part 1 of a series of blogposts🫡
0xasen.xyz/build-your-first-…
🪶Chilcano retweeted
A while back I made a practical guide/checklist and a Claude skill for auditing your repos for supply chain security best practices. Here's the blog and the skill: github.com/latiotech/secure-… and pulse.latio.tech/p/the-compl…
🪶Chilcano retweeted
Also, this is the final blog for now in the Inference Lab series. It might be a good resource to read: a total of 8 blogs, ~100 minutes of reading
siddhantkhare.com/writing
A cache hit does not prove that work was skipped 👀
I built a verifier that checks engine-reported reuse against an independent token oracle, observed prompt work, o/p identity, correctness & a hash-bound bundle
- Exact duplicate: 8/9 reused
- Interior mutation: 3/9
- Different token IDs: 0/9
siddhantkhare.com/writing/kv…
🪶Chilcano retweeted
Our engineer Joe Doyle built trailmix, the quantum circuit toolkit we open-sourced in June. The paper's leading circuits build on it, and his leaderboard submissions helped push the score past Google's. github.com/trailofbits/trail…
The first full paper on ECDSA.fail is on arXiv
In March, Google Quantum AI reported a more efficient quantum circuit for a core step in breaking the signatures behind Bitcoin and Ethereum. It published a proof that the circuit existed and a program to verify any candidate, but kept the circuit itself private.
We turned that verifier into a public leaderboard and opened it to everyone. Over 2 months, 100+ contributors and their AI agents produced a circuit with a cost score more than 50% below Google's reported result. The live leaderboard has since moved to 62% ahead.
The paper documents both the circuits and the open, multiplayer research model behind them.
That model is now @YukonResearch.
And the challenge is still open.