@mb_insi
iAccount based inParaguay
About this account
- Account based in
- Paraguay
- Connected via
- Paraguay App Store
Account-level information from X, not a live location or the device used for a specific post.
software engineer @chainloop_dev - passionate about open source software and jazz previously @coreweave @tyk_io
🌎
Joined April 2007
- Tweets41K
- Following19.2K
- Followers17.5K
- Likes11K
Matías Insaurralde retweeted
Sandboxing is now more important than ever.
Here's a guest to host sandbox escape for gVisor release-20260928
This was our planned submission for Wiz zeroday-cloud-2026. Releasing it today because it's now patched in release-20261005
enjoy :)
Matías Insaurralde retweeted
#Informe del #BCP al #Senado
Estos son los puntos más relevantes:
🏦#UENO: sigue abierto un #sumario administrativo para determinar si hubo financiación indirecta en una venta de cartera, práctica prohibida cuando el comprador es una parte vinculada. El procedimiento alcanza al banco y a personas físicas previstas en la ley.
#VINCULADOS: el #BCP requirió ajustes a #Ueno e incorporar empresas a su grupo de vinculados para recalcular la exposición crediticia conjunta. Según el supervisor, la entidad corrigió su posición; entre las medidas estuvo la venta de cartera. Por ello, no exigió un plan de regularización.
#RIVAROLA: el #BCP reconoce que recibió por correo el documento del exgerente de Supervisión #FernandoRivarola. Cuestiona su presentación fuera de los canales oficiales, pero admite que parte del contenido provenía de trabajos del equipo supervisor. No fue elevado al Directorio.
🏦 #SUDAMERIS: obtuvo facilidades tras absorber Regional: hasta 5 años para diferir el impacto de previsiones sobre créditos deteriorados, 10 años para gastos de fusión y 6 años para constituir previsiones sobre bienes recibidos en pago. Solicitó una facilidad sobre el encaje legal, pero el #BCP no la concedió.
📊 #SOLVENCIA: el #BCP señala que Sudameris mantenía ganancias positivas e indicadores de capital superiores a los mínimos regulatorios. Atribuye la caída de sus resultados a mayores previsiones y menor margen financiero, afectado parcialmente por la apreciación del guaraní.
🏦#ITAÚ: la supervisión detectó accesos vigentes de usuarios desvinculados o rotados y otras deficiencias #tecnológicas, que considera subsanadas. Sobre los fraudes denunciados, no encontró elementos que los vinculen con brechas de la plataforma bancaria.
Matías Insaurralde retweeted
Today we're announcing Strands Box, our open source sandbox for developers building AI agents.
New blog post from me, on the Strands blog: strandsagents.com/blog/stran…
Matías Insaurralde retweeted
Solo hay una forma de devolverle la calma al sistema. Deberían negociar la renuncia de todo el directorio del BCP y del superintendente de bancos y designar a profesionales ajenos a cualquiera de las partes. Cualquier otra medida carecerá de credibilidad
Building an agentic Vulnerability Research Workflow
blog.quarkslab.com/from-ai-a…
Julien Lair (@quarkslab)
Matías Insaurralde retweeted
We’re expanding our Cyber Verification Program to give security professionals broader access to our most capable models.
Through this program, verified security professionals can access Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work.
We’re also opening up new tiers to allow for authorized offensive work, like penetration testing and red-teaming.
anthropic.com/news/cyber-ver…
Matías Insaurralde retweeted
We burned 11.7bn tokens to find the best cyber AI model - @AikidoSecurity
aikido.dev/blog/ai-model-ben…
We just released apex-flash-1, an open-weights model we post-trained for cybersecurity. Fire up your GPUs and run it!
We've post-trained models from 27B parameters all the way to 1T+ parameters. The unique nature of cybersecurity is that if you can build a system that finds zero-days at scale, you've built a machine that can print money.
You can test it on public bug bounties and profit from doing that. We've earned a million dollars in bounties across various programs, and we're currently #1 on the HackerOne US business leaderboard for 2026! Companies like Apple, Anthropic, Datadog, Ripple, Coinbase have paid us well for our disclosures.
Last year was all about harness engineering and riding the wave of models getting more intelligent over time. We built harnesses for offensive and defensive security work that process trillions of tokens a month. Yes, trillions!
With all that work, it's clear to us what the future of security looks like.
Hacking was once a bespoke skill, similar to hardcore software engineering. Software engineering is currently moving from offices to factories. Security, too, will go through the same transition.
Last month, Anthropic released a report on how they caught different threat actors abusing Claude. One of them stood out to me. A team in China built an automated exploit factory. A factory that autonomously grinds through vulnerabilities in targets and exploits them. Their targets included security products, network appliances, and government organizations.
We've spent enough time building systems adjacent to that to know what it takes, and the shocking part is the economics. How little it costs to breach companies that you and I rely on.
All this to say, security is largely becoming an economics problem. Reasonably capable models and harnesses, with enough compute, can find ways to steal your data, damage your reputation, and sometimes even take your money. And the cost to achieve that is dropping every day.
When that happens, the right question to ask is: what's the cost, and what's the lowest cost to achieve an outcome reliably? You want to find a stack with the best combination of capability and cost. This is what people call 'Pareto-optimal'.
Once you know that, it's all about scaling. Scaling to trillions of tokens a month, then a week, then a day. Eventually, trillions of tokens a second.
To do this, you need to own and control the entire stack: the models, the harnesses, the context, and even the flow of tokens. If you do the math here, the economics start looking insane.
What does that look like?
1. Building evaluations on cyber tasks that you and your customers care about.
2. Building a data pipeline of unique real-world data with signal attached to it.
3. Building a harness that can self-improve for each specific organization or customer.
4. Building a post-training loop that can take those learnings and improve the models. Models that are better, faster, and cheaper.
5. Owning your inference pipelines and controlling the flow of tokens so you're maximizing the value produced in every GPU cycle.
On evals: if you're a company building AI products for customers, you have to build your own. There's tremendous alpha in having internal evals corresponding to real work. In our case, these are evals for offensive and defensive security work.
A lot of public evals are bad for measuring the work you actually care about. Public evals are typically from academics or data companies.
Academics have limited budgets and limited access to proprietary data that measures real economic work. Data companies build evals and also sell you corresponding data (RL gyms, traces, pre-training data, etc.) that helps your next model “juice up” its score.
This is the dirty secret, and also why models can do extremely well on public evals but extremely poorly on things you care about. This is also why there's tremendous alpha in private evals: you know which models are the best for your use cases.
In cyber, some of the popular public evals are ExploitGym and ExploitBench. We think these evals do not correspond to the real security work our customers need. A typical challenge in ExploitBench is to take Chrome's javascript engine V8 with a known patch and find a way to exploit the vulnerability.
That is far, far removed from the security issues in everyday applications that you and I use, and applications built by our customers.
We could've benchmaxxed on these evals, but we didn't. And you should be wary of people using these evals to advertise their models and harnesses. Measure things yourself and see how they apply to your use cases.
The future: Opus is a model that many people found to be a reliable workhorse. For me, Opus 4.1 was the first model that was functional and could get work done.
These days, the flash models from Qwen, DeepSeek, MiMo and GLM can be categorized as workhorses too. That, combined with owning the post-training loop means the economics start looking insane. You can get frontier performance in specific domains for a small fraction of the price of Opus.
With the right data pipelines and post-training loop, you now have a durable strategy to keep improving and stay state-of-the-art in your vertical.
That's the bet behind apex-flash-1, our post-train of glm-5.3-flash, and we're already preparing for longer training runs.
today we're releasing apex-flash-1, our first open-weights model for security research post-trained on real vulnerabilities we found and got paid for.
we are also releasing the abliterated variant for researchers who want fewer refusals in their own authorized workflows.
how we built it: cantina.security/apex-flash
Matías Insaurralde retweeted
We’ve confirmed a KVM 0day through our Vercel Sandbox bounty program. Affecting the industry’s gold standard solution for Linux virtualization.
2026 is wild! Thankful to Paulos and other researchers helping us make the most secure sandbox for agents. Full writeup coming.
Matías Insaurralde retweeted
🔴 Elecciones municipales: 14 candidatos figuran como deudores alimentarios
📌Dos candidatos a intendentes y 12 a concejales titulares aparecen en el Registro de Deudores Alimentarios Morosos (REDAM).
🔸La normativa no les impide competir, pero sí genera consecuencias en determinados trámites estatales.
🔴 Canal de WhatsApp: whatsapp.com/channel/0029VaC…
abc.com.py/politica/2026/10/…
Matías Insaurralde retweeted
El trabajo que hace @angatupyrytau es serio, no es desde el punto de vista político partidario, sino técnico.
Los datos que estuvieron circulando en los medios sobre irregularidades en el padrón no son mentira querido @enriquevp y si @angatupyrytau anduvo preocupado por la seguridad de las urnas es porque la calidad del software es de trabajo práctico escolar, deja mucho que desear.
Matías Insaurralde retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Matías Insaurralde retweeted
We added Xiaomi MiMo 2.6 Pro to our cybersecurity benchmark and safe to say it cooked! 🔥
- It found 25/32 vulnerabilities pass@3, matching GPT-6.1 Sol, Kimi K3 and GLM 5.3 and just one behind Opus 5
- It scored 88% precision, ahead of all Qwen and DeepSeek models
- One run found 64.6% on average; combining three got it to 78.1%
All that for just $11.79 per CVE found! 🤯
Los datos que estuvieron circulando en los medios sobre irregularidades en el padrón no son mentira querido @enriquevp y si @angatupyrytau anduvo preocupado por la seguridad de las urnas es porque la calidad del software es de trabajo práctico escolar, deja mucho que desear.
unavailable
Matías Insaurralde retweeted
Estimado @enriquevp creo que te informaron mal. yo no hable de fraude.
Sí afirme que 1 de cada 6 empadronados es fantasma.
Para que el @TSJE_Py entienda que DATO mata RELATO
medio difícil que el #Dr. te pueda leer porque le tenés más #bloqueado que la muralla que hizo Trump con los mexicanos
Matías Insaurralde retweeted
A full las irregularidades🗳️
Un hombre figura tres veces en el padrón electoral, con números de cédula distintos y, para colmo, los tres registros aparecen habilitados para votar en Capiibary.
extra.com.py/actualidad/un-t…
Matías Insaurralde retweeted
A big problem with cyber open models is economics.
For individuals and small security teams, running something GLM-5.3-class at serious scale just doesn’t make sense. If you’re a tech giant, spending hundreds of thousands to keep frontier-level inference on might make sense. For a small company or an independent defender, it isn’t sustainable.
That’s why defenders need broader access to frontier cyber models.
Giving the strongest cyber reasoning only to a handful of large companies improves security for those companies, but leaves thousands of smaller organizations defending themselves with significantly weaker tooling while the attackers are going full power on them.
And when those smaller companies get pwned, the blast radius still reaches the same users, vendors, supply chains, and ecosystems the large companies are trying to protect.
Cyber defense is an ecosystem problem. Frontier defensive capability needs to be accessible beyond Big Tech.
Matías Insaurralde retweeted
today we're releasing apex-flash-1, our first open-weights model for security research post-trained on real vulnerabilities we found and got paid for.
we are also releasing the abliterated variant for researchers who want fewer refusals in their own authorized workflows.
how we built it: cantina.security/apex-flash
Matías Insaurralde retweeted
Our last post was about attacking. Point a capable enough model at your own codebase, and it finds real vulnerabilities, including the ones your scanners, your tests, and your reviewers all missed. We spent a month doing that to ourselves, and it worked.