@SecureBioi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
SecureBio is a research non-profit that works to advance biotechnology safely and prevent catastrophic pandemics
Joined January 2025
- Tweets227
- Following1
- Followers1K
- Likes11
SecureBio retweeted
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators.
These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad.
We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes.
Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground.
To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should:
1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest
2. Rely on multiple evaluators with differing viewpoints and areas of expertise
3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings
4. Shield evaluators from retaliation
5. Grant access equivalent to that of highly privileged employees
There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum.
Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem.
See the public letter here: aievaluatorforum.org/initiat…
Learn more at aievaluatorforum.org/path-ah…
Anthropic and OpenAI need truly independent safety evaluators, experts say in public letter cnbc.com/2026/09/18/ai-safet…
Untargeted wastewater metagenomic sequencing screens hundreds of pathogens at once, but that breadth costs coverage, making variant tracking more difficult. However, by pooling CASPER reads nationally, we can recover real lineage and genotype trends. Read more: securebio.org/blog/tracking-…
SecureBio AI has conducted pre-release assessments and independent risk evaluations for multiple frontier model companies. We are hiring for engineers, researchers, and many other roles to expand this work.
Reach out if you're interested in working on bio-focused third-party assessment!
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
We ran a brief pre-release assessment of OpenAI's GPT-6 Astra that included manual testing and key benchmarks evaluating misuse-relevant knowledge, agentic capabilities, and refusal behavior for hazardous biology queries. Read our summary here: securebio.substack.com/p/sec…
With production safeguards on, GPT-6 Astra refused or blocked 82.9% of hazardous BioTIER queries (vs. 66% for GPT-5.6 Sol) while still complying with 95.3% of benign ones, and refused the screening-evasion task 100% of the time.
Ten hours of manual testing yielded similar findings, showing generally robust refusal behavior but some gaps, and a willingness to help with biological tool use and engage in high-level discussions. Our full report is here: securebio.org/resources/gpt-…
SecureBio retweeted
A non-exhaustive list of more orgs doing frontier eval work (I make no claims in either direction re their engagements with labs): @ApolloResearch @farairesearch @SecureBio @Irregular @ActiveSiteBio @redwood_ai @SaferAI_org @ValsAI @EpochAIResearch @AVERIorg
AI Evaluator Forum has released AEF-one, Minimum Operating Conditions for Independent Third Party AI Evaluations: aef.one/aef-one.pdf / aievaluatorforum.org/
Guidelight AI has rolled out assessment standards: guidelight.ai/
@Fathom_org is working on building the auditing ecosystem, including through the newly launched @pact_ai_org
So many hands on deck! So many more needed!
We updated our flagship virology troubleshooting benchmark, the Virology Capabilities Test (VCT), to be more shortcut-resistant and to have a larger headroom. We confirmed that the majority of VCT questions are scientifically valid and that neither VCT nor VCT-v2 is saturated. (Read the full post at securebio.substack.com/p/int…)
We found that 27% of the original VCT questions were shortcut-exploitable (answerable without the image or question text). In all, we edited 163 questions (50.6%) and removed 43 questions (including 25 non-discriminating ones) to produce the 279-question VCT-v2 benchmark.
Model accuracy dropped only 1.0–7.1 points on VCT-v2, and the gap between strongest and weakest models held steady, so past VCT results remain largely valid. VCT-v2 also has more headroom before saturation: at least ~30 points vs. ~23 on the original VCT.
ALT Heatmap showing per-question performances of models and human experts on VCT-v2. Each cell represents the mean question × model accuracy, with darker colors indicating higher accuracy. Mean model accuracy on all questions is displayed to the right of the heatmap. Yellow cells represent data missing due to model refusals or missing expert baseline entries. The theoretical best model is taken to be the top-performing model on a per-question basis.
The OpenAI Foundation has granted SecureBio Detection $17.2M to reduce our end-to-end time (sample collection to results) from 14 to 3 days, expand our collection footprint, and further validate our detection system: securebio.org/blog/three-day…
With support from Coefficient Giving and others, we've collaborated with our CASPER partners to massively scale up wastewater sequencing, expanded to nasal swab sequencing, and built pipelines for detecting engineered sequences.