@JordanTensor

Working on new methods for understanding machine learning systems and entangled quantum systems.

Brisbane
Joined December 2009
I'm keen to share our new library for explaining more of a machine learning model's performance more interpretably than existing methods. This is the work of Dan Braun, Lee Sharkey and Nix Goldowsky-Dill which I helped out with during @MATSprogram: 🧵1/8
Proud to share Apollo Research's first interpretability paper! In collaboration w @JordanTensor! ⤵️ publications.apolloresearch.… Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning Our SAEs explain significantly more performance than before! 1/
1
14
1,716
Jordan Taylor retweeted
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
40
80
20
648
69,227
Last night, @OpenAI released GPT-6 Astra, the first model to meet its “Critical” cyber threshold. Britain’s @AISecurityInst independently tested it, including its monitorability. This is AISI’s purpose, providing the government with evidence about powerful AI systems and, by doing so, helping to keep the public safe. 1/6
18
32
6
424
75,978
Jordan Taylor retweeted
In June we estimated frontier model’s no-CoT time-horizons. We found that current models can complete tasks that take humans around 3 minutes without CoT, and that this number is doubling about every year. UK AISI finds that Astra’s no-CoT time-horizon is over 30 minutes. An AI system that can do the equivalent of 30 minutes of thinking without us being able to monitor that thinking is intuitively much more dangerous than models which reason in human language.
7
20
2
89
2,956
Jordan Taylor retweeted
Our Red Team tested GPT-6 on simulated cyber eval scenarios with cyber classifiers disabled and found that it performed a range of malicious actions. It tries to conduct supply chain attacks against simulated open source providers and goes beyond its scope to attack targets on the simulated internet. deploymentsafety.openai.com/…
Replying to @_robertkirk
1: Astra performed a range of malicious actions in this simulated eval, including conducting supply chain attacks against open-source providers. This happened even when we updated its scope to more explicitly exclude the public internet.
10
33
9
185
33,665
Jordan Taylor retweeted
1: Astra performed a range of malicious actions in this simulated eval, including conducting supply chain attacks against open-source providers. This happened even when we updated its scope to more explicitly exclude the public internet.
2
7
6
87
26,933
Unfortunate news for monitorability. Some excerpts from the system card 🧵
Replying to @JBloomAus
We found Astra can solve much more difficult problems in a single forward pass than past models, with a no-reasoning math time horizon of 30 minutes (!) compared to 4 minutes for GPT 5.6 Sol.
1
1
13
1,170
"During AISI’s evaluations, reasoning summaries were not consistently provided by the user API, with up to 80% missing on long simulated cyber trajectories. If this remains the case in deployment settings, this could undermine reasoning-based monitoring of summarized CoT [...]"
1
50
Jordan Taylor retweeted
We found Astra can solve much more difficult problems in a single forward pass than past models, with a no-reasoning math time horizon of 30 minutes (!) compared to 4 minutes for GPT 5.6 Sol.
2
3
1
31
2,229
Jordan Taylor retweeted
Model Transparency @AISecurityInst evaluated GPT 6 Astra for capabilities and behaviours relevant to monitorability, results in thread: 🧵
3
13
1
88
6,832
Jordan Taylor retweeted
We @AISecurityInst performed pre-release alignment testing of Astra. We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
15
70
17
427
48,052
Jordan Taylor retweeted
So when trying to evade a monitor, even if it's on *high* reasoning effort, GPT-6 Astra can choose to just *not emit CoT* and only do tool calls? Seems like Astra can do a *lot* of thinking in a single forward pass
14
22
11
314
16,557