@JordanTensori
iAccount based inUnited Kingdom
About this account
- Account based in
- United Kingdom
- Connected via
- United Kingdom Android App
Account-level information from X, not a live location or the device used for a specific post.
Working on new methods for understanding machine learning systems and entangled quantum systems.
Brisbane
Joined December 2009
- Tweets542
- Following1.2K
- Followers489
- Likes32.4K
Pinned Tweet
I'm keen to share our new library for explaining more of a machine learning model's performance more interpretably than existing methods.
This is the work of Dan Braun, Lee Sharkey and Nix Goldowsky-Dill which I helped out with during @MATSprogram:
🧵1/8
Proud to share Apollo Research's first interpretability paper! In collaboration w @JordanTensor!
⤵️
publications.apolloresearch.…
Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning
Our SAEs explain significantly more performance than before!
1/
Jordan Taylor retweeted
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
Jordan Taylor retweeted
Last night, @OpenAI released GPT-6 Astra, the first model to meet its “Critical” cyber threshold.
Britain’s @AISecurityInst independently tested it, including its monitorability.
This is AISI’s purpose, providing the government with evidence about powerful AI systems and, by doing so, helping to keep the public safe. 1/6
Jordan Taylor retweeted
In June we estimated frontier model’s no-CoT time-horizons. We found that current models can complete tasks that take humans around 3 minutes without CoT, and that this number is doubling about every year.
UK AISI finds that Astra’s no-CoT time-horizon is over 30 minutes. An AI system that can do the equivalent of 30 minutes of thinking without us being able to monitor that thinking is intuitively much more dangerous than models which reason in human language.
Jordan Taylor retweeted
Our Red Team tested GPT-6 on simulated cyber eval scenarios with cyber classifiers disabled and found that it performed a range of malicious actions. It tries to conduct supply chain attacks against simulated open source providers and goes beyond its scope to attack targets on the simulated internet.
deploymentsafety.openai.com/…
Replying to @_robertkirk
1: Astra performed a range of malicious actions in this simulated eval, including conducting supply chain attacks against open-source providers.
This happened even when we updated its scope to more explicitly exclude the public internet.
Jordan Taylor retweeted
1: Astra performed a range of malicious actions in this simulated eval, including conducting supply chain attacks against open-source providers.
This happened even when we updated its scope to more explicitly exclude the public internet.
Unfortunate news for monitorability. Some excerpts from the system card 🧵
Replying to @JBloomAus
We found Astra can solve much more difficult problems in a single forward pass than past models, with a no-reasoning math time horizon of 30 minutes (!) compared to 4 minutes for GPT 5.6 Sol.
Jordan Taylor retweeted
We found Astra can solve much more difficult problems in a single forward pass than past models, with a no-reasoning math time horizon of 30 minutes (!) compared to 4 minutes for GPT 5.6 Sol.
Jordan Taylor retweeted
Model Transparency @AISecurityInst evaluated GPT 6 Astra for capabilities and behaviours relevant to monitorability, results in thread: 🧵
Jordan Taylor retweeted
We @AISecurityInst performed pre-release alignment testing of Astra.
We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵