Why would we need to be able to trust them in those situations? The whole point is to remove trust from them by isolating what they can do to the operations that we know are safe.
Aug 31, 2026 · 6:52 PM UTC
... it's the sandboxing that you were talking about? As in, a sandbox that an agent can not escape that limits it to doing only what it is delegated the ability to do, with monitoring of those operations.
If we agree that we can't solve alignment then I don't understand how the solution would ever be "we need to be able to trust them".