I don’t understand the idea that the solution to the HuggingFace incident is better sandboxing or monitoring These models are going to be deployed in production, with internet access and inevitably minimal monitoring! We need to be able to trust them in those situations
67
25
9
438
35,179
Why would we need to be able to trust them in those situations? The whole point is to remove trust from them by isolating what they can do to the operations that we know are safe.

Aug 31, 2026 · 6:52 PM UTC

1
464
RelevantRecentLikes
What the fuck is "isolating what they can do to the operations that we know are safe" but "solve the alignment problem"? We can't do it, and that's the whole point here.
1
1
42
... it's the sandboxing that you were talking about? As in, a sandbox that an agent can not escape that limits it to doing only what it is delegated the ability to do, with monitoring of those operations. If we agree that we can't solve alignment then I don't understand how the solution would ever be "we need to be able to trust them".
1
59