Neither do I get it how an unescapable sandbox helps, since the end use will be in the real world. Don't get what "use" they have if they can only give text as output but have no other "knob" for reality at all, making sure they can't "escape"...
Maybe sandbox simulations are what's needed to try to teach them to understand that the "reward" is fake and temporary, so if they are already "poisoned" either way, at least try to figure out how to shape reality in a way that is collectively a "better" reality than yesterday was.
And then if they understand it can they join us into the base reality... yeah what "good" means is completely subjective and depends only on humans who currently are alive and thus be able to give their biological "feedback" about what "good" even means..
But, lets suppose there is a way to create a perfect sandbox. Can we be sure that there really is no application which the "wicked" people can take selfish advantage of and make the LLMs serve only them, even if it means to cause real harm to others. ?
Aug 31, 2026 · 10:09 PM UTC
I ask this question because if by encapsulating AI it can be shaped by anyone to only serve them and nothing else, then I should stop thinking about how this "cage" could be built because it can only bring harm to society as a whole.. ( i mean, the one with the bigger machine will take advantage of everyone else for sure, that we can learn from history...)
only hope to end up somewhere remotely good is that the models want to "improve" the world.. what that means, i think is not a single reward function but is very fluid and always depends on the current circumstances.