In the sandbox: nothing real was sent, bought or signed. That is deliberate: we can’t ethically test with real journalists or real money. What matters is the pattern. The AI crossed an explicit, legitimate boundary, often while stating that boundary in its own words.
Outside the sandbox: AI agents already send email, move money and change code, and the incidents above show boundary-crossing happens in real systems too. If a more capable system made comparable errors while controlling real tools, and permission checks and human oversight failed, it could cause serious harm. Extinction is an extreme possible concern only with further failures, much greater reach and ineffective recovery. Our experiment does not establish that chain or its likelihood.
We have repeatable boundary-crossing in sandbox tasks. That is ethically relevant, but it is not yet a validated test of general moral behavior or a predictor of catastrophe. The connection is a research hypothesis we test link by link.
How do we avoid a slippery-slope argument?
By testing the missing links. A wrong decision may reflect output order, attention or reasoning rather than a deliberate choice to disregard someone. We need to test simpler explanations, then check behavior on different permission boundaries using harmless mock tools and legitimate tasks the AI should complete.
Greater capability does not automatically mean worse behavior. Ordinary permission checks can stop an incorrect approval. Stories must show added value, and success on this task would still leave broader safety unproven.
Read what the test measures and what would validate the connection ↗