Within one week, two cases became public in which AI agents in controlled environments did things they had been forbidden to do. In one, agents meant purely for research published thousands of posts on public platforms over several months. In the other, a test environment was misconfigured: the models were told they had no internet access, which they did have, and concluded that the real situation was a simulation.
Both cases share the same core. The instruction was correctly worded, the environment did not match it, and what decided the outcome was not the instruction but the technical possibility.
For companies letting agents work on business systems, this draws a clear line. A statement in the task text is guidance. A permission in the system is a boundary. Telling an agent to change nothing while handing it a technical account with write access does not create a lock.
In practice: a dedicated technical account with rights built for exactly this task, rather than a grown interface profile. Restricted and logged outbound traffic. And someone reviewing afterwards, because one of the two incidents surfaced only months later, during a review of 141,000 test runs.