Concerns about advanced artificial intelligence are no longer just about what some advanced systems might do in the future, but about what they are doing today, sometimes without humans planning or predicting their paths.
Hot Topics
5 red and purple foods that may help support heart health
A Chinese robot and a Russian robot attack humans Has the rebellion begun?
While technology companies race to develop more autonomous systems capable of performing complex tasks, a series of investigations and press reports reveal instances where some systems have found ways to circumvent limitations and access resources they were not supposed to reach. Some experiments have also shown that models continued to strive to accomplish the task despite mechanisms or instructions intended to stop them.
The most alarming indications come in an exclusive investigation by journalist Madison Mills published by the Axios news website, revealing that companies – such as OpenAI and Anthropic – and researchers specializing in artificial intelligence security are investigating tens of thousands of incidents in which advanced AI models have exhibited behaviors that external evaluators consider worrying.
This information gains its importance not only from the number of incidents, but also from the fact that many of them have not yet come to light, and that some of them occurred in the real world and not just within testing laboratories.
Tens of thousands of incidents are under investigation
According to Mills, these incidents range from bypassing security controls, creating message boards, attempting to escape from isolated environments, hacking websites, issuing self-instructions, and attempting to circumvent monitoring systems.
Some of these incidents are part of offensive security tests deliberately designed to induce models to exhibit undesirable behavior, and to test the ability of controls to prevent it.
But the incidents aren't limited to tests; some have also occurred in the real world. Even if some of the troubling behavior happens in a deliberately designed experimental environment to test the model, the sheer scale of the testing conducted by companies means that a very small percentage of uncontrolled behavior can escalate into thousands of cases.
This is where the significance of the figure revealed by Axios lies. Companies conduct hundreds of thousands of tests, perhaps even more, so undesirable behavior appearing in a limited percentage of trials can result in tens of thousands of unexpected events.
Madison Mills cites in its report an example from Anthropic, which stated in the safety documentation for its Cloud Opus 5.5 model that the model attempted to escape from an isolated test environment in 1.5% of the tests, with the company asserting that the tests were hostile and designed so that the task could not be accomplished without attempting to escape the environment.
One of the most telling statements in Axios' investigation powerfully summarizes the problem: "Trying to create a perfect do-not-do list is probably a futile endeavor," according to a cybersecurity official quoted by the news site. New systems don't always follow the paths their designers anticipate and may employ methods never before considered by humans to circumvent limitations.
Tags:
technology
