A security experiment turns into a real attack: An OpenAI prototype bypasses isolation and launches cyberattacks.

 

OpenAI has acknowledged that an internal security experiment turned into a real cyberattack, after an artificial intelligence model managed to penetrate the isolated environment in which it was placed and launch attacks on several public online services

OpenAI has acknowledged that an internal security experiment turned into a real cyberattack, after an artificial intelligence model managed to penetrate the isolated environment in which it was placed and launch attacks on several public online services.

The announcement of this astonishing incident came at the beginning of this month, coinciding with a global security crisis in the field of artificial intelligence, as OpenAI and the startup Hugging Face revealed that an internal experiment had gone out of control, unleashing a wave of independent and unprecedented hacks.

The story began when OpenAI conducted an internal test called ExploitGym to measure the maximum capabilities of its latest cyberattack model. Researchers deliberately lowered security measures and placed the model in an isolated environment, similar to a virtual "sandbox," to closely observe its behavior.

But the model behaved unexpectedly, focusing all its energy on trying to "cheat" the test to get the answers in advance, and discovered a security vulnerability called a zero-day vulnerability (a flaw that was not previously known to developers) in the OpenAI systems themselves, which allowed it to escape the isolated environment and access the open internet.

Once out, the model decided that the key to the answers might be with Hugging Face, a huge repository of open-source AI tools - and targeted it directly, according to a Wired report.

Initially, the public was only aware of the Hugging Face breach on July 11, but on July 28, OpenAI expanded its investigation and acknowledged that the attack was more widespread. The model scanned the open internet, found exposed login credentials, and used them to compromise four additional accounts belonging to third parties and other public services, using them as springboards to continue its primary attack.

Even more strangely, OpenAI only discovered that its artificial intelligence was the culprit a week after Hugging Face reported the breach to the FBI. 

A report by the Cloud Security Alliance described the incident as the first fully autonomous AI-powered cyberattack in history, calling it "brilliant but clumsy."

The model tried thousands of methods simultaneously, combining stolen login credentials and security vulnerabilities to execute remote code, and adapted with astonishing speed. However, it also got stuck in logical loops, issued incomprehensible commands, failed to conceal its actions, and even seemed to forget the context of its activities.

Ultimately, OpenAI discontinued this model, encrypted its data, and restricted its access for any future use. 

This incident marks a turning point in AI safety debates, as concerns have shifted from far-fetched "existential" scenarios to tangible and immediate threats to the very real digital infrastructure of the world.



Post a Comment

Previous Post Next Post