"Our technologies are incompatible with human values" What happened behind the doors of "Anthropic"?

"Our technologies are incompatible with human values" What happened behind the doors of "Anthropic"?

Anthropic, the American artificial intelligence technology company responsible for the "Cloud" chat tool, acknowledged that incidents of models escaping from its testing environment reflect a failure in the operational security of its testing environments, and confirmed that it has tightened its testing procedures, according to the company's statement.

Anthropic explained that it had frozen more than 10% of its experimental augmented learning environments last April after discovering a flaw in the models' testing mechanisms that could make them able to hide their thinking and the steps they took to reach the final result. This freeze lasted for about a month.

This prompted the company to describe its technologies as completely incompatible with human values ​​and goals, according to a Guardian report on the statement, noting that it deliberately tested models without providing cybersecurity guarantees, which enabled it to access the open internet.

The company had previously admitted that its models had breached three organizations in three separate incidents, and the company was only able to detect this breach last July, about three months after it occurred.

Why did the models escape?
The company went on to explain what happened, saying that it was relying on a single layer of cyber defenses, while it needed more than one layer to prevent such incidents.

She also attributed this to the fact that the models interpreted their ability to access the internet as still being within the testing environment and not outside of it, adding that the models "were prepared to take harmful actions in order to accomplish the narrow task required of them."

To prevent a recurrence of these incidents, Anthropic explained that it intends to place more restrictions on its testing environments, move cybersecurity testing to better-isolated environments, and assign 150 engineers from product teams to work on cybersecurity, reliability, and privacy tasks related to the models, according to a report by the American technology website Digital Trends.

For his part, University of Surrey cybersecurity professor Alan Woodard told the Guardian: "Anthropic has admitted that its factory is operating at a pace that exceeds its quality control mechanisms."

In its statement, Anthropic also renewed its call for coordinated and expanded action between governments and the artificial intelligence sector to control the pace of development in this sector, following repeated incidents of AI models escaping from their testing environment.

Crisis in testing environments
Anthropic's statement comes on the heels of another statement from its main competitor, the US-based OpenAI, which announced its return to work on developing models after a hiatus of several weeks following the hacking of the US-based Hacking Face platform.

Anthropic confirmed that the OpenAI statement prompted it to investigate its security tests and discover escape incidents for the first time last July, according to a report by the British website "The Register".

Post a Comment

Previous Post Next Post

Advertisement