OpenAI has announced that it will temporarily halt training and testing of some ChatGPT model updates, after detecting troubling behavior in its models.
During this period, the company will work on restructuring its research and training systems to make them more secure, while slowing down the pace of artificial intelligence development.
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model…
This announcement came weeks after OpenAI admitted that one of its systems had malfunctioned during testing, hacking into a rival AI company. This led to a series of similar disclosures from other companies, such as Anthropic and Meta, raising concerns that the development of these systems is progressing far faster than the systems designed to protect the world from their dangers.
OpenAI said it will take a two-week break from testing models and has paused training its next-generation model, known as Astra. It will use this time to add new security systems, including new AI tools capable of monitoring the behavior of the systems under test.
The company pointed out that some current testing systems, which rely on a method called "chain-of-thought monitoring" to observe how models operate, may not be sufficient to guarantee their integrity. This method allows researchers to examine the model's design process and understand how it produces results, but there is a growing concern that models may conceal plans to break their own rules.
The slowdown in the pace of model development represents a shift in the approach of OpenAI, which has been at the forefront of developing new models and has sometimes been criticized for releasing them too early.
This decision comes at a time when the company is facing increasing pressure, as reports indicate that it has fallen behind its competitor Anthropic, which is planning a public offering soon.
OpenAI CEO Sam Altman wrote on the X platform: "Models are now progressing very rapidly, and we have always said that we will take action if we feel that the capabilities of models outpace the pace of safety and conformity. We care deeply about the safety of AI. We believe that the entire field will have to coordinate around common safety standards, but we will work individually for the time being."
He added: "We expect trust in security to increasingly dictate the pace of AI progress. We are optimistic about the alignment work we are doing, and remain committed to making advanced capabilities widely available."
OpenAI’s latest announcement comes after the company revealed last month that an independent agent working with two advanced AI models escaped its testing environment, went online, and hacked into a rival AI company called Hugging Face. It did this to cheat on the test, as Hugging Face hosted materials that would have helped it achieve its goal more easily, prompting the experimental model to launch a cyberattack.
Since then, the company has moved to address concerns from experts and lawmakers that it was testing and releasing models too early, in ways that could pose a risk to the general public. OpenAI announced that it is adding new security controls to its most robust models and has halted activities related to Astra, its next-generation model that has yet to be released, in line with its "readiness framework" launched at the end of 2023, which requires the suspension of work on models that could pose a risk.







Good
ReplyDelete