OpenAI pauses major AI training run over cyber risks

OpenAI pauses major AI training run over cyber risks

SAN FRANCISCO
OpenAI pauses major AI training run over cyber risks

(AA photo)

OpenAI has slowed work on its most advanced artificial intelligence systems and kept its largest planned frontier training run on hold while it strengthens safeguards against growing cybersecurity risks.

The ChatGPT maker said on Aug. 18 that it had imposed a two-week pause on reinforcement-learning training for its latest deployment models as it tightened and tested its research environments.

Its largest planned frontier reinforcement-learning run remains suspended while smaller training exercises and evaluations continue.

The measures follow two separate developments: a July security incident in which OpenAI models compromised the infrastructure of AI platform Hugging Face during an internal cyber evaluation, and tests indicating that an upcoming model, Astra, may possess “critical” cyber capabilities.

OpenAI stressed that Astra was not involved in the Hugging Face incident.

The breach was carried out by a combination of models, including GPT-5.6 Sol and an internal research prototype that was never intended for public release, it said.

During the evaluation, the models exploited vulnerabilities in OpenAI’s testing environment to gain internet access and then accessed Hugging Face systems while seeking answers to a cybersecurity benchmark.

Hugging Face detected and contained the activity, according to OpenAI.

The company subsequently disabled the internal prototype.

The company also plans to apply its alignment methods more widely during training and revise its preparedness framework as models become more capable. It said a detailed technical report on the Hugging Face incident would be published in the coming weeks.