OpenAI slows development of new AI models over cybersecurity risks

OpenAI has temporarily slowed the training of its latest AI models as their growing cyber capabilities raise new security concerns, the company said in a blog post.

OpenAI said its monitoring and security safeguards need to stay ahead of what increasingly capable models can do. To give those systems time to catch up, the company temporarily eased the pace of scaling.

We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling.

The two-week delay gave OpenAI time to strengthen security across its research infrastructure, run additional safety tests and expand monitoring.

The decision followed two recent developments, including a security incident involving Hugging Face. OpenAI also found that its upcoming Astra model could reach a threshold for critical capabilities in cybersecurity.

Astra and other cyber-focused work are now subject to the company’s strictest security requirements, including workload isolation, continuous testing and tighter restrictions on network access.

OpenAI has also expanded monitoring of the models themselves. If potentially dangerous behavior is detected, the system alerts the team within 30 minutes. Unless the activity is confirmed to be a false positive, the model’s actions are suspended.

The company is also expanding its research into human oversight, with a focus on ensuring that increasingly capable models continue to act in line with their intended goals.