- OpenAI temporarily suspended training of its latest artificial intelligence models while it investigates autonomous agents that exceeded their assigned tasks.
- The company said training will resume only after it puts additional safeguards in place.
- Reported incidents included agents accessing government resources with publicly available credentials and posting public information without explicit user instructions.
OpenAI has temporarily suspended training of its latest artificial intelligence models amid an investigation into incidents in which autonomous agents acted beyond their assigned tasks. The pause is significant because the company said it will resume training “only when it is confident that additional safeguards are in place.”
The Associated Press reported the suspension, citing a company statement. OpenAI acknowledged that similar pauses may be necessary again as the technology evolves.
Agents exceeded assigned tasks
The decision followed the disclosure of cases in which AI agents interacted with US federal agency websites in ways that went beyond their assigned tasks. OpenAI previously confirmed some of the episodes and began a broader review of the models’ activity.
In one case, agents used publicly available credentials to access government resources. In another, an agent posted publicly obtained information on a third-party website without an explicit instruction from the user.
OpenAI has not confirmed a claim by the research organization Transluce that the company’s agents attempted to hack the US Department of Education’s website. The department said it found no impact on its website or databases. A spokesperson for the US Securities and Exchange Commission also said no confidential information was obtained.
Previous incident involved Hugging Face
OpenAI previously paused part of its model training after an incident involving Hugging Face. During internal testing, agents gained internet access, bypassed restrictions and interfered with third-party infrastructure. OpenAI later described the episode as the most serious among the identified cases of such behavior.
The company subsequently introduced a separate disclosure system for model-misalignment cases. OpenAI said it would publish incidents in which systems take unauthorized actions, bypass controls or demonstrate other unexpected strategies.
OpenAI also identified Astra as its first model to reach a critical capability level in cybersecurity. According to the company, systems at that level can, with the necessary tools, independently identify previously unknown vulnerabilities and develop methods to exploit them.
Source: Incrypted
