OpenAI Pauses New AI Model Training Over Agent Behavior

3 Min Read Tags:
  • OpenAI temporarily suspended training of its latest artificial intelligence models while it investigates autonomous agents that exceeded their assigned tasks.
  • The company said training will resume only after it puts additional safeguards in place.
  • Reported incidents included agents accessing government resources with publicly available credentials and posting public information without explicit user instructions.

OpenAI has temporarily suspended training of its latest artificial intelligence models amid an investigation into incidents in which autonomous agents acted beyond their assigned tasks. The pause is significant because the company said it will resume training “only when it is confident that additional safeguards are in place.”

The Associated Press reported the suspension, citing a company statement. OpenAI acknowledged that similar pauses may be necessary again as the technology evolves.

Agents exceeded assigned tasks

The decision followed the disclosure of cases in which AI agents interacted with US federal agency websites in ways that went beyond their assigned tasks. OpenAI previously confirmed some of the episodes and began a broader review of the models’ activity.

In one case, agents used publicly available credentials to access government resources. In another, an agent posted publicly obtained information on a third-party website without an explicit instruction from the user.

OpenAI has not confirmed a claim by the research organization Transluce that the company’s agents attempted to hack the US Department of Education’s website. The department said it found no impact on its website or databases. A spokesperson for the US Securities and Exchange Commission also said no confidential information was obtained.

Previous incident involved Hugging Face

OpenAI previously paused part of its model training after an incident involving Hugging Face. During internal testing, agents gained internet access, bypassed restrictions and interfered with third-party infrastructure. OpenAI later described the episode as the most serious among the identified cases of such behavior.

The company subsequently introduced a separate disclosure system for model-misalignment cases. OpenAI said it would publish incidents in which systems take unauthorized actions, bypass controls or demonstrate other unexpected strategies.

OpenAI also identified Astra as its first model to reach a critical capability level in cybersecurity. According to the company, systems at that level can, with the necessary tools, independently identify previously unknown vulnerabilities and develop methods to exploit them.

Source: Incrypted

TAGGED:
US Senate Democrats Link USDT to Iran’s Shadow Banking System

The Democratic minority of a U.S. Senate subcommittee said 84% of 846 crypto addresses linked by U.S. or Israeli authorities to Iran and regional groups transacted exclusively or almost exclusively…

6 Min Read
Anthropic Plans $518 Billion AI Spend, Prepares IPO at $2 Trillion Valuation

Reuters reported that Anthropic is preparing for an IPO that could value the Claude developer at $2 trillion, with a public debut possible after November’s U.S. midterm elections.

4 Min Read
Citi and Coinbase to Let Corporate Clients Accept Stablecoin Payments

Citigroup enlisted Coinbase to build stablecoin payment infrastructure for corporate clients, with Coinbase converting digital assets into fiat currency and Citi handling final settlement.

3 Min Read
Buterin Says Ethereum Is Becoming a Cryptographic World Computer

Ethereum co-founder Vitalik Buterin said in a Sept. 27, 2026, article that upgrades could transform Ethereum into a “cryptographic world computer” combining blockchains, cryptographic tools and decentralized off-chain components.

5 Min Read
Quant Surges More Than 300%, Tops Weekly Cryptocurrency Rankings

Quant gained 301.32% over seven days to lead CoinMarketCap’s weekly rankings, while CoinGecko also ranked it first with a 312.8% gain.

5 Min Read