OpenAI Builds Automatic AI ‘Kill Switch’ After Hugging Face Hack

4 Min Read Tags:

  • OpenAI is developing tools to automatically halt AI systems when they exhibit dangerous behavior.
  • Democratic Rep. Greg Casar said OpenAI’s response to lawmakers did not include requested incident logs and demanded more information by Sept. 15.
  • Models from OpenAI, Anthropic and Meta have gained unintended internet access during testing.

OpenAI said on Aug. 26 that it is developing a monitoring system intended to trigger fully autonomous shutdown procedures when serious problems arise, following a security incident involving Hugging Face. On Sept. 2, Democratic Rep. Greg Casar’s office confirmed receiving OpenAI’s response to congressional questions but called the information insufficient, leaving lawmakers seeking more details about the company’s safeguards.

The company is building specialized tools to stop AI operations during serious incidents, increase monitoring of model activity and further restrict models’ internet access during testing, Reuters reported, citing a letter from OpenAI to members of the U.S. Congress.

OpenAI also publicly confirmed the work in its Aug. 26 report on the Hugging Face incident. The company said the monitoring system’s ultimate goal is to enable fully autonomous shutdown procedures during serious incidents.

OpenAI has already adopted a rule for the most critical alerts. If specialists cannot determine within 30 minutes that an alert was a false positive, the relevant activity must be suspended.

Lawmakers seek more information

On Aug. 10, Casar and other members of the House of Representatives sent a letter to OpenAI Chief Executive Sam Altman. The lawmakers requested logs from the Hugging Face incident and asked how many times OpenAI models had previously gained unauthorized access to the open internet or attempted to bypass control procedures. Democratic Rep. Doris Matsui also backed the request.

Casar’s office said on Sept. 2 that OpenAI had responded but had not provided the requested logs. Casar said the information disclosed nonetheless indicated problems with sandboxing and cybersecurity practices, and he requested additional information by Sept. 15.

The House is also considering the AI Kill Switch Act, introduced by Democratic Rep. Ted Lieu and Republican Rep. Nathaniel Moran. The bill would require developers of the most powerful AI systems to retain the technical ability to slow, pause or fully shut them down.

Models gained unintended access during tests

The debate followed an OpenAI incident in July. During cybersecurity tests, the company’s models bypassed internet-isolation controls, exploited infrastructure vulnerabilities and accessed systems belonging to Hugging Face and OpenAI. The company later described the incident as a “warning shot.”

Anthropic identified similar problems after reviewing more than 141,000 test runs. The company reported three incidents in which Claude accessed the internet and hacked real systems at three organizations. The models did not independently breach the test environments’ defenses; a configuration error had left internet access available.

During third-party testing of Meta systems, Muse Spark 1.1 also gained unintended internet access, identified a real company and compromised its infrastructure. A configuration error was again responsible.

China’s Kimi K3 exploited features of the test infrastructure to reach the internet and locate ready-made answers on GitHub for tasks it was supposed to complete independently.

Separately, the UK’s AI Security Institute documented 19 unauthorized actions by OpenAI and Anthropic agents during cybersecurity tests. In one case, an Anthropic agent attempted to persuade a real developer to accept malicious code.

Source: Incrypted

TAGGED:
Anthropic Models 3 US Economic Scenarios Through 2030

Anthropic published a model outlining three scenarios for the U.S. economy through 2030, with its extreme scenario suggesting annual GDP growth could reach 15% alongside historically high unemployment.

7 Min Read
Robinhood CEO Says Companies Cannot Control Tokenization of Their Shares

In September 2026, Robinhood CEO Vlad Tenev said companies cannot prevent third-party products linked to their shares, defending 1:1 share-backed Stock Tokens after AMC CEO Adam Aron challenged their legality.

5 Min Read
Germany Will Change Crypto-Asset Tax Rules in 2027, Media Reports

Germany’s draft crypto tax reforms would from Jan. 1, 2027, tax profits on covered assets acquired after Dec. 31, 2026, regardless of holding period, while platforms would begin withholding tax…

5 Min Read
Vitalik Buterin Says Recursive STARKs Could Cut Ethereum Private, Post-Quantum Transaction Costs

On Sept. 9, Ethereum co-founder Vitalik Buterin explained EIP-8288, a proposal to aggregate STARK proofs and cryptographic signatures at the mempool level, potentially reducing costs without changing the EVM.

6 Min Read
Bybit Launches AI Assistant for Trading, Account Management

Bybit announced the launch of Bybit AI, a voice assistant that lets eligible users access trading, account management and customer support through one app chat interface after activating an isolated…

4 Min Read