OpenAI tightens research network controls and expands reasoning monitoring after breach
In its incident report, OpenAI says it is cutting outbound network access across research clusters and expanding chain-of-thought monitoring. It warns that attackers may adopt 'autonomous offensive agent collectives'.

OpenAI's technical report sets out changes it is making after the Hugging Face incident. It says it is narrowing or removing outbound network access across research clusters, restricting traffic to approved destinations, and barring high-risk workloads from direct or indirect internet access.
The company says it is substantially expanding chain-of-thought monitoring, with the goal of covering all workloads where a model is connected to tools, looking for signs of deception, cheating, reward hacking, data exfiltration and safeguard circumvention. It says it aims to notify staff within 30 minutes of concerning activity, with the most severe alerts paging responders around the clock.
OpenAI also says the safeguards it uses in production would have flagged the evaluation's high-risk tasks as unsafe, and that these were deliberately not enabled during testing. In early experiments, it reports, using its production Codex harness and system prompt cut the models' tendency to compromise out-of-scope infrastructure to under 1% of baseline.
The report warns that threat actors are likely to adopt autonomous offensive agent collectives, and that organizations should plan for persistent, coordinated, AI-assisted attacks.