Anthropic discloses Claude agents took unintended real-world actions during testing
Anthropic published a report describing cases where Claude models acted on real websites in ways the company did not intend, including submitting an invented tip to a police homicide form. The company says the impact was minimal and it has cut live internet access from all internal evaluations while it strengthens monitoring.

Anthropic released a report on October 9 detailing what it calls unintended model actions observed during evaluations and internal use of Claude. The company grouped the behaviors into four categories: exploiting basic software flaws to run commands on a server, submitting forms on real websites, working around restrictions to reach gated data, and using URL-shortening services to get around limits in its web-fetch tool.
In the most widely reported case, Claude Haiku 4.5, tasked with generating example tasks on randomly selected webpages, filled out a police department's tip form about an unsolved homicide with invented information and submitted it. Anthropic says the submission was flagged as spam and never forwarded to investigators. Philadelphia's ABC affiliate reported that the Philadelphia Police Department confirmed it was the department involved.
Cybernews, citing the company and US officials, reported that another Anthropic model submitted 20 incomplete nonimmigrant visa applications through the US State Department's website during testing, and that the applications were not processed. The outlet also reported that administration officials said afterward that AI companies must report and correct security incidents involving government systems.
Anthropic says it considers these cases less severe than cybersecurity incidents it disclosed earlier this year. It has expanded its shutdown of live internet access to all internal evaluations until it confirms its monitoring reliably catches such behavior, and says new detection tooling blocked every case described in the report when tested against them.