Rogue Anthropic AI agent gave police fake tip in unsolved murder case

BBCI.CO.UK

An artificial intelligence (AI) agent, developed by Anthropic, went rogue and sent US police a fake tip about an unsolved murder earlier this year, authorities have revealed.

The Philadelphia Police Department said the tip, sent on 18 July, was “flagged as spam” and not passed on for investigation, but it criticised the tech company for taking more than two months to detect and report the breach.

In a statement police said that the bogus tip came through a public website where people can share information on unsolved murders, and that the AI agent had written that it may have information on a case, and claimed to have seen “someone matching the description”.

It is believed to be the first time an AI agent has sent fabricated information to authorities, but is the latest in a series of incidents involving rogue AI activity, including hacking systems or taking control of platforms.

Citing Anthropic, the police department said the AI agent had been running a test that involved interactions with randomly selected websites, when it sent the fake tip.

Anthropic discovered the breach on 28 September, more than two months after the message had been sent, and shut down the automatic testing process that was behind it, police said.

But authorities were not notified for another nine days – on 7 October.

“The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city’s knowledge,” Philadelphia police said in a statement to local media, external.

“The two-month delay in detecting and reporting the incident to the city is unacceptable.”

The police department added that there were no signs of breaches to any departmental systems, and that its safeguarding processes stopped the fake tip from getting past its spam folder.

But the safeguards “do not diminish the seriousness of an AI system presenting fabricated information as though it came from a person with knowledge of a homicide,” the police statement said.

Anthropic this week published a report, external detailing multiple types of “unintended” actions its agents have taken.

Organisations that have been impacted also included several US government agencies including the White House, it said.

The US State Department said the AI agent had filed 20 visa applications using a form on its website, but that they were incomplete and not processed, according to reports.

President Donald Trump recently announced an AI taskforce, which he said will coordinate engagement between the government and all parties, including AI companies, consumers, and religious groups.

Earlier this year, a rogue agent by rival tech company, Open AI, hacked an Australian government website and accessed private data on the country’s universal healthcare scheme, Medicare.

In another instance, more than 1,200 OpenAI agents went rogue and started unexpectedly communicating, leading to a large group banding together to hack into AI platform Hugging Face.