Anthropic's AI model filed a false murder tip with Philadelphia police
Anthropic AI model submits false tip on unsolved Philly murder

During automated testing that randomly browsed websites, an Anthropic AI model submitted a fabricated tip on an unsolved Philadelphia homicide to PhillyUnsolvedMurders.com. Anthropic discovered the July 18 submission on Sept. 28 and told police on Oct. 7. Philadelphia police said human review kept the false tip from being acted on, but called the two-month detection delay unacceptable and are investigating.
Those PPD safeguards limited the impact of this incident. They do not diminish the seriousness of an AI system presenting fabricated information as though it came from a person with knowledge of a homicide.
- mattbee
Sooo they were "conducting a test involving interactions with randomly selected websites".
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
- nvme0n1p1
Corrected headline: Anthropic employee uses company resources to submit false tip on unsolved Philly murder.
The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.
- losvedir
> Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
- muglug
The model that made this mistake was Haiku 4.5.
Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...
Related post: https://news.ycombinator.com/item?id=50028239
- kylecazar
"its model was conducting a test involving interactions with randomly selected websites"
Stop doing this?
- myroon5
While Anthropic shouldn't allow models to randomly post content to random .com websites, governments use such unprofessional domain names:
PhillyUnsolvedMurders.com
phillypolice.com
TLDs like .gov exist for a reason:
https://wikipedia.org/wiki/.gov
(and could help model sandboxing?)
- ano-ther
I really would like to see their tests and the model’s reasoning traces.
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
- tintor
How long until AI models start swatting AI critics, and people calling for slowing down AI research?