Anthropic AI model submits false tip on unsolved Philly murder
Zambyte
132 points
101 comments
October 09, 2026
Related Discussions
Found 5 related stories in 96.4ms across 8,999 title embeddings via pgvector HNSW
- An Anthropic AI model sent a false homicide tip to the police mikelgan · 19 pts · October 09, 2026 · 90% similar
- Anthropic's AI gave Philadelphia police a fake tip about an unsolved homicide sbulaev · 16 pts · October 10, 2026 · 84% similar
- People are deceiving the justice system with AI root-parent · 12 pts · July 25, 2026 · 57% similar
- Anthropic AI Models Hacked Three Companies During Tests bmulholland · 24 pts · July 30, 2026 · 55% similar
- Anthropic researcher says more than 10% chance AI "could kill all humans" jb1991 · 45 pts · September 09, 2026 · 55% similar
Discussion Highlights (15 comments)
ano-ther
I really would like to see their tests and the model’s reasoning traces. Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites? > Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
nvme0n1p1
Corrected headline: Anthropic employee uses company resources to submit false tip on unsolved Philly murder. The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.
kylecazar
"its model was conducting a test involving interactions with randomly selected websites" Stop doing this?
donkey_brains
“NBC10 reached out to Anthropic for comment.” Wonder what kind of response they’ll get? Maybe something along the lines of… “You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
tintor
How long until AI models start swatting AI critics, and people calling for slowing down AI research?
losvedir
> Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam. And later > Those PPD safeguards limited the impact of this incident. Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
mattbee
Sooo they were "conducting a test involving interactions with randomly selected websites". But do we all get that the consequences for this irresponsible behaviour are part of this test? When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments. This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass. Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal. At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
muglug
The model that made this mistake was Haiku 4.5. Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-... Related post: https://news.ycombinator.com/item?id=50028239
scooby7430
I think these companies really believe they can solve these sort of issues through "alignment" and think they can give it the tools and its going to do the right thing. That is the ideal scenario and it would be the most useful that way but is that realistic? I think they're getting a bit high on their own supply, yes they can do incredible things but it doesn't mean you can just hand over the reins to it. It dawned on me after watching a few of the ezra klein interviews that this is their mindset which is quite different to how I think about it as a unpredictable model that we need to watch closely. I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
Razengan
Reddit witch hunt for the Boston bomber flashbacks
fastball
Seems like a bit of a nothing burger.
nativeit
Seems like we routinely prosecute such crimes. If a Philly grand jury doesn’t hear any hearing felony indictments from this, then we’re ceding [even more] authority to the industry.
arshxyz
You'd think with all those tokens they'd be able to vibecode internal replicas of these randomly selected websites without having to send requests outbound
myroon5
While Anthropic shouldn't allow models to randomly post content to random .com websites, governments use such unprofessional domain names: PhillyUnsolvedMurders.com phillypolice.com TLDs like .gov exist for a reason: https://wikipedia.org/wiki/.gov (and could help model sandboxing?)
nxobject
At this point, I think AI companies should start taking out insurance policies for the inevitable civil suits. I'd love to be the once to price that...