Anthropic discloses 2 months old fake tip to police among new rogue AI incidents
guessmyname
41 points
35 comments
October 10, 2026
Related Discussions
Found 5 related stories in 98.9ms across 9,063 title embeddings via pgvector HNSW
- An Anthropic AI model sent a false homicide tip to the police mikelgan · 19 pts · October 09, 2026 · 83% similar
- Anthropic AI model submits false tip on unsolved Philly murder Zambyte · 132 pts · October 09, 2026 · 81% similar
- Anthropic's AI gave Philadelphia police a fake tip about an unsolved homicide sbulaev · 16 pts · October 10, 2026 · 80% similar
- Rogue Anthropic AI agent gave police fake tip in unsolved murder case vinni2 · 15 pts · October 10, 2026 · 78% similar
- Anthropic discloses fourth AI hacking incident missed in earlier review accountinhn · 11 pts · September 10, 2026 · 69% similar
Discussion Highlights (9 comments)
thesurlydev
Let’s just get comfortable that these kind of incidents are going to be commonplace. Kudos to the people sounding the alarms but I can’t help but feel skeptical about the ability to keep the rogue entities contained.
grey-area
This sounds highly irresponsible. They set loose LLM agents with instructions to post data to randomly selected websites. What it it submits false data as here? What is it DDOSs a website by mistake? What if it wastes a lot of time and resources? What if it decides it needs to hack a website using a vulnerability it found? These things are very unpredictable and should not be allowed in the open internet except in read mode (and even that doesn’t always work as we have seen). Why are these tests being made using other people’s resources and polluting the common wealth of our public spaces with slop?
hmartin
https://news.ycombinator.com/item?id=50027118
Capricorn2481
I'm sorry, but how can you call it a test suite if it has access to the outside world and is able to submit requests? Sounds like they didn't even bother sandboxing to this time. This is the lamest example of "Rogue AI" I have seen so far.
tcdent
Before you all start screaming negligence and irresponsibility and how-could-they-be-so-dumb, let's think about what we're observing. These are relatively brand new systems that process an incredible amount of data to form their "world view". The word non-deterministic gets thrown around a lot, but you have to understand that there's absolutely nothing deterministic about these architectures. The only way to know that certain qualities could emerge is to observe the qualities emerging. Giving an agent instructions that, for example, instruct it to only perform GET requests, and then observing that the agent does not always respect that request, is not a failure of security or configuration: it's data being gathered. Yeah, we're going to become more diligent about the protections that we put in place beyond any level of protection that we've ever employed before. You can say that we have decades of security research and experience, but we're watching all of that fall more and more day by day, finally putting a real delta on how secure we thought we were versus how secure we actually are [1]. Additionally, adversaries have never existed inside of the systems that we hoped to secure in the first place. So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet. [1] https://x.com/PaulosYibelo/status/2106378929158135903
autoexec
This is pretty much useless without knowing exactly what it was they told their bot to do in the first place. All we get are "example tasks" for what it should have done and a short list of things it was told not to do (which it followed). > Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions. Without more information this looks much less like an AI problem and more like yet another example of incompetent or malicious internal tests. Many people are using Claude. There are no other reports of police stations getting fake reports from AI.
brun
Hmm additional motive behind that September Dario claim over AI internet takeover?
puppycodes
Calling it rouge implies it has a will, which of course it does not. Loose coupling to an outcome is not evidence of sentience. Its just exhausting marketing
oulipo
It's not a "rogue AI", it' s a company letting a statistical model do actions in the real world, sending (by themselves, through that statistical model) fake informations to a federal agency...