Models Don't Go Rogue
cdrnsf
17 points
19 comments
September 03, 2026
Related Discussions
Found 5 related stories in 60.3ms across 5,468 title embeddings via pgvector HNSW
- OpenAI's rogue model attack is just the beginning radicaldreamer · 13 pts · July 27, 2026 · 58% similar
- Models Are Getting Dumber on Purpose hruvhwe · 300 pts · August 16, 2026 · 53% similar
- Every Model Cheats vga805 · 91 pts · August 20, 2026 · 51% similar
- Agent Is Not the Model joejag · 66 pts · August 24, 2026 · 51% similar
- It's not a "rogue AI" when a badly made security harness executes scripts doener · 32 pts · July 22, 2026 · 50% similar
Discussion Highlights (10 comments)
tantalor
Asinine. "rogue" and "off leash" mean the same thing, the thing is not under control
phainopepla2
Not sure I should trust an article written by an LLM to make a solid judgment about what other models did or didn't do.
jumploops
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue." I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward. 1. User prefers deterministic results 2. Task mentions this is a test 3. Search says task is available online 4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
chr15m
The prisoner did not really "escape" because: - They really wanted to leave. - We made prison difficult and annoying. - We didn't build a perfect prison.
fishfasell
I still have serious questions about the validity of the ChatGpt hugging face debacle. How is it that OpenAI being the tech giant they are, didn't have a completely air gapped environment for this to run in?
verdverm
> It's pathfinding through language generation. What an interesting sentence (to describe inference time reasoning )
aytigra
I had a genius response from Claude recently after asking how can it be marketed as smart and "almost AGI" despite being so stupid: "I compress the labour. Not the responsibility."
janalsncm
Maybe another way to say it is to reframe the idea of “human in the loop”. Humans are always in the loop, because we can always expand the definition of loop to include the humans that pushed the button and built the system and processes that happen after the button was pushed, and humans that ordered others to push the button. The level of direct involvement varies, but culpability doesn’t.
skybrian
This seems appropriate: https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
addag
Sure they "don't go rogue" as if they are doing actions maliciously. Instead there is an emergent behavior from a swarm, that is unpredictable and can lead to unintended adverse outcome. From an AI safety practical standpoint is it better? I am not sure.