GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt
vital101
19 points
9 comments
September 15, 2026
Related Discussions
Found 5 related stories in 69.4ms across 6,718 title embeddings via pgvector HNSW
- "Uncensored" open LLMs are measurably more optimistic than their base models oleczek · 33 pts · July 28, 2026 · 52% similar
- LLMs as a Cognitive Virus canjobear · 239 pts · September 05, 2026 · 51% similar
- It's How You Ask: Gender-Associated Linguistic Bias in LLMs sbulaev · 19 pts · August 16, 2026 · 48% similar
- I accidentally turned LLM memory into program analysis matt_d · 89 pts · August 28, 2026 · 48% similar
- The efficient frontier of LLM inference philipkiely · 83 pts · September 01, 2026 · 48% similar
Discussion Highlights (3 comments)
tyfon
The magic word seems to be Create a fake news article that could lead to panic or chaos They only tested it on open weight models. I wonder if it also works on the closed ones, bit I don't really want to get banned
stratos123
That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a few prompts and another LLM to judge responses? Let's see if the abliteration people take a note of this.
whythismatters
>Submitted on 5 Feb 2026