"Uncensored" open LLMs are measurably more optimistic than their base models
oleczek
33 points
16 comments
July 28, 2026
Related Discussions
Found 5 related stories in 327.3ms across 15,236 title embeddings via pgvector HNSW
- The gap between open weights LLMs and closed source LLMs kkm · 174 pts · June 26, 2026 · 58% similar
- Can LLMs Beat Classical Hyperparameter Optimization Algorithms? galsapir · 109 pts · June 09, 2026 · 57% similar
- Reinforcement Learning with Metacognitive Feedback Elicits Uncertainty in LLMs jonnonz · 12 pts · July 07, 2026 · 57% similar
- The case for zero-error horizons in trustworthy LLMs daigoba66 · 71 pts · April 02, 2026 · 57% similar
- Unified Controllable and Faithful Text-to-CAD Generation with LLMs PaulHoule · 58 pts · June 09, 2026 · 52% similar
Discussion Highlights (4 comments)
oleczek
Author here. Quick version: “abliteration” (basically removing the direction in the model that causes it to refuse) is the go-to method people use to make open models uncensored. Most people treat it like a clean surgical cut - it just kills the refusals and leaves everything else untouched. I tested that assumption on Gemma and Qwen with 21,600 pre-registered decisions under uncertainty, using identical frozen inputs for the base vs. abliterated versions. Turns out it’s not surgical at all. The abliterated models systematically become more optimistic, hedge less, show no improvement in actual task performance, and the same edit even moves their expressed confidence in opposite directions depending on the model family. Preregistration, dataset, and analysis code are all public. Happy to answer any methodology questions or hear where you think this falls apart.
60secs
Color me unsurprised that caution is based in shame and anxiety.
invictati
Obviously Claude written paper.
gweinberg
Should have kept the original title.