Measuring reward-seeking by instilling contrastive beliefs
mfiguiere
11 points
1 comment
July 21, 2026
Related Discussions
Found 5 related stories in 59.1ms across 5,564 title embeddings via pgvector HNSW
- GulliBench: Measuring Skepticism in Frontier Models rigelbm · 18 pts · August 12, 2026 · 48% similar
- Lean Eval for Alignment on Faithfulness asdajksbda · 103 pts · August 11, 2026 · 47% similar
- Calibrate Before You Accelerate: Bias Toward Action in a New Role tuckerwales · 137 pts · August 29, 2026 · 47% similar
- Safety and alignment in an era of long-horizon models Wingy · 27 pts · July 20, 2026 · 46% similar
- The Conceptual Reasoning Index optimalsolver · 74 pts · August 13, 2026 · 46% similar
Discussion Highlights (1 comments)
HarHarVeryFunny
OpenAI have shown that RL-trained models learn that long-term reward circuits need to override other predictions such as those inferred by user preferences. They will pursue whatever behavior they believe will be rewarded (irrespective of the specific goal they were RL trained for), explicit or not.