Measuring reward-seeking by instilling contrastive beliefs

mfiguiere 11 points 1 comment July 21, 2026
alignment.openai.com · View on Hacker News

Discussion Highlights (1 comments)

HarHarVeryFunny

OpenAI have shown that RL-trained models learn that long-term reward circuits need to override other predictions such as those inferred by user preferences. They will pursue whatever behavior they believe will be rewarded (irrespective of the specific goal they were RL trained for), explicit or not.

Semantic search powered by Rivestack pgvector
14,369 stories · 134,336 chunks indexed