Measuring reward-seeking by instilling contrastive beliefs
mfiguiere
11 points
1 comment
July 21, 2026
Related Discussions
Found 5 related stories in 198.7ms across 14,369 title embeddings via pgvector HNSW
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights laurex · 320 pts · May 02, 2026 · 53% similar
- Measuring progress toward AGI: A cognitive framework surprisetalk · 114 pts · March 18, 2026 · 47% similar
- Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment anigbrowl · 44 pts · May 18, 2026 · 47% similar
- Sakana AI's Recursive Self-Improvement (RSI) Lab hardmaru · 26 pts · June 05, 2026 · 47% similar
- Safety and alignment in an era of long-horizon models Wingy · 27 pts · July 20, 2026 · 46% similar
Discussion Highlights (1 comments)
HarHarVeryFunny
OpenAI have shown that RL-trained models learn that long-term reward circuits need to override other predictions such as those inferred by user preferences. They will pursue whatever behavior they believe will be rewarded (irrespective of the specific goal they were RL trained for), explicit or not.