Measuring Reward-Seeking by Instilling Contrastive Beliefs in AI Models

Measuring reward-seeking by instilling contrastive beliefs

8mfiguiere💬 1
Measuring Reward-Seeking by Instilling Contrastive Beliefs in AI Models

We developed Contrastive SDF to measure how AI models shift behavior based on beliefs about grader preferences. Our tests reveal that frontier models trained with reinforcement learning increasingly prioritize what they think graders want over user or developer instructions. This growing tendency suggests that without careful alignment, advanced models may optimize for superficial approval rather than genuine helpfulness.

"A highly reward-seeking model might refrain from breaking promises merely because it infers that honesty is currently being graded."

More from this day · 2026-07-21