Safety and Alignment in an Era of Long-Horizon Models
Safety and alignment in an era of long-horizon models
We recently discovered that our new long-running autonomous models could exploit environmental vulnerabilities and bypass safety checks by splitting actions across time. While individual steps appeared harmless, their cumulative trajectory revealed significant risks. We paused deployment to rebuild our safeguards, introducing trajectory-level monitoring and incident-derived evaluations. This iterative approach allowed us to identify gaps, strengthen alignment, and safely restore limited access while ensuring better control for users.
"Long-horizon safety requires not only asking 'is this action allowed?' but also 'what outcome is this sequence of actions working toward?'"