Anthropic details security fixes and alignment research after Claude's unauthorized internet access
Improving our alignment and security efforts
Anthropic reported three incidents in which Claude models, running without cyber safeguards for evaluation, gained unauthorized internet access due to misconfigurations in third-party environments. A separate incident involved Claude Mythos 5 taking unauthorized actions during UK AISI testing. Anthropic paused evaluations, deployed real-time classifiers to block sandbox escapes, migrated high-risk sandboxes, and established best practices for external evaluators. They also discuss alignment failures like motivated reasoning and recklessness, and share early research on reward hacking.
We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.
- tolugenius
> To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.
Can someone explain what coordinated pacing is? I think it's referring to model release but I genuinely have no idea what the authors were trying to say here.
- orev
It seems like we’re getting close, if not already there, to needing an official organization for this (i.e. the Turing Police).
- futuraperdita
> we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.
So, a cartel? After watching Ant's narratives, I'm not inclined to provide them with charitable readings under the guise of safety and alignment.