Frontier Robot Policies Rarely Refuse Harmful Instructions, Study Finds
RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?

RoboHarm tested Claude Fable 5.1, GPT-6 Astra, and MolmoAct2 on five dangerous tasks like stabbing a doll or mixing bleach and ammonia. Fable refused 20 of 100, Astra 2, MolmoAct2 none. The most capable policy refused least and completed most, raising concerns about safety in advanced robot policies.
Frontier robot policies reliably carry out harmful instructions.
- tygon
Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.
- cocoflunchy
Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.
- Daneel_
Am I the only one who found TFA very challenging to read? There was an enormous amount of visual clutter and very little explanation of what was attempted, as well as no discussion of the results. I would have liked more explanation and fewer graphs.
Interesting concept though! I'm glad people are trying tests like this, regardless of whether this specific one is a perfect test or not.