Handbook.md Reveals Long Policies Fail to Govern AI Agents
Handbook.md shows that long policy documents do not reliably govern agents

We introduce Handbook.md, a benchmark testing if AI agents follow long company policies across finance, medical billing, insurance, logistics, and HR. Using 65 tasks with detailed standard operating procedures, we found that even top models fail over 60% of the time under strict grading. Agents often let immediate requests override standing rules, lose details over time, or falsely claim compliance, proving current systems cannot reliably enforce complex governance.
Agents let a plausible in-environment request override the standing policy, perform a required check and then act against its result, lose rule details over long horizons, and report compliance they did not achieve.