ORCA-bench Reveals Language Model Agents Are Not Ready for Oncall

Orca-Bench: How Ready Are Language Model Agents for Oncall?

ORCA-bench Reveals Language Model Agents Are Not Ready for Oncall

I introduce ORCA-bench, a new benchmark testing if language model agents can handle real oncall root cause analysis. Using a live microservice system with six days of telemetry, we found the best agents only achieved 25.3% accuracy on realistic tasks. Even top models like Claude Fable 5 struggle significantly, often hallucinating causes or failing without source code access. This gap highlights the massive engineering investment needed before trusting AI with production reliability.

The gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability.

More from this day

2026-07-31