Project Arena - Vendor-Neutral Kubernetes Incident Benchmark for AI SRE Agents
Show HN: AI SRE Arena, an Open Benchmark for AI SRE Agents on Kubernetes
Project Arena is an open benchmark from Edge Delta for testing AI investigation tools on Kubernetes. It deploys a disposable cluster, injects faults, and scores how well any product detects, diagnoses, and mitigates incidents. With 21 scenarios and a configurable AI judge, it compares Edge Delta, Grafana, and Claude across root cause, blast radius, and mitigation readiness. Python 3.10+ is the only dependency, and it works with local kind or AWS EKS. Run the smoke suite in minutes or the full suite for rigorous evaluation.
Detection was not independently measured for Claude, so those cells are shown as —.
- smithclay
More benchmarks comparing effectiveness of various AI SREs is welcome and overdue: especially vendor-vs-vendor comparisons. One area that I think is going to be really interesting and important is the best way to emulate complex IT environments for evals.
Some related work I recommend checking out:
- https://arxiv.org/abs/2609.33023 (new last week!)
- https://github.com/SREGym/SREGym
- https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-... (Clickstack's version)
- https://github.com/grafana/o11y-bench (Grafana's version)
- nikhilunni
Cool that you guys created this, but interpreting the results: why wouldn't I just use Claude instead of an AI SRE tool?
Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?
- tkkiran
Cool, as I see your product react proactively rather than Claude being reactive so that with your tools automated investigations happens without me asking explicitly the problem and root cause as I understand? Do you guys also have mitigations?