A frontier model with only grep matches a tuned RAG pipeline at 0.61, but takes five times as long

Benchmarking retrieval for agents on messy real-world company knowledge

A frontier model with only grep matches a tuned RAG pipeline at 0.61, but takes five times as long

Kapa built Company Knowledge Bench, 1,000 eval cases from real production data, to measure retrieval on messy company knowledge. They scored seven retrievers: hybrid search scores 0.41, adding a reranker lifts it to 0.50, query decomposition reaches 0.56, and an optimized agentic retriever hits 0.65 in about five seconds. A frontier model with only grep matches a tuned pipeline at 0.61 but takes five times as long. Public benchmarks were too narrow, so they built their own.

A frontier model with nothing but grep matches a tuned modern retrieval pipeline at 0.61, but takes five times as long.

More from this day

2026-10-02