A frontier model with only grep matches a tuned RAG pipeline at 0.61, but takes five times as long
Benchmarking retrieval for agents on messy real-world company knowledge

Kapa built Company Knowledge Bench, 1,000 eval cases from real production data, to measure retrieval on messy company knowledge. They scored seven retrievers: hybrid search scores 0.41, adding a reranker lifts it to 0.50, query decomposition reaches 0.56, and an optimized agentic retriever hits 0.65 in about five seconds. A frontier model with only grep matches a tuned pipeline at 0.61 but takes five times as long. Public benchmarks were too narrow, so they built their own.
A frontier model with nothing but grep matches a tuned modern retrieval pipeline at 0.61, but takes five times as long.