JetBrains on Building a Semantic Code Search RAG Pipeline: Parsing, Chunking, and Vectorization
Building a RAG Pipeline for Semantic Code Search

JetBrains engineers share field notes from building Air Context, a RAG pipeline for semantic code search. They explain why grep fails for abstract queries like where session tokens get refreshed, and detail structure-aware chunking using JetBrains parsers for nine languages. They also cover vectorization, storage trade-offs between dimension count and precision, and using an LLM-as-a-judge to evaluate chunk quality.
To reason through abstract domains, the agent needs the ability to search for code by meaning, also known as semantic search.
- keeda
I think something like this would be key to improving the quality of coding agents. A very common issue (maybe the biggest one) observed by many people is that agents often produce a lot of duplicate and redundant code; multiple abstractions, methods, classes, data structures, etc. serving minor variations of the same purpose... sometimes within the same file!
My theory is that this is due to a kind of "tunnel vision" these models have as they execute on a given task, because engineers new to a company do the same thing until they learn the "lay of the land" and figure out that similar problems have been solved elsewhere.
In a past job my team owned the internal multi-repo codesearch tool, which was by far the most popular internal tool, and later another team added a similar semantic search capability. This was very exciting, but I left before I could see how well it worked out in real-life.
Like, you'd do a keyword search and explore if you need some major piece of functionality that would require significant work, or whenever you encounter an abstraction whose code does not exist in your repo and you want to learn more about it. But when you're in the flow and inventing smaller abstractions, like a class or utility method, you don't necessarily think to search for it. Worse, even if you did, you could not search for it effectively because something similar may exist with slightly different naming or terminology or a typo that a keyword search would miss. Predictably, at sc […]
- duhhhhh1212
https://www.pangram.com/history/c901e80e-9cb7-46e4-bf10-7348...
I don't want to say don't waste your time since the first half is human written. Questions for the authors: did y'all just get tired of writing and said "fuck it let's have the LLM finish the rest"? Or did one of you use LLM to write the last half and the other used their own words?
- simianwords
Here we go again, the industry largely gave on up RAG. In fact I have hardly seen any case where grep doesn't work as well as RAG.