DocuBrowser: Turning Your Document Pile into a Local AI Knowledge Base

Turning a pile of documents into a searchable useable knowledge base

DocuBrowser: Turning Your Document Pile into a Local AI Knowledge Base

I built DocuBrowser to transform messy piles of PDFs, ebooks, and notes into a searchable knowledge base using local AI. It combines keyword and semantic search to find documents by meaning, not just exact words, while generating instant summaries. Running entirely offline on your machine with Ollama, it ensures your data stays private with no API costs or cloud dependencies.

DocuBrowse turns a messy pile of documents into something you can actually search.
  1. linuxrebe1

    I had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted:

    1) To keep my data local

    2) be able to filter out PII and other data

    3) Be able to find and delete duplicates

    4) Get short synopsis of what a document is

    5) Semantic and keyword search

    6) All of this kept local to me requiring no internet access and no tokens spent to train someone elses AI.

    The result I call DocuBrowser and in it's current form is FOSS (GPL-3) licensed for your personal use. The UI is in your browser. The AI models used are held local and are tiny, Available for Linux(RPM,Deb, and tgz) Windows and Mac. Let me know what you think and thanks for taking the time to try it out.

  2. karmakaze

    I learned a solution is to turn the documents into vectors in say PostgreSQL (with pgvector) and do a cosine similarity search with a search vector. Doing a search for embed models on HuggingFace shows nomic-ai/nomic-embed-text-v1.5 and Qwen/Qwen3-Embedding-0.6B. I might have used a larger one like Qwen/Qwen3-Embedding-4B.

    There's some info for AnythingLLM[0] which supports RAG. AnythingLLM has LanceDB out of the box but also supports others including pgvector.

    [0] https://docs.anythingllm.com/features/embedding-models

  3. asciimoo

    We need projects like this. Automatically classifying the files is smart.

    I'm working on a similar application called Hister (https://github.com/asciimoo/hister). I should borrow some of your ideas. =]

  4. kamranjon

    Wanted to share Antfly which I think serves a similar niche:

    https://antfly.io/

    https://github.com/antflydb/antfly

    They’ve put a lot of effort into optimizing the local llm pipelines and I have a lot of faith in the devs working on it.

  5. mune2gu-chan

    Not a fan of pushing every personal document to someone else's cloud. Nice to see a tool that keeps everything on disk instead.

More from this day

2026-07-08