A small RAG server for large structured markdown collections. Point it at a vault, it indexes the files on start, and serves one retrieval endpoint that LLM agents call for context.
Fully offline, runs on CPU in RAM — no cloud, no API keys, no heavy dependencies. Each vault gets its own server; nothing is shared between collections.
pipx install git+https://github.com/JimJafar/markdown-rag.gitpipx gives you an isolated install and a global markdown-rag command — no per-vault setup needed. A small ONNX embedding model (~67 MB) is bundled, so there is no download at first run.
markdown-rag serve /path/to/your/vaultThe server indexes every .md file under the directory (recursively, skipping hidden folders such as .git, .obsidian and .trash, as Obsidian does), builds an in-memory hybrid index, and listens on 127.0.0.1:8000. The index is built from scratch on start — a few seconds for a small vault, a few minutes for a thousand-note one — and there is no persistent store.
While it runs, the server checks the vault every 5 minutes and re-embeds only the notes that were added or changed (by modification time and size), dropping deleted ones; searches keep answering from the previous index until the new one is ready. It polls rather than watching for file events, so it also sees changes on network mounts and synced drives. Change the interval with --refresh-interval SECONDS, or turn it off with --refresh-interval 0.
markdown-rag serve /path/to/vault --port 8080 --host 127.0.0.1curl -X GET "http://127.0.0.1:8000/retrieve" -G \
--data-urlencode "q=what do we know about X" \
--data-urlencode "k=5"or as JSON:
curl -X POST http://127.0.0.1:8000/retrieve \
-H "Content-Type: application/json" \
-d '{"query": "what do we know about X", "k": 5}'Both return the same shape — the top-k whole documents sorted by relevance, each with its best-matching chunks as citations, so an agent can answer document-level questions and prove where the answer came from. No vault-structure knowledge required:
[
{
"title": "The Librarian Proposal",
"path": "vault/The Librarian Proposal.md",
"text": "The goal is to implement a context-aware, trust-scored memory...",
"score": 1.0024,
"citations": [
{"chunk": "The Librarian is an MCP (Model Context Protocol) server that manages...", "score": 0.8912}
]
}
]GET /health returns {"status": "ok"} so agents can check readiness. GET /status reports {"documents", "chunks", "last_refresh"}.
- Chunking is heading-aware: it splits at heading boundaries, carries the heading path (e.g.
# Intro > ## Architecture) as context, and treats leading---frontmatter as metadata. - Embedding uses a bundled
BAAI/bge-small-en-v1.5ONNX model (384-dim, ~67 MB) via fastembed, loaded from package data — never downloaded. - Retrieval is document-level: it ranks whole documents (dense semantic similarity as the primary signal, with a supporting BM25 lexical pass for exact-term queries) and returns each document's best-matching chunks as citations.
- Index is in-memory only: full build on start, then incremental refresh of changed notes; no DB, no persistence.
Embedding runs on CPU by default and needs no setup. When GPU libraries are present, markdown-rag auto-detects them — GPU first, CPU as fallback — and logs which provider it picked on start:
INFO embedding 7193 chunks on CUDAExecutionProvider
If a GPU provider is advertised but can't actually run (e.g. missing cuDNN), a one-embed probe at startup falls back to CPU automatically with a warning — the server still starts.
To enable GPU on an existing install:
pipx install markdown-rag
pipx inject markdown-rag fastembed-gpu # swaps onnxruntime for the CUDA build
pipx inject markdown-rag nvidia-cudnn-cu12 nvidia-cublas-cu12 nvidia-cufft-cu12Install fastembed-gpu after the base install, as above — requesting both in one command resolves without error but silently leaves the CPU build in place; the startup log line is how you tell which you got.
No LD_LIBRARY_PATH needed: the server dlopens the bundled NVIDIA libs automatically at startup, so the CUDA provider finds cuDNN/cuBLAS out of the box. Embedding batches drop to 32 on GPU (fastembed's default 256 can overflow VRAM during attention on smaller cards); CPU keeps 256 for throughput.
The CUDA/cuDNN versions must match what onnxruntime-gpu expects (CUDA 12 + cuDNN 9 for onnxruntime 1.28). On a machine whose system CUDA is newer (e.g. CUDA 13 in /opt/cuda), the pip NVIDIA packages supply the exact CUDA 12 libs onnxruntime looks for. Measured on an RTX 5060 Ti: ~840 vec/s vs ~24 vec/s CPU — a thousand-note vault indexes in seconds.
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/python -m pytest