A hybrid retrieval-augmented generation (RAG) pipeline that answers questions about the vLLM codebase, with an honest, source-grounded evaluation harness. Built to run fully locally on a single 6 GB GPU.
- Hybrid retrieval — BGE-M3 dense + learned-sparse vectors, fused with RRF in Qdrant
- Cross-encoder reranking — bge-reranker-v2-m3, with scope-aware dedup
- Grounded generation — Qwen3.5-4B on llama.cpp; answers cite file paths or refuse
- Trustworthy eval — 30 questions, references read from source (not the model), LLM-judged, retrieval and generation scored separately
The code is organised so each module maps to exactly one box in the diagram above.
rag/
config.py shared settings: models, collection, endpoints, system prompt
indexing/ ── Offline Indexing Pipeline ──
load.py Raw Documents — load & clean vLLM source
chunk.py Chunking — AST split (class bodies grouped)
embed.py Embedding Model — BGE-M3 dense + learned sparse
index.py Vector Database — Qdrant create + upsert
pipeline.py orchestrates load → chunk → embed → index
retrieval/ ── Online Retrieval-Generation Pipeline ──
query_encoder.py Query Encoder — embed the query (same BGE-M3)
search.py Similarity Search — Qdrant hybrid dense+sparse, RRF fusion
rerank.py Re-ranker — bge-reranker-v2-m3 + scope-aware dedup
generate.py LLM Generator — llama.cpp (Qwen3.5-4B)
answer.py Answer + Citations — orchestrates the online pipeline
server.py FastAPI service exposing the online pipeline
eval/ evaluation harness (dataset, runner, reference gate, model probe)
scripts/serve_llm.sh LLM Generator backend (llama.cpp launcher)
docs/ architecture + eval visualisations
data/ eval result snapshots (index data is regenerated, not committed)
This repo is the pipeline code. The large external pieces are not committed — set them up once:
- vLLM source (the corpus) into
./vllm:git clone --depth 1 https://github.com/vllm-project/vllm.git vllm
- llama.cpp (the LLM backend) built into
./llama.cpp:git clone https://github.com/ggml-org/llama.cpp.git cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON && cmake --build llama.cpp/build -j - The model — a Qwen3.5-4B GGUF into
./models/Qwen3.5-4B-Q4_K_M.gguf(or setLLAMA_MODELto your path). - Qdrant via Docker (data persists in a named volume):
docker run -d --name qdrant -p 6333:6333 -v qdrant_data:/qdrant/storage qdrant/qdrant
- Python deps (Python 3.12; CUDA torch — see note in
requirements.txt):uv venv --python 3.12 && source .venv/bin/activate uv pip install -r requirements.txt
- Eval key (optional, only for the judge):
cp .env.example .envand add yourGOOGLE_API_KEY.
# 1. LLM Generator backend (:8080)
scripts/serve_llm.sh
# 2. Build the index once (load → chunk → embed → upsert)
python -m rag.indexing.pipeline --recreate
# 3. Online service (:8000)
python -m uvicorn rag.server:app --host 127.0.0.1 --port 8000
# 4. Query it
curl -s :8000/query -H 'content-type: application/json' \
-d '{"question":"What is the default KV cache block_size?","top_k":5}'Every component is also runnable standalone, e.g. python -m rag.retrieval.search "<query>".
The eval set has 30 questions across six scenarios: A–D answerable (architectural, API-surface, factual, cross-component), E out-of-corpus (must refuse), F false-premise (must reject). A–D references are read from the vLLM source on disk, so answer-correctness is independent of the system's own output. Retrieval (did the ground-truth file get retrieved?) and generation (is the answer right?) are scored separately.
python -m rag.eval.check_references # gate: A-D references filled
JUDGE_MODEL=gemini-2.5-flash python -m rag.eval.run --no-resume # full 30-question runA chunking bug was found where oversized config-class bodies were shattered into micro-chunks that all shared one signature, so retrieval-side dedup made ~15% of the corpus unreachable. Grouping class-body statements and making the dedup key scope/part-aware lifted answer accuracy:
| Metric | Before | After |
|---|---|---|
| Answer accuracy (A–D) | 0.62 | 0.85 |
| Scenario D (cross-component) | 0.40 | 0.80 |
| Out-of-corpus refusal (E) | 10/10 | 10/10 |
| False-premise rejection (F) | 10/10 | 10/10 |
Notably, context-recall stayed flat (0.75 → 0.72) while accuracy rose +0.22 — the bug was never which files were retrieved, but what was inside the chunks and dedup discarding the good ones.
Loading BGE-M3 emits, on Python 3.10+, a DeprecationWarning from sentencepiece:
builtin type SwigPyObject has no __module__ attribute. It originates in SWIG-generated
bindings and is fixed only by SWIG 4.4 + a sentencepiece rebuild
(swig/swig#2881,
google/sentencepiece#1150). No
released sentencepiece fixes it yet. It is a benign DeprecationWarning (ignored by
Python's default filter, so it does not appear in normal runs) and is left visible
rather than filtered out.

