syedazeez337/raglearn

Local hybrid RAG over the vLLM codebase — BGE-M3 + Qdrant (dense+sparse, RRF) + cross-encoder rerank + Qwen3.5-4B, with a source-grounded evaluation harness.

★ 0Forks 0PythonGitHub ↗Compare

README

raglearn

A hybrid retrieval-augmented generation (RAG) pipeline that answers questions about the vLLM codebase, with an honest, source-grounded evaluation harness. Built to run fully locally on a single 6 GB GPU.

  • Hybrid retrieval — BGE-M3 dense + learned-sparse vectors, fused with RRF in Qdrant
  • Cross-encoder reranking — bge-reranker-v2-m3, with scope-aware dedup
  • Grounded generation — Qwen3.5-4B on llama.cpp; answers cite file paths or refuse
  • Trustworthy eval — 30 questions, references read from source (not the model), LLM-judged, retrieval and generation scored separately

Pipeline architecture

The code is organised so each module maps to exactly one box in the diagram above.

rag/
  config.py                  shared settings: models, collection, endpoints, system prompt
  indexing/                  ── Offline Indexing Pipeline ──
    load.py                  Raw Documents     — load & clean vLLM source
    chunk.py                 Chunking          — AST split (class bodies grouped)
    embed.py                 Embedding Model   — BGE-M3 dense + learned sparse
    index.py                 Vector Database   — Qdrant create + upsert
    pipeline.py              orchestrates load → chunk → embed → index
  retrieval/                 ── Online Retrieval-Generation Pipeline ──
    query_encoder.py         Query Encoder     — embed the query (same BGE-M3)
    search.py                Similarity Search — Qdrant hybrid dense+sparse, RRF fusion
    rerank.py                Re-ranker         — bge-reranker-v2-m3 + scope-aware dedup
    generate.py              LLM Generator     — llama.cpp (Qwen3.5-4B)
    answer.py                Answer + Citations — orchestrates the online pipeline
  server.py                  FastAPI service exposing the online pipeline
  eval/                      evaluation harness (dataset, runner, reference gate, model probe)

scripts/serve_llm.sh         LLM Generator backend (llama.cpp launcher)
docs/                        architecture + eval visualisations
data/                        eval result snapshots (index data is regenerated, not committed)

Prerequisites

This repo is the pipeline code. The large external pieces are not committed — set them up once:

  1. vLLM source (the corpus) into ./vllm:
    git clone --depth 1 https://github.com/vllm-project/vllm.git vllm
  2. llama.cpp (the LLM backend) built into ./llama.cpp:
    git clone https://github.com/ggml-org/llama.cpp.git
    cmake -S llama.cpp -B llama.cpp/build -DGGML_CUDA=ON && cmake --build llama.cpp/build -j
  3. The model — a Qwen3.5-4B GGUF into ./models/Qwen3.5-4B-Q4_K_M.gguf (or set LLAMA_MODEL to your path).
  4. Qdrant via Docker (data persists in a named volume):
    docker run -d --name qdrant -p 6333:6333 -v qdrant_data:/qdrant/storage qdrant/qdrant
  5. Python deps (Python 3.12; CUDA torch — see note in requirements.txt):
    uv venv --python 3.12 && source .venv/bin/activate
    uv pip install -r requirements.txt
  6. Eval key (optional, only for the judge): cp .env.example .env and add your GOOGLE_API_KEY.

Running

# 1. LLM Generator backend (:8080)
scripts/serve_llm.sh

# 2. Build the index once  (load → chunk → embed → upsert)
python -m rag.indexing.pipeline --recreate

# 3. Online service (:8000)
python -m uvicorn rag.server:app --host 127.0.0.1 --port 8000

# 4. Query it
curl -s :8000/query -H 'content-type: application/json' \
  -d '{"question":"What is the default KV cache block_size?","top_k":5}'

Every component is also runnable standalone, e.g. python -m rag.retrieval.search "<query>".

Evaluation

The eval set has 30 questions across six scenarios: A–D answerable (architectural, API-surface, factual, cross-component), E out-of-corpus (must refuse), F false-premise (must reject). A–D references are read from the vLLM source on disk, so answer-correctness is independent of the system's own output. Retrieval (did the ground-truth file get retrieved?) and generation (is the answer right?) are scored separately.

python -m rag.eval.check_references                                  # gate: A-D references filled
JUDGE_MODEL=gemini-2.5-flash python -m rag.eval.run --no-resume      # full 30-question run

Results

Evaluation results

A chunking bug was found where oversized config-class bodies were shattered into micro-chunks that all shared one signature, so retrieval-side dedup made ~15% of the corpus unreachable. Grouping class-body statements and making the dedup key scope/part-aware lifted answer accuracy:

Metric Before After
Answer accuracy (A–D) 0.62 0.85
Scenario D (cross-component) 0.40 0.80
Out-of-corpus refusal (E) 10/10 10/10
False-premise rejection (F) 10/10 10/10

Notably, context-recall stayed flat (0.75 → 0.72) while accuracy rose +0.22 — the bug was never which files were retrieved, but what was inside the chunks and dedup discarding the good ones.

Known upstream warning (not suppressed)

Loading BGE-M3 emits, on Python 3.10+, a DeprecationWarning from sentencepiece: builtin type SwigPyObject has no __module__ attribute. It originates in SWIG-generated bindings and is fixed only by SWIG 4.4 + a sentencepiece rebuild (swig/swig#2881, google/sentencepiece#1150). No released sentencepiece fixes it yet. It is a benign DeprecationWarning (ignored by Python's default filter, so it does not appear in normal runs) and is left visible rather than filtered out.

License

MIT

Contributors

syedazeez337

Issues