Shadowwfall/Ketrel-Labs-Agent

A multi-agent RAG system that answers questions about Kestrel Labs' internal documentation — grounded, cited, and honest about uncertainty.

★ 0Forks 0PythonGitHub ↗Compare

README

Kestrel Labs Research Assistant

A multi-agent RAG system that answers questions about Kestrel Labs' internal documentation — grounded, cited, and honest about uncertainty.

Architecture

Three agents orchestrated by LangGraph:

User Query → [Planner] → [Retriever-Synthesizer] → [Verifier] → Cited Answer
  • Planner: Rewrites follow-up questions into standalone queries using conversation history
  • Retriever-Synthesizer: Searches FAISS vector index and drafts a cited answer
  • Verifier: Cross-checks every claim against retrieved chunks, detects document conflicts, assigns a verdict

See DESIGN.md for the full architecture and design rationale.

Setup

1. Prerequisites

2. Clone and install

git clone https://github.com/Shadowwfall/Ketrel-Labs-Agent.git
cd Ketrel-Labs-Agent
pip install -r requirements.txt

3. Configure environment

cp .env.example .env
# Edit .env and fill in your API keys

Required variables:

GROQ_API_KEY=your_groq_key
LANGCHAIN_API_KEY=your_langsmith_key
LANGCHAIN_TRACING_V2=true
LANGCHAIN_PROJECT=kestrel-research-assistant
# Optional: if your LangSmith account is in APAC or EU
# LANGCHAIN_ENDPOINT=https://apac.api.smith.langchain.com

4. Build the vector index

python indexer.py

This embeds the 154 corpus chunks using all-MiniLM-L6-v2 (runs locally, ~30s on CPU) and writes to index/.

5. Run the app

streamlit run app.py

Open http://localhost:8501 in your browser.

Single-command run

After setting environment variables:

python indexer.py && streamlit run app.py

Evaluation

python evaluate.py

Runs 18 test questions (covering single-hop, multi-hop, conflicting, unsupported, and follow-up types) and writes:

  • results/eval_results.jsonl — Per-question results with scores
  • results/metrics_summary.json — Aggregate metrics by type

Run Evaluation to LangSmith Dataset

To evaluate against the LangSmith Dataset and attach live feedback metrics to your dashboard:

python run_langsmith_eval.py

This synchronizes all 18 benchmark questions to the kestrel-eval-suite dataset in LangSmith, runs the multi-agent pipeline, and attaches custom score columns (retrieval_recall, citation_precision, verdict_accuracy, end_to_end_correct) to each experiment run.

Models

Component Model Why
LLM generation llama-3.1-8b-instant via Groq Fast free tier, low latency
Embeddings all-MiniLM-L6-v2 (sentence-transformers) Local, fast, ~80MB, good quality

Project Structure

.
├── corpus.jsonl            # Kestrel Labs wiki — 154 chunks, unchanged
├── indexer.py              # Build FAISS index from corpus
├── tools.py                # search_corpus() tool
├── state.py                # Shared AgentState TypedDict
├── llm_provider.py         # LLM init + rate-limit retry
├── graph.py                # LangGraph pipeline
├── agents/
│   ├── planner.py          # Query rewriting + follow-up resolution
│   ├── retriever_synthesizer.py  # Retrieval + draft answer
│   └── verifier.py         # Claim verification + conflict detection
├── app.py                  # Streamlit chat UI
├── evaluate.py             # Evaluation script
├── results/
│   ├── eval_questions.jsonl
│   ├── eval_results.jsonl
│   ├── metrics_summary.json
│   └── improvement.md
├── DESIGN.md               # Architecture + design rationale
├── REFLECTION.md           # Trade-offs + lessons learned
├── requirements.txt
├── .env.example
└── .gitignore

LangSmith

The LangSmith project kestrel-research-assistant contains:

  • Full traces for every agent invocation
  • Evaluation dataset run with all 18 questions (kestrel-eval-suite-2)
  • Custom feedback scores attached (retrieval_recall, citation_precision, verdict_accuracy, end_to_end_correct)
  • Three representative traced runs (single-hop, multi-hop, unsupported)

Contributors

Shadowwfall

Issues