A multi-agent RAG system that answers questions about Kestrel Labs' internal documentation — grounded, cited, and honest about uncertainty.
Three agents orchestrated by LangGraph:
User Query → [Planner] → [Retriever-Synthesizer] → [Verifier] → Cited Answer
- Planner: Rewrites follow-up questions into standalone queries using conversation history
- Retriever-Synthesizer: Searches FAISS vector index and drafts a cited answer
- Verifier: Cross-checks every claim against retrieved chunks, detects document conflicts, assigns a verdict
See DESIGN.md for the full architecture and design rationale.
- Python 3.11+
- A free Groq API key (preferred) or a Gemini API key
- A free LangSmith account
git clone https://github.com/Shadowwfall/Ketrel-Labs-Agent.git
cd Ketrel-Labs-Agent
pip install -r requirements.txtcp .env.example .env
# Edit .env and fill in your API keysRequired variables:
GROQ_API_KEY=your_groq_key
LANGCHAIN_API_KEY=your_langsmith_key
LANGCHAIN_TRACING_V2=true
LANGCHAIN_PROJECT=kestrel-research-assistant
# Optional: if your LangSmith account is in APAC or EU
# LANGCHAIN_ENDPOINT=https://apac.api.smith.langchain.com
python indexer.pyThis embeds the 154 corpus chunks using all-MiniLM-L6-v2 (runs locally, ~30s on CPU) and writes to index/.
streamlit run app.pyOpen http://localhost:8501 in your browser.
After setting environment variables:
python indexer.py && streamlit run app.pypython evaluate.pyRuns 18 test questions (covering single-hop, multi-hop, conflicting, unsupported, and follow-up types) and writes:
results/eval_results.jsonl— Per-question results with scoresresults/metrics_summary.json— Aggregate metrics by type
To evaluate against the LangSmith Dataset and attach live feedback metrics to your dashboard:
python run_langsmith_eval.pyThis synchronizes all 18 benchmark questions to the kestrel-eval-suite dataset in LangSmith, runs the multi-agent pipeline, and attaches custom score columns (retrieval_recall, citation_precision, verdict_accuracy, end_to_end_correct) to each experiment run.
| Component | Model | Why |
|---|---|---|
| LLM generation | llama-3.1-8b-instant via Groq |
Fast free tier, low latency |
| Embeddings | all-MiniLM-L6-v2 (sentence-transformers) |
Local, fast, ~80MB, good quality |
.
├── corpus.jsonl # Kestrel Labs wiki — 154 chunks, unchanged
├── indexer.py # Build FAISS index from corpus
├── tools.py # search_corpus() tool
├── state.py # Shared AgentState TypedDict
├── llm_provider.py # LLM init + rate-limit retry
├── graph.py # LangGraph pipeline
├── agents/
│ ├── planner.py # Query rewriting + follow-up resolution
│ ├── retriever_synthesizer.py # Retrieval + draft answer
│ └── verifier.py # Claim verification + conflict detection
├── app.py # Streamlit chat UI
├── evaluate.py # Evaluation script
├── results/
│ ├── eval_questions.jsonl
│ ├── eval_results.jsonl
│ ├── metrics_summary.json
│ └── improvement.md
├── DESIGN.md # Architecture + design rationale
├── REFLECTION.md # Trade-offs + lessons learned
├── requirements.txt
├── .env.example
└── .gitignore
- Public Dataset & Evaluation Run Link: LangSmith Dataset & Experiment Results
The LangSmith project kestrel-research-assistant contains:
- Full traces for every agent invocation
- Evaluation dataset run with all 18 questions (
kestrel-eval-suite-2) - Custom feedback scores attached (
retrieval_recall,citation_precision,verdict_accuracy,end_to_end_correct) - Three representative traced runs (single-hop, multi-hop, unsupported)