Danielyan86/KnowledgeWeaver

★ 0Forks 0PythonGitHub ↗Compare

README

KnowledgeWeaver

Knowledge Graph + RAG Hybrid Question Answering System

中文文档 | English

Overview

KnowledgeWeaver is an intelligent question-answering system that combines Knowledge Graph (KG) and Retrieval-Augmented Generation (RAG) technologies. The system automatically extracts entities and relationships from documents, builds a knowledge graph, and provides accurate, context-aware answers through a combination of vector retrieval and graph reasoning.

Core Features

  • ✅ Async Concurrent Processing: Process document chunks concurrently using Claude CLI/Gemini API, 3-5x faster
  • ✅ Resume from Checkpoint: Support interruption recovery and resume processing
  • ✅ Progress Tracking: Real-time progress monitoring and status updates
  • ✅ Neo4j Graph Storage: High-performance graph database with built-in GDS algorithms
  • ✅ Vector Search: ChromaDB vector storage for semantic search
  • ✅ Hybrid Retrieval: Combines vector similarity and graph structure retrieval
  • ✅ Modular Architecture: Layered package structure for easy maintenance and extension
  • ✅ Document-Scoped Q&A: Ask questions about specific documents or entire knowledge base
  • ✅ Hierarchical Visualization: Multi-level graph views (overview, document, entity)
  • ✅ Bilingual Support: Native support for Chinese and English content
  • ✅ Observability: Phoenix/Langfuse integration for LLM call tracking and performance monitoring

System Architecture

Application Architecture

System Architecture

Interactive Version (Recommended):

AWS Deployment Architecture

AWS Architecture

Learn More:

Processing Pipeline

  1. Input Layer:

    • Document upload (TXT/PDF via API or CLI)
    • Automatic file queuing and management
  2. Processing Layer:

    • Document Chunking: Smart chunking with 50% overlap for context preservation
    • Async Concurrent Extraction: Process multiple chunks in parallel (3-5x faster)
      • Gemini API for extraction (fast, free tier available)
      • Entity and relationship extraction using LLM
    • Knowledge Normalization:
      • Entity deduplication and merging
      • Relationship standardization
      • Cross-document entity linking
    • Progress Tracking: Real-time status updates and checkpoint saves
  3. Storage Layer:

    • Neo4j Graph Database:
      • Entity nodes with doc_ids array (cross-document sharing)
      • Relationship edges with doc_id (document-specific)
      • Support for 60+ built-in GDS algorithms
    • ChromaDB Vector Database:
      • Document chunks for semantic search
      • Entity embeddings for entity retrieval
      • Bilingual embedding support
  4. Service Layer:

    • FastAPI Backend: High-performance async API
    • Hybrid Retriever:
      • Vector similarity search (RAG)
      • Graph structure traversal (KG)
      • Dynamic weight fusion
    • QA Engine:
      • Multiple query modes (auto, kg_only, rag_only, hybrid)
      • Document-scoped or global queries
      • Context-aware answer generation
    • Session Tracking: Phoenix/Langfuse observability (optional)
  5. Frontend Layer:

    • Interactive Visualization:
      • D3.js force-directed graph
      • Hierarchical views (overview, document, entity)
      • Real-time search and filtering
    • Chat Interface: Natural language Q&A
    • Document Selector: Per-document or global view

Tech Stack

Backend

  • FastAPI: High-performance web framework with async support
  • LLM Providers:
    • Gemini API: Fast and free document extraction (gemini-2.0-flash)
    • Custom Endpoints: Q&A system (DeepSeek, GPT-4, etc.)
    • Claude CLI: Subprocess-based concurrent processing (optional)
  • Neo4j: Graph database with GDS algorithm library
  • ChromaDB: Vector database for semantic embedding search
  • asyncio: Asynchronous concurrent processing
  • Bilingual Support: Native Chinese and English processing

Frontend

  • D3.js: Interactive knowledge graph visualization
    • Hierarchical views (overview, document, entity)
    • Force-directed graph layout
    • Real-time search and filtering
  • JavaScript: Interactive logic and chat interface

Observability (Optional)

  • Phoenix: AI observability and evaluation platform (OpenTelemetry-based)
    • Session-level tracing
    • LLM call monitoring
    • Performance analysis and evaluation
    • Integration Guide
  • Langfuse: Production LLM monitoring

Quick Start

1. Install Dependencies

pip install -r requirements.txt

2. Configure Environment Variables

Copy .env.example to .env and configure:

# === Document Extraction - Gemini API (Recommended: Fast & Free) ===
EXTRACTION_LLM_BACKEND=gemini
GEMINI_API_KEY=your_gemini_api_key_here
GEMINI_MODEL=gemini-2.0-flash

# === Embedding - Gemini (Recommended: Fast & Free) ===
EMBEDDING_BACKEND=gemini
GEMINI_EMBEDDING_MODEL=text-embedding-004

# === Q&A System - Custom LLM Endpoint ===
LLM_BINDING_HOST=https://space.ai-builders.com/backend/v1
LLM_BINDING_API_KEY=your_api_key_here
LLM_MODEL=deepseek  # or gpt-4, claude-3-sonnet, etc.

# === Concurrent Processing ===
CONCURRENT_REQUESTS=5  # Concurrent document chunk processing
MAX_RETRIES=3
CHUNK_SIZE=800
CHUNK_OVERLAP_RATIO=0.5

# === Neo4j Configuration ===
USE_NEO4J=true
NEO4J_URI=bolt://localhost:7687
NEO4J_USER=neo4j
NEO4J_PASSWORD=your_neo4j_password
NEO4J_MAX_POOL_SIZE=50
NEO4J_BATCH_SIZE=500

# === Service Configuration ===
HOST=0.0.0.0
PORT=9621

# === Observability (Optional) ===
# Phoenix
PHOENIX_ENABLED=false
PHOENIX_COLLECTOR_ENDPOINT=http://localhost:4317
PHOENIX_PROJECT_NAME=knowledge-weaver

# Langfuse
LANGFUSE_ENABLED=false
LANGFUSE_PUBLIC_KEY=pk-lf-xxx
LANGFUSE_SECRET_KEY=sk-lf-xxx
LANGFUSE_HOST=https://cloud.langfuse.com

3. Start Service

Option 1: Quick Start (Recommended for Development)

./scripts/start_dev.sh

Option 2: Full Start with Checks (Recommended for First Run)

./scripts/start.sh

Option 3: Manual Start

python -m backend.server

Other Commands:

# Check service status
./scripts/status.sh

# Stop service
./scripts/stop.sh

# Restart service
./scripts/restart.sh

See scripts/README.md for detailed documentation.

4. Access the Service

5. Process Documents

Recommended: Async API (Fast, Concurrent Processing)

curl -X POST "http://localhost:9621/documents/upload-async" \
  -F "file=@your_document.txt"

# Response: {"doc_id": "abc123", "status": "processing"}

# Check progress
curl "http://localhost:9621/documents/progress/abc123"

Alternative: Command Line

python -m backend.extraction.async_extractor path/to/document.txt

See User Guide below for more details.

User Guide

Starting Local Services

1. Start Neo4j Database

Option A: Docker (Recommended)

docker run -d \
  --name neo4j \
  -p 7474:7474 -p 7687:7687 \
  -e NEO4J_AUTH=neo4j/your_password \
  -v $(pwd)/data/neo4j/data:/data \
  neo4j:latest

Access Neo4j Browser at: http://localhost:7474

Option B: Native Installation

Download and install from: https://neo4j.com/download/

2. Start FastAPI Backend

python -m backend.server

The API will start at: http://localhost:9621

3. (Optional) Start Observability Tools

Option A: Langfuse - For Production Monitoring

Langfuse provides detailed LLM call tracking, cost analysis, and production monitoring.

# 1. Configure environment variables in .env
LANGFUSE_ENABLED=true
LANGFUSE_PUBLIC_KEY=pk-lf-xxx
LANGFUSE_SECRET_KEY=sk-lf-xxx
LANGFUSE_HOST=https://cloud.langfuse.com  # or your self-hosted URL

# 2. Install Langfuse
pip install langfuse>=2.0.0

# 3. Restart backend service
python -m backend.server

Visit https://cloud.langfuse.com to view traces.

For self-hosted deployment, see: Langfuse Complete Guide

Option B: Phoenix - For Development and Evaluation

Phoenix offers zero-code tracing, experiment tracking, and prompt optimization.

# 1. Start Phoenix server (Docker)
docker run -d \
  --name phoenix \
  -p 6006:6006 \
  -p 4317:4317 \
  -v $(pwd)/data/phoenix:/data \
  arizephoenix/phoenix:latest

# 2. Configure environment variables in .env
PHOENIX_ENABLED=true
PHOENIX_ENDPOINT=http://localhost:6006/v1/traces

# 3. Install Phoenix packages
pip install arize-phoenix arize-phoenix-otel openinference-instrumentation-openai

# 4. Restart backend service
python -m backend.server

Access Phoenix UI at: http://localhost:6006

For detailed setup and comparison, see: Phoenix Integration Guide

Using the System

Upload and Process Documents

Via API (Asynchronous - Recommended)

curl -X POST "http://localhost:9621/documents/upload-async" \
  -F "file=@your_document.txt"

# Response: {"doc_id": "abc123", "status": "processing"}

# Check progress
curl "http://localhost:9621/documents/progress/abc123"

Via API (Synchronous)

curl -X POST "http://localhost:9621/documents/upload" \
  -F "file=@your_document.pdf"

Via Command Line

python backend/extraction/async_extractor.py path/to/document.txt

Query the Knowledge Graph

Ask Questions

curl -X POST "http://localhost:9621/qa" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What is knowledge graph?",
    "mode": "auto",
    "n_hops": 2,
    "top_k": 5
  }'

Query modes:

  • auto: Automatically choose best strategy
  • kg_only: Knowledge graph only
  • rag_only: Vector retrieval only
  • hybrid: Combine both KG and RAG
  • kg_first: Try KG first, fallback to RAG
  • rag_first: Try RAG first, fallback to KG

Semantic Search

curl -X POST "http://localhost:9621/search" \
  -H "Content-Type: application/json" \
  -d '{
    "query": "investment strategies",
    "search_type": "all",
    "top_k": 10
  }'

Search types:

  • all: Search both chunks and entities
  • chunks: Search document chunks only
  • entities: Search entities only

Visualize Knowledge Graph

Get Full Graph

curl "http://localhost:9621/graphs"

Get Document-Specific Graph

curl "http://localhost:9621/documents/abc123"

View in Browser

Open frontend/index.html in your browser to interact with the D3.js visualization.

Monitor and Manage

View Statistics

# Knowledge graph stats
curl "http://localhost:9621/stats"

# Vector store stats
curl "http://localhost:9621/vector-stats"

List Documents

curl "http://localhost:9621/documents"

Delete Document

curl -X DELETE "http://localhost:9621/documents/abc123"

Configuration Reference

Edit .env file to customize settings:

# === LLM Configuration ===
LLM_BINDING_HOST=https://space.ai-builders.com/backend/v1
LLM_BINDING_API_KEY=your_api_key
LLM_MODEL=deepseek  # or other model

# === Neo4j Configuration ===
USE_NEO4J=true
NEO4J_URI=bolt://localhost:7687
NEO4J_USER=neo4j
NEO4J_PASSWORD=your_password
NEO4J_MAX_POOL_SIZE=50
NEO4J_BATCH_SIZE=500

# === Processing Configuration ===
CONCURRENT_REQUESTS=5  # Concurrent LLM requests
MAX_RETRIES=3
CHUNK_SIZE=800
CHUNK_OVERLAP_RATIO=0.5

# === Observability (Optional) ===
# Langfuse
LANGFUSE_ENABLED=false
LANGFUSE_PUBLIC_KEY=pk-lf-xxx
LANGFUSE_SECRET_KEY=sk-lf-xxx
LANGFUSE_HOST=https://cloud.langfuse.com

# Phoenix
PHOENIX_ENABLED=false
PHOENIX_ENDPOINT=http://localhost:6006/v1/traces

# === Service Configuration ===
HOST=0.0.0.0
PORT=9621

Troubleshooting

Neo4j Connection Failed

# Check if Neo4j is running
docker ps | grep neo4j

# Check logs
docker logs neo4j

# Verify credentials in .env match your Neo4j setup

API Server Not Starting

# Check port availability
lsof -i :9621

# Check logs for detailed error messages
python -m backend.server

Document Processing Stuck

# Check progress
curl "http://localhost:9621/documents/progress/YOUR_DOC_ID"

# Check checkpoint files
ls -la data/checkpoints/YOUR_DOC_ID/

# Restart processing (will resume from checkpoint)
curl -X POST "http://localhost:9621/documents/upload-async" \
  -F "file=@same_document.txt"

LLM API Errors

# Verify API key is correct
# Check LLM_BINDING_HOST is accessible
# Review rate limits and quotas

For more detailed guides, see the Documentation Index.

Project Structure

KnowledgeWeaver/
├── backend/                      # Backend code
│   ├── core/                    # Core modules
│   │   ├── config.py           # Configuration and prompt templates
│   │   ├── language_utils.py   # Bilingual support utilities
│   │   ├── embeddings/         # Embedding services
│   │   │   └── service.py      # Text embedding service (Gemini/OpenAI)
│   │   ├── storage/            # Storage layer
│   │   │   ├── neo4j.py       # Neo4j graph storage
│   │   │   └── vector.py      # ChromaDB vector storage
│   │   ├── observability.py    # Langfuse integration (optional)
│   │   └── phoenix_observability.py  # Phoenix integration (optional)
│   │
│   ├── extraction/              # Knowledge extraction module
│   │   ├── async_extractor.py  # Async extractor (concurrent processing)
│   │   ├── extractor.py        # Sync extractor (backup)
│   │   ├── normalizer.py       # Graph normalization (dedup, merge)
│   │   └── entity_filter.py    # Entity filtering
│   │
│   ├── retrieval/               # Retrieval module
│   │   ├── hybrid_retriever.py # Hybrid retriever (KG + RAG)
│   │   ├── qa_engine.py        # Q&A engine
│   │   ├── overview_generator.py  # Hierarchical graph generator
│   │   └── prompts/            # Prompt management
│   │       ├── extraction_prompts.py  # Extraction prompts
│   │       ├── qa_prompts.py         # Q&A prompts
│   │       └── prompt_loader.py      # Prompt loader
│   │
│   ├── management/              # Management module
│   │   ├── kg_manager.py       # Knowledge graph unified manager
│   │   └── progress_tracker.py # Progress tracking
│   │
│   └── server.py                # FastAPI service entry point
│
├── frontend/                    # Frontend code
│   ├── index.html              # Main page
│   ├── kg-config.js            # Graph configuration
│   ├── kg-core.js              # Core graph logic
│   ├── kg-sidebar.js           # Sidebar UI
│   ├── kg-chat.js              # Chat interface
│   └── kg-*.js                 # Other modules
│
├── data/                        # Data directory (gitignored)
│   ├── storage/
│   │   └── vector_db/          # Vector database (ChromaDB)
│   ├── checkpoints/            # Processing checkpoints (resume capability)
│   ├── progress/               # Progress tracking data
│   └── inputs/                 # User uploaded files
│       └── __enqueued__/       # Processing queue
│
├── docs/                        # Documentation
│   ├── architecture/           # Architecture diagrams and design docs
│   ├── deployment/             # AWS and deployment guides
│   ├── database/               # Neo4j setup and optimization
│   ├── observability/          # Phoenix/Langfuse integration
│   ├── development/            # Development and testing guides
│   └── README.md               # Documentation index
│
├── deploy/                      # Deployment configuration
│   ├── docker/                 # Docker images
│   │   └── api/                # API service Dockerfile
│   ├── kubernetes/             # Kubernetes manifests
│   │   ├── base/               # Base configuration
│   │   └── README.md           # Deploy guide
│   └── terraform/              # AWS EKS infrastructure (IaC)
│
├── scripts/                     # Utility scripts
│   ├── start.sh                # Full start (with checks)
│   ├── start_dev.sh            # Quick start (dev mode)
│   ├── start_with_phoenix.sh   # Start with Phoenix monitoring
│   ├── stop.sh                 # Stop service
│   ├── restart.sh              # Restart service
│   ├── status.sh               # Check service status
│   └── README.md               # Script documentation
│
├── tests/                       # Test cases
│   ├── test_*.py               # Test scripts
│   └── data/                   # Test data
│
├── logs/                        # Log files (gitignored, runtime)
├── .env                         # Environment config (not committed)
├── .env.example                 # Environment config template
├── requirements.txt             # Python dependencies
├── CLAUDE.md                    # Project config for Claude Code
└── README.md                    # Project documentation (this file)

Key Features

Async Concurrent Processing

  • 3-5x Speed Improvement: Process document chunks in parallel
  • Checkpoint & Resume: Automatic checkpoint saves for interruption recovery
  • Progress Tracking: Real-time progress monitoring via API
  • Configurable Concurrency: Adjust concurrent requests based on rate limits

Neo4j Graph Storage

  • Cross-Document Entity Sharing: Entities shared via doc_ids array
  • Document-Specific Relationships: Each relationship tagged with doc_id
  • Smart Incremental Updates: Add/update/delete without affecting shared entities
  • 60+ Built-in Algorithms: PageRank, Louvain, Node Similarity, Shortest Path, etc.
  • Efficient Querying: Cypher query language with full-text search support

Knowledge Graph Normalization

  • Entity Deduplication: Merge similar entities across documents
  • Relationship Standardization: Normalize relationship types and labels
  • Information Island Detection: Connect isolated knowledge clusters
  • Bilingual Support: Native Chinese and English entity extraction

Hybrid Retrieval Strategy

  • Multiple Query Modes:
    • auto: Automatically choose best strategy
    • kg_only: Graph traversal only
    • rag_only: Vector search only
    • hybrid: Combine KG and RAG results
    • kg_first: Try KG first, fallback to RAG
    • rag_first: Try RAG first, fallback to KG
  • Document-Scoped Queries: Ask questions about specific documents
  • Dynamic Weight Fusion: Intelligently balance KG and RAG results
  • Multi-Hop Graph Traversal: Configurable traversal depth (n_hops)

Intelligent Q&A

  • Context-Aware Answers: Leverages both graph structure and semantic meaning
  • Multi-Source Integration: Combines entities, relationships, and document chunks
  • Structured Output: Clear citations and context references
  • Bilingual Processing: Seamless Chinese and English query handling

Hierarchical Visualization

  • Overview Graph: High-level view of entire knowledge base
  • Document Graph: Per-document entity and relationship view
  • Entity Context: Detailed view of entity connections and properties
  • Interactive Features:
    • Force-directed graph layout
    • Real-time search and filtering
    • Node/edge highlighting on hover
    • Zoom and pan support

Observability & Monitoring

  • Session Tracking: Track multi-turn conversations with session IDs
  • LLM Call Monitoring: Detailed traces of all LLM interactions
  • Performance Analysis: Latency, token usage, and cost tracking
  • Experiment Tracking: Compare different prompts and configurations
  • Integration Options: Phoenix (dev/eval) or Langfuse (production)

License

MIT License

Contributors

Danielyan86

Issues