hthadicherla/DocMinerMCP

โ˜… 0Forks 0PythonGitHub โ†—Compare

README

Document Miner MCP Server

๐Ÿง  An intelligent MCP server that transforms your documents into a searchable knowledge base with AI-powered note generation and Obsidian integration.

Python MCP License Obsidian

๐ŸŒŸ Overview

The Document Miner MCP Server is a powerful tool that automatically processes your PDF documents, creates semantic embeddings, and generates intelligent study notes. It seamlessly integrates with Obsidian and provides a robust API for knowledge retrieval and note creation.

โœจ Key Features

  • ๐Ÿ” Semantic Search: Advanced vector-based search using sentence transformers
  • ๐Ÿ“ AI Note Generation: Automatically create comprehensive study notes from PDFs
  • ๐Ÿ”— Obsidian Integration: Direct integration with Obsidian for seamless note management
  • ๐Ÿ“š Multi-Format Support: Process PDFs, textbooks, and research papers
  • ๐Ÿง  Intelligent Chunking: Advanced text chunking with LangChain for optimal context preservation
  • โšก MCP Protocol: Built on the Model Context Protocol for easy integration with AI assistants
  • ๐Ÿ—„๏ธ Vector Database: ChromaDB integration for fast similarity search
  • ๐Ÿ”ง Configurable: Flexible configuration for different use cases

๐Ÿ—๏ธ Architecture Overview

graph LR
    A[๐Ÿ“„ PDF Documents] --> B[๐Ÿ”„ PDF Processing]
    B --> C[๐Ÿงฉ Text Chunking]
    C --> D[๐Ÿง  Vector Embeddings]
    D --> E[๐Ÿ—„๏ธ ChromaDB]
    
    F[๐Ÿค– AI Assistant] <--> G[๐Ÿ”Œ MCP Server]
    G <--> H{๐Ÿ“‹ MCP Tools}
    
    H <--> I[๐Ÿ” Search Tool]
    H <--> J[๐Ÿ“ Note Creator]
    H <--> K[๐Ÿ“š Study Assistant]
    
    I <--> E
    J <--> E
    J --> L[๐Ÿ““ Obsidian Vault]
    K <--> E
    
    style A fill:#e1f5fe
    style E fill:#f3e5f5
    style G fill:#e8f5e8
    style L fill:#fff3e0
Loading

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.8 or higher
  • Git

Installation

  1. Clone the repository

    git clone https://github.com/yourusername/knowledge-base-mcp-server.git
    cd knowledge-base-mcp-server
  2. Install dependencies

    python install_deps.py

    Or manually:

    pip install -r requirements.txt
  3. Set up configuration

    cp config.env.template .env
    # Edit .env with your settings
  4. Run the setup

    python quick_setup.py

๐Ÿ“– Usage

Adding Documents

  1. Place PDFs in the knowledge base directory:

    knowledge_base/
    โ”œโ”€โ”€ pdfs/              # General documents
    โ””โ”€โ”€ textbooks/         # Academic textbooks
    
  2. Process documents, convert to chunks and push to vector database:

    from integrations.pdf_integration import PDFIntegration
    from config import Config
    
    config = Config()
    pdf_integration = PDFIntegration(config)
    
    # Process a directory
    await pdf_integration.process_pdf_directory("path/to/pdfs")

Searching Knowledge Base

# Search for information
results = await pdf_integration.search_knowledge_base(
    query="machine learning algorithms",
    max_results=10
)

for result in results:
    print(f"Source: {result.source_document}")
    print(f"Content: {result.content}")
    print(f"Similarity: {result.similarity_score}")

Creating Notes

# Create Obsidian notes from topics
note = await pdf_integration.create_note_from_topic(
    topic="Neural Networks",
    note_type="detailed",
    focus_areas=["architecture", "training", "applications"]
)

โš™๏ธ Configuration

The system uses a hierarchical configuration system. Key settings include:

Environment Variables (.env)

# Obsidian Integration
OBSIDIAN_VAULT_PATH="/path/to/your/vault"
OBSIDIAN_NOTES_FOLDER="Knowledge_Base_Notes"

# PDF Processing
PDF_DIRECTORY="knowledge_base/pdfs"
TEXTBOOK_DIRECTORY="knowledge_base/textbooks"

# Vector Database
VECTOR_DB_PATH="knowledge_base/vector_db"
SIMILARITY_THRESHOLD=0.7
MAX_SEARCH_RESULTS=10

# AI Model Settings
EMBEDDING_MODEL="all-MiniLM-L6-v2"
CHUNK_SIZE=1000
CHUNK_OVERLAP=200

Advanced Configuration

See config.py for detailed configuration options including:

  • Chunking strategies
  • Embedding models
  • Search parameters
  • Obsidian settings

๐Ÿ”Œ MCP Integration

This server implements the Model Context Protocol, making it compatible with various AI assistants:

Available Tools

  • search_knowledge_base - Search for information in the knowledge base
  • create_note_from_topic - Generate notes on specific topics
  • create_note_with_content - Create custom notes with LLM-generated content
  • get_semantic_chunks - Retrieve raw semantic chunks
  • process_pdf_directory - Process PDFs from a directory
  • get_study_suggestions - Get study recommendations

Example MCP Usage

{
  "method": "tools/call",
  "params": {
    "name": "search_knowledge_base",
    "arguments": {
      "query": "data structures and algorithms",
      "max_results": 5
    }
  }
}

Steps to integrate it into Cursor/Claude Desktop

  1. Navigate to Settings->Tools and Integrations and click new MCP server

  2. Add this in mcp.json and save it.

{
  "mcpServers": {
    "knowledge-base": {
      "command": "python",
      "args": ["<PATH_TO_PROJECT_FOLDER>/run_mcp_server.py"]
    }
  }
}
  1. Toggle the MCP and enable it

  2. Now the code assistant in client can access all the MCP tools.

๐Ÿ“ Project Structure

doc-miner-mcp-server/
โ”œโ”€โ”€ ๐Ÿ“„ server.py                 # Main MCP server
โ”œโ”€โ”€ ๐Ÿ“„ config.py                 # Configuration management
โ”œโ”€โ”€ ๐Ÿ“„ run_mcp_server.py         # MCP server runner
โ”œโ”€โ”€ ๐Ÿ“ integrations/
โ”‚   โ”œโ”€โ”€ ๐Ÿ“„ pdf_integration.py    # PDF processing and search
โ”‚   โ””โ”€โ”€ ๐Ÿ“„ obsidian_integration.py # Obsidian note creation
โ”œโ”€โ”€ ๐Ÿ“ models/
โ”‚   โ””โ”€โ”€ ๐Ÿ“„ knowledge_models.py   # Data models
โ”œโ”€โ”€ ๐Ÿ“ knowledge_base/
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ pdfs/                 # PDF documents (gitignored)
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ textbooks/            # Textbook PDFs (gitignored)
โ”‚   โ””โ”€โ”€ ๐Ÿ“ vector_db/            # Vector database (gitignored)
โ”œโ”€โ”€ ๐Ÿ“„ requirements.txt          # Python dependencies
โ”œโ”€โ”€ ๐Ÿ“„ install_deps.py           # Dependency installer
โ”œโ”€โ”€ ๐Ÿ“„ quick_setup.py            # Quick setup script
โ””โ”€โ”€ ๐Ÿ“„ config.env.template       # Configuration template

๐Ÿ”ง Development

Adding New Features

  1. Create a new integration in integrations/
  2. Add configuration options to config.py
  3. Update the MCP server tools in server.py
  4. Add tests for new functionality

๐Ÿค Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Development Setup

# Clone your fork
git clone https://github.com/yourusername/knowledge-base-mcp-server.git
cd knowledge-base-mcp-server

# Install development dependencies
pip install -r requirements.txt
python install_deps.py

# Set up pre-commit hooks (optional)
pip install pre-commit
pre-commit install

๐Ÿ“š Use Cases

๐ŸŽ“ Students

  • Process lecture PDFs and textbooks
  • Generate study notes automatically
  • Create comprehensive study guides
  • Build a searchable knowledge base

๐Ÿ”ฌ Researchers

  • Process research papers and publications
  • Extract key concepts and methodologies
  • Create literature review notes
  • Build domain-specific knowledge bases

๐Ÿ“– Knowledge Workers

  • Process technical documentation
  • Create training materials
  • Build organizational knowledge bases
  • Generate summarized reports

๐Ÿ› ๏ธ Troubleshooting

Common Issues

  1. Import Errors

    python install_deps.py
    # or
    pip install --upgrade sentence-transformers chromadb
  2. ChromaDB Issues

    # Clear the database and restart
    rm -rf knowledge_base/vector_db/
    python quick_setup.py
  3. Obsidian Integration Not Working

    • Check OBSIDIAN_VAULT_PATH in .env
    • Ensure the vault directory exists
    • Verify folder permissions

Getting Help

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

๐ŸŒŸ Star History

If you find this project useful, please consider giving it a star! โญ

Contributors

hthadicherla

Issues