๐ง An intelligent MCP server that transforms your documents into a searchable knowledge base with AI-powered note generation and Obsidian integration.
The Document Miner MCP Server is a powerful tool that automatically processes your PDF documents, creates semantic embeddings, and generates intelligent study notes. It seamlessly integrates with Obsidian and provides a robust API for knowledge retrieval and note creation.
- ๐ Semantic Search: Advanced vector-based search using sentence transformers
- ๐ AI Note Generation: Automatically create comprehensive study notes from PDFs
- ๐ Obsidian Integration: Direct integration with Obsidian for seamless note management
- ๐ Multi-Format Support: Process PDFs, textbooks, and research papers
- ๐ง Intelligent Chunking: Advanced text chunking with LangChain for optimal context preservation
- โก MCP Protocol: Built on the Model Context Protocol for easy integration with AI assistants
- ๐๏ธ Vector Database: ChromaDB integration for fast similarity search
- ๐ง Configurable: Flexible configuration for different use cases
graph LR
A[๐ PDF Documents] --> B[๐ PDF Processing]
B --> C[๐งฉ Text Chunking]
C --> D[๐ง Vector Embeddings]
D --> E[๐๏ธ ChromaDB]
F[๐ค AI Assistant] <--> G[๐ MCP Server]
G <--> H{๐ MCP Tools}
H <--> I[๐ Search Tool]
H <--> J[๐ Note Creator]
H <--> K[๐ Study Assistant]
I <--> E
J <--> E
J --> L[๐ Obsidian Vault]
K <--> E
style A fill:#e1f5fe
style E fill:#f3e5f5
style G fill:#e8f5e8
style L fill:#fff3e0
- Python 3.8 or higher
- Git
-
Clone the repository
git clone https://github.com/yourusername/knowledge-base-mcp-server.git cd knowledge-base-mcp-server -
Install dependencies
python install_deps.py
Or manually:
pip install -r requirements.txt
-
Set up configuration
cp config.env.template .env # Edit .env with your settings -
Run the setup
python quick_setup.py
-
Place PDFs in the knowledge base directory:
knowledge_base/ โโโ pdfs/ # General documents โโโ textbooks/ # Academic textbooks -
Process documents, convert to chunks and push to vector database:
from integrations.pdf_integration import PDFIntegration from config import Config config = Config() pdf_integration = PDFIntegration(config) # Process a directory await pdf_integration.process_pdf_directory("path/to/pdfs")
# Search for information
results = await pdf_integration.search_knowledge_base(
query="machine learning algorithms",
max_results=10
)
for result in results:
print(f"Source: {result.source_document}")
print(f"Content: {result.content}")
print(f"Similarity: {result.similarity_score}")# Create Obsidian notes from topics
note = await pdf_integration.create_note_from_topic(
topic="Neural Networks",
note_type="detailed",
focus_areas=["architecture", "training", "applications"]
)The system uses a hierarchical configuration system. Key settings include:
# Obsidian Integration
OBSIDIAN_VAULT_PATH="/path/to/your/vault"
OBSIDIAN_NOTES_FOLDER="Knowledge_Base_Notes"
# PDF Processing
PDF_DIRECTORY="knowledge_base/pdfs"
TEXTBOOK_DIRECTORY="knowledge_base/textbooks"
# Vector Database
VECTOR_DB_PATH="knowledge_base/vector_db"
SIMILARITY_THRESHOLD=0.7
MAX_SEARCH_RESULTS=10
# AI Model Settings
EMBEDDING_MODEL="all-MiniLM-L6-v2"
CHUNK_SIZE=1000
CHUNK_OVERLAP=200See config.py for detailed configuration options including:
- Chunking strategies
- Embedding models
- Search parameters
- Obsidian settings
This server implements the Model Context Protocol, making it compatible with various AI assistants:
search_knowledge_base- Search for information in the knowledge basecreate_note_from_topic- Generate notes on specific topicscreate_note_with_content- Create custom notes with LLM-generated contentget_semantic_chunks- Retrieve raw semantic chunksprocess_pdf_directory- Process PDFs from a directoryget_study_suggestions- Get study recommendations
{
"method": "tools/call",
"params": {
"name": "search_knowledge_base",
"arguments": {
"query": "data structures and algorithms",
"max_results": 5
}
}
}-
Navigate to Settings->Tools and Integrations and click new MCP server

-
Add this in mcp.json and save it.
{
"mcpServers": {
"knowledge-base": {
"command": "python",
"args": ["<PATH_TO_PROJECT_FOLDER>/run_mcp_server.py"]
}
}
}doc-miner-mcp-server/
โโโ ๐ server.py # Main MCP server
โโโ ๐ config.py # Configuration management
โโโ ๐ run_mcp_server.py # MCP server runner
โโโ ๐ integrations/
โ โโโ ๐ pdf_integration.py # PDF processing and search
โ โโโ ๐ obsidian_integration.py # Obsidian note creation
โโโ ๐ models/
โ โโโ ๐ knowledge_models.py # Data models
โโโ ๐ knowledge_base/
โ โโโ ๐ pdfs/ # PDF documents (gitignored)
โ โโโ ๐ textbooks/ # Textbook PDFs (gitignored)
โ โโโ ๐ vector_db/ # Vector database (gitignored)
โโโ ๐ requirements.txt # Python dependencies
โโโ ๐ install_deps.py # Dependency installer
โโโ ๐ quick_setup.py # Quick setup script
โโโ ๐ config.env.template # Configuration template
- Create a new integration in
integrations/ - Add configuration options to
config.py - Update the MCP server tools in
server.py - Add tests for new functionality
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
# Clone your fork
git clone https://github.com/yourusername/knowledge-base-mcp-server.git
cd knowledge-base-mcp-server
# Install development dependencies
pip install -r requirements.txt
python install_deps.py
# Set up pre-commit hooks (optional)
pip install pre-commit
pre-commit install- Process lecture PDFs and textbooks
- Generate study notes automatically
- Create comprehensive study guides
- Build a searchable knowledge base
- Process research papers and publications
- Extract key concepts and methodologies
- Create literature review notes
- Build domain-specific knowledge bases
- Process technical documentation
- Create training materials
- Build organizational knowledge bases
- Generate summarized reports
-
Import Errors
python install_deps.py # or pip install --upgrade sentence-transformers chromadb -
ChromaDB Issues
# Clear the database and restart rm -rf knowledge_base/vector_db/ python quick_setup.py -
Obsidian Integration Not Working
- Check
OBSIDIAN_VAULT_PATHin.env - Ensure the vault directory exists
- Verify folder permissions
- Check
- ๐ Report bugs in Issues
- ๐ฌ Join discussions in Discussions
This project is licensed under the MIT License - see the LICENSE file for details.
- Model Context Protocol for the MCP specification
- LangChain for text processing capabilities
- ChromaDB for vector database functionality
- Sentence Transformers for embeddings
- Obsidian for the amazing note-taking platform
If you find this project useful, please consider giving it a star! โญ
