A Model Context Protocol (MCP) server for converting documents to markdown using the Docling library. This server enables Claude and other AI assistants to process and extract content from various document formats.
- Convert documents from URLs or local files to markdown
- Extract tables from documents
- Convert documents with embedded images
- Support for OCR (Optical Character Recognition)
- Batch processing of multiple documents
- Caching of conversion results for improved performance
- Hardware acceleration support (MPS on macOS)
- Python 3.10 or higher
- Docling library
- MCP library
-
Clone the repository:
git clone https://github.com/yourusername/mcp-docling.git cd mcp-docling -
Create a virtual environment:
python -m venv .venv source .venv/bin/activate # On Windows: .venv\Scripts\activate
-
Install the package:
pip install -e .
Run the server in development mode:
mcp dev mcp_docling/server.pyRun the server as a module:
python -m mcp_doclingTo use this server with Claude Desktop, add the following configuration to your Claude Desktop config file (located at ~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"docling": {
"command": "/path/to/your/python/environment/bin/python",
"args": [
"-m",
"mcp_docling"
],
"env": {
"PYTHONPATH": "/path/to/your/project/directory"
}
}
}
}Replace the paths with your actual Python environment and project directory paths.
Converts a document from a URL or local path to markdown format.
convert_document("https://arxiv.org/pdf/2408.09869")
convert_document("/path/to/document.pdf", enable_ocr=True, ocr_language=["en"])Converts a document and returns both markdown text and embedded images.
convert_document_with_images("https://arxiv.org/pdf/2408.09869")Extracts tables from a document and returns them as structured data.
extract_tables("https://arxiv.org/pdf/2408.09869")Converts multiple documents in batch mode.
convert_batch(["https://arxiv.org/pdf/2408.09869", "/path/to/document.pdf"])Returns information about the system configuration and acceleration status.
get_system_info()You can test the server using the provided test script:
python test_docling_server.pyMake sure to set the required environment variables before running the test:
export INFERENCE_MODEL="your-model-id"
export LLAMA_STACK_PORT="8080"The server supports various configuration options:
- OCR support with language selection
- Hardware acceleration (MPS on macOS)
- Caching of conversion results
- Batch processing settings
If you encounter issues:
- Check the logs for error messages
- Verify that your Python environment has all required dependencies
- Ensure the PYTHONPATH is correctly set in your configuration
- For hardware acceleration issues, check that your system supports the configured accelerator
[Your License Here]
Contributions are welcome! Please feel free to submit a Pull Request.