This tool converts PDF files to Markdown format using Mistral OCR API, preserving images and text structure.
Install the package using pip:
pip install pdf2mdYou need to set your Mistral API key to use this tool. You can do this in two ways:
- Set it as an environment variable:
export MISTRAL_API_KEY=your-api-key-here- Pass it directly as a command-line argument:
pdf2md --api-key your-api-key-here document.pdfBasic usage:
pdf2md document.pdfConvert multiple files:
pdf2md document1.pdf document2.pdf document3.pdfConvert all PDF files in a directory:
pdf2md *.pdfDisable progress bars:
pdf2md document.pdf --no-progress- Converts PDF documents to Markdown format
- Extracts and saves images from PDF files
- Provides progress bars for better user feedback
- Handles multiple files in a single command
- Robust error handling that continues processing other files even if one fails
- URL-encoded image paths for proper Markdown rendering
For each input PDF file, the tool will create:
- A Markdown file with the same name (but .md extension)
- A directory containing extracted images, named after the PDF file with
_imagessuffix
For example, for document.pdf, the output will be:
document.md- the main Markdown filedocument_images/- directory containing all extracted images