CruzR/ocr_pdftotext

A replacement for pdftotext that calls tesseract instead.

★ 1Forks 0PythonGitHub ↗Compare

README

ocr_pdftotext

This is a drop-in for the part of the pdftotext command line interface used by Zotero. It auto-detects when pdftotext fails on a PDF file and instead performs OCR using the open source tesseract OCR engine.

Contributors

CruzR

Issues