A powerful multi-format Document โ Text converter built using:
- Tesseract OCR
- EasyOCR
- PyMuPDF
- pdf2image
- python-docx / python-pptx
Supports:
- ๐ผ๏ธ Images: JPG, PNG, TIFF, BMP, WEBP
- ๐ PDFs: Text-based & Scanned PDFs
- ๐ DOCX files (Office Word)
- ๐ PPTX files (Office PowerPoint)
Also includes:
- ๐ Automatic Important Details Extractor (Emails, Phones, Dates, Key:Value pairs)
- ๐ Automatic output folder generation
- ๐งน Optional spell correction
- ๐ฅ๏ธ Full Streamlit Web App UI
- Multi-OCR merge: Tesseract + EasyOCR
- PDF text-mode detection โ uses direct extraction when possible
Automatically extracts:
- Emails
- Phone numbers
- Date formats
- Key:Value structured text
Built with Streamlit, featuring:
- File upload
- OCR progress
- Download TXT output
- Download JSON details
- Text preview panel
Image preprocessing:
- Grayscale
- Denoising
- Adaptive thresholding
Install system packages:
- Tesseract OCR โ https://github.com/UB-Mannheim/tesseract/wiki
- Poppler for Windows โ add
bin/to PATH
sudo apt install tesseract-ocr poppler-utils