LAN-SHLOK/document_converter_full

A comprehensive utility tool for seamless, full-scale document format conversion.

โ˜… 2Forks 0PythonGitHub โ†—Compare
easyocrpdf2imagepymupdfpython-docxpython-pptxtesseract-ocr

README

๐Ÿ“„ Document Converter โ€“ OCR + Text Extraction (Tesseract + EasyOCR)

A powerful multi-format Document โ†’ Text converter built using:

  • Tesseract OCR
  • EasyOCR
  • PyMuPDF
  • pdf2image
  • python-docx / python-pptx

Supports:

  • ๐Ÿ–ผ๏ธ Images: JPG, PNG, TIFF, BMP, WEBP
  • ๐Ÿ“„ PDFs: Text-based & Scanned PDFs
  • ๐Ÿ“ DOCX files (Office Word)
  • ๐Ÿ“Š PPTX files (Office PowerPoint)

Also includes:

  • ๐Ÿ” Automatic Important Details Extractor (Emails, Phones, Dates, Key:Value pairs)
  • ๐Ÿ“ Automatic output folder generation
  • ๐Ÿงน Optional spell correction
  • ๐Ÿ–ฅ๏ธ Full Streamlit Web App UI

๐Ÿš€ Features

โœ” Convert ANY document to .txt

  • Multi-OCR merge: Tesseract + EasyOCR
  • PDF text-mode detection โ†’ uses direct extraction when possible

โœ” Important details extraction

Automatically extracts:

  • Emails
  • Phone numbers
  • Date formats
  • Key:Value structured text

โœ” Clean Web UI

Built with Streamlit, featuring:

  • File upload
  • OCR progress
  • Download TXT output
  • Download JSON details
  • Text preview panel

โœ” High accuracy

Image preprocessing:

  • Grayscale
  • Denoising
  • Adaptive thresholding

โœ” Works locally or in Docker


๐Ÿ“ฆ Requirements

Install system packages:

Windows

Linux

sudo apt install tesseract-ocr poppler-utils

Contributors

LAN-SHLOK

Issues