rjoydip/try-tesseract

A powerful TypeScript-based OCR (Optical Character Recognition) tool that extracts text from images, specifically optimized for old newspapers and scanned documents. Built with Bun.js, Tesseract.js, and Sharp.

โ˜… 0Forks 0TypeScriptGitHub โ†—Compare

README

Tesseract OCR Text Extractor

A powerful TypeScript-based OCR (Optical Character Recognition) tool that extracts text from images, specifically optimized for old newspapers and scanned documents. Built with Bun.js, Tesseract.js, and Sharp.

Features

  • ๐Ÿ” Multiple OCR methods for best results
  • ๐Ÿ“ฐ Optimized for old newspapers and poor-quality scans
  • ๐Ÿ–ผ๏ธ Advanced image preprocessing (contrast enhancement, noise reduction, sharpening)
  • ๐Ÿ“ Automatic image upscaling for small images
  • ๐ŸŽฏ Multiple page segmentation modes
  • ๐Ÿ“Š Confidence scoring and method comparison
  • โšก Fast execution with Bun.js runtime

Project Structure

try-tesseract/
โ”œโ”€ test-images/             # Test images directory (auto-generated during tests)
โ”‚  โ””โ”€ test-image.jpg        # Sample test image
โ”œโ”€ .github/                 # GitHub Actions workflows
โ”‚  โ””โ”€ workflows/
โ”‚     โ””โ”€ ci.yml             # CI/CD pipeline configuration
โ”œโ”€ .gitignore               # Git ignore file
โ”œโ”€ bun.lockb                # Bun lock file for dependencies
โ”œโ”€ eng.traineddata          # Tesseract English language data (auto-downloaded)
โ”œโ”€ ocr.test.ts              # Unit tests for OCR functionality
โ”œโ”€ ocr.ts                   # Main OCR application script
โ”œโ”€ utils.ts                 # Utility functions for image processing
โ”œโ”€ package.json             # Node package configuration and scripts
โ”œโ”€ README.md                # Project documentation (this file)
โ””โ”€ tsconfig.json            # TypeScript configuration

Prerequisites

1. Install Tesseract OCR Engine

This application requires Tesseract OCR to be installed on your system. Tesseract.js uses the native Tesseract engine under the hood.

macOS

brew install tesseract

Ubuntu/Debian

sudo apt update
sudo apt install tesseract-ocr
sudo apt install libtesseract-dev

Windows

  1. Download the installer from GitHub Releases
  2. Run the installer and add Tesseract to your PATH
  3. Or use Chocolatey:
choco install tesseract

Arch Linux

sudo pacman -S tesseract

Fedora/RHEL

sudo dnf install tesseract

Verify Installation:

tesseract --version

For more installation details, visit: https://tesseract-ocr.github.io/tessdoc/Installation.html

2. Install Bun.js

# macOS/Linux
curl -fsSL https://bun.sh/install | bash

# Windows
powershell -c "irm bun.sh/install.ps1 | iex"

Verify Installation:

bun --version

Installation

  1. Clone or download this repository

  2. Install dependencies:

bun install tesseract.js sharp
  1. Save the OCR script as ocr.ts

Usage

Basic Usage

Extract text with default settings (aggressive preprocessing):

bun run ocr.ts image.png

Try All Methods (Recommended for Old Newspapers)

This will try all preprocessing and segmentation combinations and return the best result:

bun run ocr.ts newspaper.png --all

Specify Custom Settings

Use specific page segmentation mode and preprocessing:

# Single column text with aggressive preprocessing (good for newspapers)
bun run ocr.ts newspaper.png --psm 4 --prep aggressive

# Single block with threshold preprocessing
bun run ocr.ts newspaper.png --psm 6 --prep threshold

# Automatic segmentation with basic preprocessing
bun run ocr.ts document.png --psm 3 --prep basic

Help

bun run ocr.ts --help

Command Line Options

Option Description Values
--all Try all methods and return best result -
--psm <mode> Tesseract page segmentation mode 1, 3, 4, 6, 11
--prep <type> Image preprocessing method basic, aggressive, threshold

Page Segmentation Modes (PSM)

  • PSM 1: Automatic page segmentation with orientation detection
  • PSM 3: Fully automatic page segmentation (default)
  • PSM 4: Single column of text (best for newspapers)
  • PSM 6: Single uniform block of text
  • PSM 11: Sparse text detection

Preprocessing Methods

  • basic: Light enhancement (normalize + sharpen)
  • aggressive: Heavy processing for poor quality (contrast boost + noise reduction)
  • threshold: Black and white conversion (best for very old documents)

Examples

Example 1: Old Newspaper Scan

bun run ocr.ts old-newspaper.png --all

Output:

==========================================================
Trying: PSM 3 (Auto) + Aggressive preprocessing
==========================================================
Progress: 100%
Confidence: 87.45%
Characters extracted: 1247

==========================================================
Best method: PSM 4 (Single column) + Basic preprocessing
Confidence: 89.32%
==========================================================

==========================================================
EXTRACTED TEXT:
==========================================================
Union Negro ....
...

Example 2: Document with Single Column

bun run ocr.ts document.png --psm 4 --prep basic

Example 3: Poor Quality Scan

bun run ocr.ts poor-quality.png --psm 6 --prep threshold

How It Works

  1. Image Preprocessing: The image is enhanced using Sharp:

    • Upscaling (if needed)
    • Grayscale conversion
    • Contrast normalization
    • Sharpening
    • Noise reduction (optional)
    • Thresholding (optional)
  2. OCR Processing: Tesseract analyzes the preprocessed image with specified segmentation mode

  3. Result Selection: When using --all, the tool compares all methods and selects the one with the best score (length ร— confidence)

Tips for Best Results

  • Old Newspapers: Use --all or --psm 4 --prep aggressive
  • Clean Documents: Use --psm 3 --prep basic
  • Poor Quality Scans: Use --prep threshold or --prep aggressive
  • Small Images: The tool automatically upscales, but higher resolution inputs work better
  • Multiple Columns: Try --psm 3 for automatic detection

Troubleshooting

Error: "Tesseract couldn't load any languages"

  • Ensure Tesseract OCR is properly installed
  • Verify with tesseract --version
  • Check that language data files are installed

Error: "Cannot find module 'sharp'"

bun install sharp

Test Failures

# Make sure Tesseract is installed and working
tesseract --version

# Clean install dependencies
rm -rf node_modules bun.lock
bun install

Linting/Formatting Issues

# Auto-fix most issues
bun run lint:fix
bun run format

Poor OCR Results

  • Try --all to find the best method
  • Ensure input image is high resolution (at least 300 DPI)
  • Try different preprocessing methods
  • Check if the image is properly oriented

Low Confidence Scores

  • Image quality may be too poor
  • Try --prep aggressive or --prep threshold
  • Consider manually cleaning the image first

Running Tests

The test suite includes:

  • โœ… Image preprocessing tests
  • โœ… Text extraction tests
  • โœ… Multiple PSM mode tests
  • โœ… Image format support tests
  • โœ… Edge case handling
  • โœ… Performance benchmarks
# Run all tests
bun test

# Run with coverage
bun test --coverage

# Run specific test
bun test ocr.test.ts

Note: Tests create temporary images in ./test-images/ directory and clean up automatically.

CI/CD Pipeline

This project includes a GitHub Actions workflow that automatically:

  • โœ… Lints code on every push and pull request
  • โœ… Checks formatting to ensure code consistency
  • โœ… Runs unit tests to verify functionality
  • โœ… Type checks TypeScript code
  • โœ… Generates test coverage reports

GitHub Actions Workflow

The CI pipeline runs on:

  • Push to main or develop branches
  • Pull requests targeting main or develop

Jobs:

  1. lint-and-format: Checks code quality and formatting
  2. test: Runs unit tests with Tesseract OCR
  3. build: Verifies TypeScript compilation
  4. coverage: Generates test coverage report (on push only)

To view CI status, check the Actions tab in your GitHub repository.

Adding CI Badge

Add this badge to your README to show CI status:

![CI](https://github.com/rjoydip/try-tesseract/workflows/CI/badge.svg)

Dependencies

Contributing

Contributions are welcome! Here's how you can help:

Getting Started

  1. Fork the repository

  2. Clone your fork:

git clone https://github.com/yourusername/try-tesseract.git
cd try-tesseract
  1. Install dependencies:
bun install
  1. Create a new branch:
git checkout -b feature/your-feature-name

Development Workflow

  1. Make your changes

  2. Run linter and formatter:

bun run lint:fix
bun run format
  1. Add tests for new features

  2. Run tests:

bun test
  1. Commit your changes:
git add .
git commit -m "feat: add your feature description"
  1. Push to your fork:
git push origin feature/your-feature-name
  1. Create a Pull Request

Commit Message Convention

Follow conventional commits:

  • feat: - New feature
  • fix: - Bug fix
  • docs: - Documentation changes
  • test: - Test changes
  • refactor: - Code refactoring
  • style: - Code style changes (formatting)
  • chore: - Build process or auxiliary tool changes

Code Style

  • Use TypeScript strict mode
  • Follow the existing code style
  • Add JSDoc comments for public functions
  • Keep functions small and focused
  • Write meaningful variable names

Testing Guidelines

  • Write unit tests for all new functions
  • Ensure all tests pass before submitting PR
  • Aim for high test coverage
  • Test edge cases and error conditions

Pull Request Guidelines

  • Keep PRs focused on a single feature/fix
  • Update documentation if needed
  • Ensure CI checks pass
  • Respond to review feedback promptly

License

MIT

Additional Resources

Support

If you find this project helpful, please consider:

  • โญ Starring the repository
  • ๐Ÿ› Reporting bugs
  • ๐Ÿ’ก Suggesting new features
  • ๐Ÿ“– Improving documentation

Contributors

rjoydip

Issues