A powerful TypeScript-based OCR (Optical Character Recognition) tool that extracts text from images, specifically optimized for old newspapers and scanned documents. Built with Bun.js, Tesseract.js, and Sharp.
- ๐ Multiple OCR methods for best results
- ๐ฐ Optimized for old newspapers and poor-quality scans
- ๐ผ๏ธ Advanced image preprocessing (contrast enhancement, noise reduction, sharpening)
- ๐ Automatic image upscaling for small images
- ๐ฏ Multiple page segmentation modes
- ๐ Confidence scoring and method comparison
- โก Fast execution with Bun.js runtime
try-tesseract/
โโ test-images/ # Test images directory (auto-generated during tests)
โ โโ test-image.jpg # Sample test image
โโ .github/ # GitHub Actions workflows
โ โโ workflows/
โ โโ ci.yml # CI/CD pipeline configuration
โโ .gitignore # Git ignore file
โโ bun.lockb # Bun lock file for dependencies
โโ eng.traineddata # Tesseract English language data (auto-downloaded)
โโ ocr.test.ts # Unit tests for OCR functionality
โโ ocr.ts # Main OCR application script
โโ utils.ts # Utility functions for image processing
โโ package.json # Node package configuration and scripts
โโ README.md # Project documentation (this file)
โโ tsconfig.json # TypeScript configurationThis application requires Tesseract OCR to be installed on your system. Tesseract.js uses the native Tesseract engine under the hood.
brew install tesseractsudo apt update
sudo apt install tesseract-ocr
sudo apt install libtesseract-dev- Download the installer from GitHub Releases
- Run the installer and add Tesseract to your PATH
- Or use Chocolatey:
choco install tesseractsudo pacman -S tesseractsudo dnf install tesseractVerify Installation:
tesseract --versionFor more installation details, visit: https://tesseract-ocr.github.io/tessdoc/Installation.html
# macOS/Linux
curl -fsSL https://bun.sh/install | bash
# Windows
powershell -c "irm bun.sh/install.ps1 | iex"Verify Installation:
bun --version-
Clone or download this repository
-
Install dependencies:
bun install tesseract.js sharp- Save the OCR script as
ocr.ts
Extract text with default settings (aggressive preprocessing):
bun run ocr.ts image.pngThis will try all preprocessing and segmentation combinations and return the best result:
bun run ocr.ts newspaper.png --allUse specific page segmentation mode and preprocessing:
# Single column text with aggressive preprocessing (good for newspapers)
bun run ocr.ts newspaper.png --psm 4 --prep aggressive
# Single block with threshold preprocessing
bun run ocr.ts newspaper.png --psm 6 --prep threshold
# Automatic segmentation with basic preprocessing
bun run ocr.ts document.png --psm 3 --prep basicbun run ocr.ts --help| Option | Description | Values |
|---|---|---|
--all |
Try all methods and return best result | - |
--psm <mode> |
Tesseract page segmentation mode | 1, 3, 4, 6, 11 |
--prep <type> |
Image preprocessing method | basic, aggressive, threshold |
- PSM 1: Automatic page segmentation with orientation detection
- PSM 3: Fully automatic page segmentation (default)
- PSM 4: Single column of text (best for newspapers)
- PSM 6: Single uniform block of text
- PSM 11: Sparse text detection
- basic: Light enhancement (normalize + sharpen)
- aggressive: Heavy processing for poor quality (contrast boost + noise reduction)
- threshold: Black and white conversion (best for very old documents)
bun run ocr.ts old-newspaper.png --allOutput:
==========================================================
Trying: PSM 3 (Auto) + Aggressive preprocessing
==========================================================
Progress: 100%
Confidence: 87.45%
Characters extracted: 1247
==========================================================
Best method: PSM 4 (Single column) + Basic preprocessing
Confidence: 89.32%
==========================================================
==========================================================
EXTRACTED TEXT:
==========================================================
Union Negro ....
...
bun run ocr.ts document.png --psm 4 --prep basicbun run ocr.ts poor-quality.png --psm 6 --prep threshold-
Image Preprocessing: The image is enhanced using Sharp:
- Upscaling (if needed)
- Grayscale conversion
- Contrast normalization
- Sharpening
- Noise reduction (optional)
- Thresholding (optional)
-
OCR Processing: Tesseract analyzes the preprocessed image with specified segmentation mode
-
Result Selection: When using
--all, the tool compares all methods and selects the one with the best score (length ร confidence)
- Old Newspapers: Use
--allor--psm 4 --prep aggressive - Clean Documents: Use
--psm 3 --prep basic - Poor Quality Scans: Use
--prep thresholdor--prep aggressive - Small Images: The tool automatically upscales, but higher resolution inputs work better
- Multiple Columns: Try
--psm 3for automatic detection
- Ensure Tesseract OCR is properly installed
- Verify with
tesseract --version - Check that language data files are installed
bun install sharp# Make sure Tesseract is installed and working
tesseract --version
# Clean install dependencies
rm -rf node_modules bun.lock
bun install# Auto-fix most issues
bun run lint:fix
bun run format- Try
--allto find the best method - Ensure input image is high resolution (at least 300 DPI)
- Try different preprocessing methods
- Check if the image is properly oriented
- Image quality may be too poor
- Try
--prep aggressiveor--prep threshold - Consider manually cleaning the image first
The test suite includes:
- โ Image preprocessing tests
- โ Text extraction tests
- โ Multiple PSM mode tests
- โ Image format support tests
- โ Edge case handling
- โ Performance benchmarks
# Run all tests
bun test
# Run with coverage
bun test --coverage
# Run specific test
bun test ocr.test.tsNote: Tests create temporary images in ./test-images/ directory and clean up automatically.
This project includes a GitHub Actions workflow that automatically:
- โ Lints code on every push and pull request
- โ Checks formatting to ensure code consistency
- โ Runs unit tests to verify functionality
- โ Type checks TypeScript code
- โ Generates test coverage reports
The CI pipeline runs on:
- Push to
mainordevelopbranches - Pull requests targeting
mainordevelop
Jobs:
- lint-and-format: Checks code quality and formatting
- test: Runs unit tests with Tesseract OCR
- build: Verifies TypeScript compilation
- coverage: Generates test coverage report (on push only)
To view CI status, check the Actions tab in your GitHub repository.
Add this badge to your README to show CI status:
- Bun.js - Fast JavaScript runtime
- Tesseract.js - OCR library
- Sharp - Image processing
- Tesseract OCR - OCR engine
Contributions are welcome! Here's how you can help:
-
Fork the repository
-
Clone your fork:
git clone https://github.com/yourusername/try-tesseract.git
cd try-tesseract- Install dependencies:
bun install- Create a new branch:
git checkout -b feature/your-feature-name-
Make your changes
-
Run linter and formatter:
bun run lint:fix
bun run format-
Add tests for new features
-
Run tests:
bun test- Commit your changes:
git add .
git commit -m "feat: add your feature description"- Push to your fork:
git push origin feature/your-feature-name- Create a Pull Request
Follow conventional commits:
feat:- New featurefix:- Bug fixdocs:- Documentation changestest:- Test changesrefactor:- Code refactoringstyle:- Code style changes (formatting)chore:- Build process or auxiliary tool changes
- Use TypeScript strict mode
- Follow the existing code style
- Add JSDoc comments for public functions
- Keep functions small and focused
- Write meaningful variable names
- Write unit tests for all new functions
- Ensure all tests pass before submitting PR
- Aim for high test coverage
- Test edge cases and error conditions
- Keep PRs focused on a single feature/fix
- Update documentation if needed
- Ensure CI checks pass
- Respond to review feedback promptly
MIT
If you find this project helpful, please consider:
- โญ Starring the repository
- ๐ Reporting bugs
- ๐ก Suggesting new features
- ๐ Improving documentation