knaaptime/ocrbot

vscode extension to covert a pdf document into latex math using ollama/GLM

★ 0Forks 0TypeScriptGitHub ↗Compare

README

ocrbot

A VS Code Copilot chat extension that extracts text and LaTeX equations from academic PDFs using glm-ocr running locally via Ollama.

How it works

Most academic PDFs contain selectable text, but mathematical equations are often garbled or missing when extracted. ocrbot uses a hybrid approach:

  1. Extract text from all pages using pdfjs-dist
  2. Scan each page for equation indicators (Greek letters, math operators, non-ASCII clusters)
  3. Send only equation-heavy pages to glm-ocr via Ollama for full OCR
  4. Merge results back into page order

This means a 200-page paper with equations on 20 pages only sends 20 images to glm-ocr instead of 200.

Requirements

Host machine

brew install graphicsmagick ghostscript

or

sudo apt install graphicsmagick ghostscript

Ollama

ollama pull glm-ocr
ollama serve

Ollama must be running on localhost:11434 when using the extension.

Installation

  1. Download the .vsix file
  2. Install it:
code --install-extension ocrbot-0.2.0.vsix

Or via the Extensions panel: ... menu → Install from VSIX.

Usage

In the VS Code Copilot chat panel, use the @ocr agent:

@ocr /path/to/paper.pdf

To save the output as a markdown file next to the original PDF:

@ocr /path/to/paper.pdf save

This produces paper_ocr.md in the same directory as the PDF.

You can also use relative paths if you have a workspace open:

@ocr ref/paper.pdf save

Output

Pages extracted via plain text are output as-is. Pages processed by glm-ocr are marked with (glm-ocr) and contain LaTeX-rendered equations using $...$ for inline math and $$...$$ for display math.

A summary is shown on completion:

✅ Done — 26 pages, 25 via glm-ocr

💾 Saved to: /path/to/paper_ocr.md

Development

# Install dependencies
npm install

# Compile
npm run build

# Package
npm run package

# Install
npm run install:vsix

Project structure

src/
├── extension.ts      # VS Code chat participant, file path parsing, output streaming
├── ollamaClient.ts   # Ollama API calls to glm-ocr
└── pdfHandler.ts     # PDF text extraction, equation detection, image conversion

Troubleshooting

glm-ocr failed: image: unknown format The model received a raw PDF instead of an image. Make sure graphicsmagick and ghostscript are installed so pdf2pic can convert pages correctly.

No activated agent with id "ocr-agent.glm" The extension failed to load. Run Developer: Reload Window and check the Extensions panel for errors.

Ollama not responding Make sure Ollama is running: ollama serve

Issues