In our daily life, sometimes we want to convert some pdf files to txt files, especially when these pdf files are old scanned version. This script is designed to finish this task.
Several days ago (before 2018.11.11), Professor Yan Feng asked a related question in Weibo. He asked if there is some free software that can convert scanned version PDF to txt. Together with the reason mentioned above, I decided to write such a script to do this task.
The script originally relied on the Tencent AI open platform OCR API (ai.qq.com), which has since been shut down. It has been refactored to follow the current best practice for PDF-to-text conversion:
- Text-layer fast path: if a page contains an embedded text layer, the text is extracted directly with PyMuPDF — no OCR at all, which is orders of magnitude faster.
- Local OCR for scanned pages: pages without a text layer are rendered in memory with PyMuPDF and recognized locally with Tesseract. There are no network round-trips, no API rate limits (the old script had to sleep 1 second per page), and no intermediate image/json files written to disk.
- Parallel processing: all pages are processed concurrently in a process pool, using every CPU core.
Install the Tesseract OCR engine first (only needed for scanned PDFs without a text layer):
- Linux:
sudo apt install tesseract-ocr tesseract-ocr-chi-sim - macOS:
brew install tesseract tesseract-lang - Windows: download the installer from UB-Mannheim/tesseract and select the Chinese (Simplified) language pack.
Then install the Python packages:
pip install -r requirements.txtfor each page (in parallel): extract embedded text layer if present, otherwise render the page in memory and OCR it locally -> rebuild paragraphs from line positions -> concatenate pages into one txt file
Before using this script, you can set some variables at the top of pdf2txt.py:
-
pdf_file: Pdf file needed to be converted. Save it in the same directory with the script, or pass it on the command line (see Running below).
-
crop_area: Crop location, or
Noneto disable cropping. In some scanned version pdf, you may see watermark, header or footer in the paper. They will generate some noise during OCR. To avoid this, you should set this variable to crop the margin. The crop location is in the form of (xmin, ymin, xmax, ymax) in pixels at the configured dpi, which is left top x axis, left top y axis, right down x axis, right down y axis. -
method: When rebuilding paragraphs, we need to confirm if two adjacent lines belong to the same paragraph. If the paragraph in the pdf file has indentation, you should choose the 'space' method; or else you should choose the 'interval' method, which means the script will use intervals between two lines to confirm.
-
dpi: The resolution used to render scanned pages before OCR (default 200).
-
ocr_lang: The Tesseract language packs to use (default 'chi_sim+eng').
-
max_workers: Number of worker processes (default: all CPU cores).
Note: the platform variable of the old script is gone — the file is written in text mode, so the line ending automatically matches the operating system the script runs on.
As shown in the picture below, before ocr you need to crop out the area outside the red line. The area within the red line is actually what we need. Using software to get the x/y axis value, you will get the value of variable crop_area.
After finish setting the variables, you can simply type python pdf2txt.py in the command line to run this script, or pass the pdf file directly:
python pdf2txt.py your.pdfPDFs with an embedded text layer are converted almost instantly. Scanned PDFs are OCRed locally in parallel, so the conversion speed scales with the number of CPU cores instead of being capped at 1 page per second by a remote API. When the script is done, you could find the txt file in the current directory with the prefix of the pdf file.
前几天(2018.11.11之前)严锋老师在微博上询问有没有什么软件可以完成扫描版PDF OCR转录成TXT的免费软件。从网友的回复来看,似乎仍然没有一款软件能满足上述的全部需求(免费且好用)。所以就有了这么一个小项目。
本脚本最初依赖腾讯AI开放平台的通用OCR API(ai.qq.com),该平台现已停止服务。脚本已按照当前PDF转文本的最佳实践重构:
- 文本层快速路径:如果页面含有内嵌文本层,直接使用PyMuPDF提取文本,完全无需OCR,速度提升数个量级;
- 本地OCR:无文本层的扫描页在内存中渲染后使用本地Tesseract识别,没有网络往返和API调用频率限制(旧版每页强制sleep一秒),也不再向磁盘写入中间图片/json文件;
- 并行处理:所有页面通过进程池并行处理,充分利用多核CPU。
首先安装Tesseract OCR引擎(仅处理无文本层的扫描版PDF时需要):
- Linux:
sudo apt install tesseract-ocr tesseract-ocr-chi-sim - macOS:
brew install tesseract tesseract-lang - Windows:从UB-Mannheim/tesseract下载安装包,并勾选中文简体语言包。
然后安装Python依赖:
pip install -r requirements.txt并行处理每一页:优先提取内嵌文本层,否则在内存中渲染页面并进行本地OCR -> 根据文本行位置信息重组段落 -> 跨页拼接得到最终的TXT文件
在使用本脚本之前,你可以在pdf2txt.py顶部设置以下参数:
-
pdf_file: 你需要识别转换的PDF文件(需和当前脚本保持同一目录),也可以通过命令行参数传入(见下方“运行”)。
-
crop_area: 需要剪裁的区域大小,设置为
None表示不剪裁。由于某些扫描版本的pdf中存在额外添加的水印,为消除此水印以及页眉页脚带来的噪音干扰,请设定剪裁区域。剪裁区域的坐标为dpi分辨率下的像素值,分别为(左上角x轴数值,左上角y轴数值,右下角x轴数值,右下角y轴数值)。你可以使用shutter或者画图等软件进行像素值的确认。 -
method: 重组段落时,需要确认识别结果中紧邻两行是否属于同一个段落。如果PDF文件中的文字段落有缩进,请选择'space';否则请选择'interval',即通过行间距判断文字的归属。
-
dpi: 扫描页渲染为图像时的分辨率(默认200)。
-
ocr_lang: Tesseract OCR使用的语言包(默认'chi_sim+eng')。
-
max_workers: 并行处理的进程数(默认使用全部CPU核心)。
注意:旧版脚本中的platform参数已经移除——文本以文本模式写入,换行符会自动适配脚本运行的操作系统。
如下图所示,你需要将正文的内容从PDF扫描页中剪裁出来;红框中的区域才是我们所需要的。通过相关软件确认了左上角和右下角的坐标点后即可得到crop_area参数。
在完成参数的设置后,你可以通过在命令行中输入python pdf2txt.py运行脚本,或者直接传入PDF文件名:
python pdf2txt.py your.pdf含有文本层的PDF几乎瞬间即可完成转换;扫描版PDF会在本地并行OCR,转换速度随CPU核心数扩展,而不再受远程API每秒一页的限制。当脚本最终运行结束后,你会在当前文件夹中发现一个和PDF文件同名的TXT文件。
