Build a chapterized M4B audiobook from a scanned PDF with resumable OCR, text cleanup, chapter detection, and audio synthesis stages.
The preferred entry point is the single-file uv script:
OPENAI_API_KEY=... uv run make_audiobook.py book.pdfuv reads the dependency metadata embedded in make_audiobook.py, downloads the
needed Python packages, and runs the pipeline.
Markdown is still produced, but it is an intermediate artifact: the main output is a resumable audiobook build ending in chapter WAV files and, optionally, an M4B package.
For book.pdf, files are written under output/book/:
page_images/ rendered PDF pages
markdown_pages/ raw OCR Markdown, one file per page
cleaned_pages/ narration-cleaned text, one file per page
chapters/ chapter-aware narration text, one file per chapter
chapter_audio/ per-engine chapter WAV output
audio_chunks/ generated WAV chunks
book.md merged raw OCR Markdown
cleaned_book.md merged narration-cleaned Markdown
audiobook_text.md chapter-aware text used for speech synthesis
chapters.json detected chapters, front/back bounds, and section headings
audiobook.wav final concatenated audiobook
audiobook.m4b optional chapterized audiobook package
config.json effective run configuration
cost_report.json estimated text/vision API cost from logged usage
run_manifest.jsonl page/chunk processing log
Existing files are reused by default, so reruns resume where they left off.
Use --overwrite to regenerate existing page text and audio chunks.
Chapter detection is based on the raw OCR pages, then validated against page
order and chapter sequence before headings are inserted into the final
audiobook_text.md. The detector handles literal Chapter N headings and can
also derive chapters from an early contents page with numbered entries such as
1 Title, matching those titles back to opener-shaped pages. This keeps
narration chapter markers even when the cleanup stage removes or normalizes
them. It also looks for early narratable front matter such as a foreword,
preface, introduction, prologue, or author note, and drops cover, copyright,
contents, and other pages before that point from the audiobook text. When an
Index section is detected after the book body, narration stops before it so
index pages and trailing ads are not read aloud. The cleaned page files are also
repaired from the validated chapter list so chapter titles remain visible during
audits. If no chapters are detected, the whole narration is synthesized as one
book-level WAV and the M4B is packaged without chapter markers.
The cost report uses actual token usage returned by text/vision API calls and a small rate table sourced from OpenAI's published pricing page. It separates the current run from all logged work, and breaks text/vision cost down by OCR and cleanup stages. Exact billed spend is still reconciled through the OpenAI dashboard or organization Costs endpoint, especially for audio generation.
run_manifest.jsonl is append-only and records major pipeline events as they
happen: run start/completion, stage start/completion/skips, rendered pages, OCR
pages, cleaned pages, chapter detection, chapter text creation, TTS chunks, WAV
concatenation, M4B packaging, and cost report creation.
uv run make_audiobook.py book.pdf --text-only
uv run make_audiobook.py book.pdf --voice nova
uv run make_audiobook.py book.pdf --raw-ocr
uv run make_audiobook.py book.pdf --image-size 1400
uv run make_audiobook.py book.pdf --ocr-model gpt-5.4-mini
uv run make_audiobook.py book.pdf --text-only --refresh clean
uv run make_audiobook.py book.pdf --audio-engine kokoro --chapters 1
uv run make_audiobook.py book.pdf --audio-engine kokoro --m4bThe CLI intentionally exposes only options that change produced files, model choices, voice, image quality/cost, output location, or resume behavior.
List known Kokoro voice IDs:
uv run make_audiobook.py --list-voicesGenerate a short voice sample without processing a PDF:
uv run make_audiobook.py \
--voice af_nicole \
--sample-text "This is a short narration sample for choosing an audiobook voice."Samples are written to voice_samples/ by default. Use --sample-output to
choose a specific WAV path or output directory.
Generate a separate WAV file for each detected chapter:
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice am_adamFiles are written to output/book/chapter_audio/:
chapter_001.wav
chapter_002.wav
chapter_003.wav
Kokoro uses the first character of the voice ID as the language by default:
a for American English, b for British English, p for Brazilian Portuguese,
and so on. You can override that when needed:
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice bf_emma
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice af_nicole --kokoro-speed 0.95
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice pm_alex --kokoro-language p
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice af_heart --refresh audioGenerate only chapter 1:
uv run make_audiobook.py book.pdf --audio-engine kokoro --voice am_adam --chapters 1Chapter announcements are synthesized separately from the chapter body and get
real silence around them by default: 0.4 seconds before the announcement and
1.0 second after it. Tune that spacing when regenerating audio:
uv run make_audiobook.py book.pdf \
--audio-engine kokoro \
--voice am_adam \
--chapter-announcement-lead-silence 0.6 \
--chapter-announcement-trail-silence 1.4 \
--refresh audio \
--m4bGenerate all chapter WAV files and package them into a chapterized M4B:
uv run make_audiobook.py book.pdf \
--audio-engine kokoro \
--voice am_adam \
--m4bThe M4B is written to output/book/audiobook.m4b. By default, the first
rendered PDF page is attached as the cover. Override it with --cover cover.jpg,
or disable cover art with --no-cover.
When chapter text contains Markdown section headings, M4B packaging adds those
headings as additional navigation markers inside the chapter. These section
marker timestamps are estimated from text position within the generated WAV; for
exact subchapter timings, regenerate audio as smaller section-level files before
packaging. The detected headings are also written to chapters.json under
section_headings so you can audit what will become M4B navigation.
Intro WAV files can be prepended before chapter 1:
uv run make_audiobook.py book.pdf \
--audio-engine kokoro \
--m4b \
--intro-wav intro.wav \
--intro-wav dedication.wav \
--cover cover.jpgIf chapter WAV files already exist, reruns reuse them. When --m4b is present,
the M4B package is rebuilt every time so changes to intro files, cover art,
bitrate, or chapter metadata are picked up without an extra refresh flag.
mlx-chatterbox uses mlx-audio with the mlx-community/chatterbox-fp16
model. It can synthesize speech with the model's shipped conditionals, so a
reference clip is not required.
uv run make_audiobook.py book.pdf \
--audio-engine mlx-chatterbox \
--chapters 1 \
--mlx-speed 0.88 \
--mlx-chunk-chars 350 \
--mlx-max-tokens 4000 \
--refresh audioFor voice cloning, pass a short reference WAV and its transcript with
--mlx-ref-audio and --mlx-ref-text. Keep the reference clip short and
transcript-matched; using a full chapter as the reference can produce unusable
audio.
Chatterbox output is written to output/book/chapter_audio/mlx-chatterbox/.
Long chapters are split into smaller MLX-Audio segments and stitched into one
chapter WAV to avoid single-call truncation. --mlx-speed post-processes the
chapter WAV after generation; values below 1.0 slow speech while preserving
pitch. Use --refresh audio when changing speed for an already-generated
chapter. This backend is much slower than Kokoro in local testing, so start with
a short sample before regenerating a full chapter.