synmux/yt-pdlp

Parallel YouTube (and other sites) downloader

★ 0Forks 0PythonGitHub ↗Compare

README

yt-pdlp

Download a large YouTube playlist with several yt-dlp workers running concurrently, showing each worker's live progress in a Textual terminal UI. It supports a planning dry-run and a post-run flush report that reconciles how many videos actually landed versus failed.

Why

yt-dlp has no built-in option to download multiple playlist items in parallel — its -N / --concurrent-fragments flag only parallelises fragments within a single video, so playlist items are processed one at a time. Downloading a ~1,000-item playlist (such as Watch Later) therefore crawls. yt-pdlp runs several downloads at once in one process, with a live UI and resumability.

Requirements

  • Python ≥ 3.11
  • ffmpeg on your PATH (needed to remux to mp4)
  • deno (or Node.js) on your PATH — yt-dlp runs YouTube's JS challenges through it. On first use yt-dlp fetches its challenge-solver script from GitHub (the ejs:github remote component, which yt-pdlp allows); without a JS runtime, cookie-authenticated downloads fail with "Requested format is not available".
  • A browser you are signed in to YouTube with (default: Chrome) — Watch Later is private, so cookies are required.

Install & run

This project uses uv:

uv sync                       # create the environment
uv run yt-pdlp --help  # see all options

Usage

yt-pdlp is a command group with a default command, so a bare invocation downloads:

# Download Watch Later with 4 workers (the defaults)
uv run yt-pdlp

# Six workers, a specific playlist, into ./videos
uv run yt-pdlp -j 6 -u "https://www.youtube.com/playlist?list=PLxxxx" -o ./videos

# Sources from a file — playlists, channels and single videos, one per line
uv run yt-pdlp -a sources.txt -o ./videos

# Plan only — contacts YouTube read-only, downloads nothing
uv run yt-pdlp --dry-run

# Reconcile an existing output directory without downloading
uv run yt-pdlp flush -o ./videos

download options (the default command)

Option Short Default Meaning
--jobs -j 4 Number of concurrent workers.
--url -u Watch Later Playlist, channel or video URL; combines with --batch-file.
--batch-file -a (none) File of source URLs, one per line — each a playlist, channel or video. Blank lines and #/;/] comment lines are skipped. Watch Later is used only when neither --url nor --batch-file is given.
--output -o ./downloads Output directory (created if absent).
--browser -b chrome Browser to read cookies from.
--format -f None Remux container; an empty string or None disables remux.
--fragments -N 8 concurrent_fragment_downloads per worker (intra-video).
--dry-run off Plan only; download nothing.
--plain auto Disable the Textual UI; emit line-based progress (auto-on when stdout is not a terminal).

flush options

Option Short Default Meaning
--output -o ./downloads Locate and reconcile this output's state directory.
--url -u (optional) Re-flatten this playlist to define the "requested" set instead of using the cached list.

How it works

  • A shared work queue holds every playlist entry. --jobs workers each pull the next entry and download it, so the work load-balances naturally across videos of very different lengths.
  • Each worker owns its own yt_dlp.YoutubeDL instance and reports progress through a UI-agnostic event stream. A Textual front-end renders per-worker panels and an overall counter; a plain front-end prints line-based progress when there is no terminal.
  • A source that fails to flatten — a deleted or private video, a dead playlist — is warned about and skipped; the run aborts only when every source fails.
  • A shared download archive records only successful downloads, so re-running the same command skips what is already done and retries what failed.
  • The browser cookie store is read once and written to a reusable cookie file; every worker then reads the file (browsers lock their cookie database, so reading it from many workers at once would contend).

State directory

Everything lives under <output>/.ytdlp-state/:

File Purpose
cookies.txt Cookies exported once from the browser. Sensitive — it holds a live session; keep it private.
entries.json The flattened playlist: a list of {id, url, title}.
archive.txt The yt-dlp download archive (resume + skip).
failed.txt Outstanding URLs after a run or flush, one per line — retry with yt-dlp -a failed.txt … or just re-run.
report.txt The last completion report.

Downloaded media files go directly under <output>/.

A note on concurrency

The realistic sweet spot is 4–8 workers. The bottleneck is YouTube's per-account/IP throttling (HTTP 429), not your machine — beyond a handful of workers, total throughput usually drops. yt-pdlp warns when you ask for more than 8 but does not stop you.

Development

uv run pytest -q                       # tests
uv run ruff check src tests            # lint
uv run ruff format --check src tests   # formatting
uv run ty check src                    # type-check

The download engine is deliberately decoupled from the UI, and yt-dlp, the network, and the terminal are faked in tests only so the whole tool runs offline in CI.

Contributors

synmuxdependabot[bot]

Issues