Playwright-based scraper for Pickup Music Learning Pathway courses.
The project does three things:
- logs into Pickup Music and saves an authenticated browser session,
- scrapes the course structure and lesson content,
- downloads lesson and exercise videos from the real player media URLs.
It is designed for courses/playlists for Pickup Music learning pathways:
- Class overview
- Grade 1, Grade 2, ...
- Day 1, Day 2, ...
- lesson tabs such as Lesson, Exercise 1, Exercise 2, Jam
- video delivered through the in-page player, typically Mux/HLS (
.m3u8)
For each lesson, the scraper writes files into:
output/lessons/<grade-slug>/<lesson-slug>/
Typical outputs:
lesson.md— markdown version of the lessonlesson.html— HTML export of the lesson*.pdf— downloaded notation PDF, when available*.png— notation / fretboard screenshots, when capturedvideo-lesson.mp4— main lesson videovideo-exercise-1.mp4,video-exercise-2.mp4, ...video-jam.mp4output/course.json— the scraped course structure used later by the video downloaderoutput/storage.json— authenticated Playwright session
- Node.js 20+
- ffmpeg available in
PATH - ffprobe available in
PATH - Chromium installed through Playwright
Install dependencies:
npm install
npx playwright install chromiumCreate a .env file in the project root.
Example:
BASE_URL=https://my.pickupmusic.com
PICKUP_EMAIL=[email protected]
PICKUP_PASSWORD=your_password
COURSE_TITLE=X Learning Pathway
COURSE_PATH=/guitar/class/<class-id>/grade/<grade-id>
CONCURRENCY=1BASE_URL— Pickup Music base URLPICKUP_EMAIL— account emailPICKUP_PASSWORD— account passwordCOURSE_PATH— path to the target course / learning pathway
COURSE_TITLE— used inoutput/course.jsonSCRAPE_GRADE— scrape only a specific grade, for exampleGrade 5SCRAPE_DAY— scrape only a specific day, for exampleDay 3CONCURRENCY— currently safe to keep at1
COURSE_PATH can point either to:
- a class URL, or
- a grade URL inside the class
The scraper still works as long as the page exposes the standard Class overview sidebar and the course uses the normal learning pathway structure.
Logs into Pickup Music and saves browser state to output/storage.json.
npm run authnpm run scrapeThis creates or updates:
output/course.json- lesson markdown / HTML exports
- notation PDFs and screenshots
npm run videosThis requires output/course.json from the scrape step.
npm run fullThis runs:
- authentication,
- scraping,
- video download.
npm run auth
npm run scrape
npm run videosUsually you only need:
npm run scrape
npm run videosThe scraper writes output/course.json incrementally, and the video downloader is built to skip files that already exist and look valid.
If this happens, consider manually helping the scraper once it's done with the grade, open the specific grade tab and from then it should automatically scrape all the content from the grade. If it doesn't try to click on day 1 or specify SCRAPE_DAY/SCRAPE_GRADE in .env
Run the scrape step first:
npm run scrapeYour login session does not exist yet. Run:
npm run authThe downloader opened the tab but did not capture a valid media URL.
Typical causes:
- the player did not load in time,
- the session expired,
- the page structure changed,
- that lesson uses a slightly different player flow.
Start by re-authenticating:
npm run authThen retry the failing lesson or grade.
That means the media URL filter is too permissive or the player media URL was not captured. The downloader should only pass real video candidates such as Mux .m3u8 / .mp4 URLs into ffmpeg.
Use ffprobe to validate durations. Broken files can be re-downloaded by deleting them and rerunning npm run videos.
Example PowerShell check:
Get-ChildItem -Recurse -Filter *.mp4 | ForEach-Object {
$d = ffprobe -v error -show_entries format=duration -of csv=p=0 "$($_.FullName)"
if ($d) {
[timespan]::FromSeconds([double]$d).ToString("hh\:mm\:ss") + " " + $_.FullName
} else {
"INVALID " + $_.FullName
}
}output/course.json is written incrementally, so partial progress is often preserved. Rerun the scraper or continue with grade/day filters.
In most cases, no code changes are needed. Just change the .env link, title and get rid of previous output files.