A Flask application to track the average rating and rating count of films on Letterboxd over time. It scrapes film pages periodically and stores snapshots of the rating data, which can be viewed on a public dashboard with historical charts.
- Admin Dashboard: Add, remove, and manage films to be tracked.
- Automated Scraping: A background scheduler (standalone APScheduler) periodically fetches new rating data for tracked films.
- Historical Data: Stores rating snapshots to visualize trends over time.
- Public View: A simple public interface to view tracked films and their rating history.
- API Endpoint: Provides JSON data for charts.
- Manual Scraping: Trigger immediate scraping from the admin dashboard.
- Compare Page: Compare two films on a scatter plot (X: rating count, Y: average rating) with autocomplete search.
- Backend: Flask
- Database: SQLAlchemy with SQLite (default) or PostgreSQL.
- Migrations: Flask-Migrate
- Scheduling: Standalone APScheduler (not Flask-APScheduler)
- Deployment: Gunicorn, Nginx, Systemd
-
Clone the repository:
git clone <your-repo-url> cd letterboxd-tracker
-
Create and activate a virtual environment:
python -m venv venv # On Windows .\venv\Scripts\activate # On macOS/Linux source venv/bin/activate
-
Install dependencies:
pip install -r requirements.txt
-
Configure environment variables: Create a
.envfile in the project root with the following content:# Flask Configuration SECRET_KEY=your-secret-key-here-change-this-in-production FLASK_ENV=development # Database Configuration (SQLite by default) DATABASE_URL=sqlite:///instance/letterboxd_tracker.sqlite3 # Admin Dashboard Credentials ADMIN_USERNAME=admin ADMIN_PASSWORD=your-secure-password-here # Scheduler Configuration SCHEDULER_INTERVAL_MINUTES=60 # Optional: Enable scheduler API in main Flask app SCHEDULER_API_ENABLED=False
-
Initialize and migrate the database:
flask db init # Only run this the very first time flask db migrate -m "Initial migration" flask db upgrade
-
Run the application:
flask run
The app will be available at
http://127.0.0.1:5000.
The application now uses a standalone APScheduler script for scraping jobs. This is more robust and works reliably on all platforms.
Open a new terminal and run:
python run_scheduler.pyYou should see output like:
Starting standalone APScheduler for scraping (interval: 15 min)...
Scheduler started! Press Ctrl+C to exit.
============================================================
SCHEDULED SCRAPE JOB FIRED at 2025-06-22 04:30:00
============================================================
The scraping job will run at the interval set in your .env file (SCHEDULER_INTERVAL_MINUTES).
Create a systemd service for the scheduler:
/etc/systemd/system/letterboxd-tracker-scheduler.service
[Unit]
Description=Letterboxd Tracker Standalone Scheduler
After=network.target
[Service]
User=your_user
Group=www-data
WorkingDirectory=/path/to/your/letterboxd-tracker
Environment="PATH=/path/to/your/letterboxd-tracker/venv/bin"
ExecStart=/path/to/your/letterboxd-tracker/venv/bin/python run_scheduler.py
Restart=always
RestartSec=10
[Install]
WantedBy=multi-user.targetEnable and start the service:
sudo systemctl daemon-reload
sudo systemctl enable letterboxd-tracker-scheduler
sudo systemctl start letterboxd-tracker-schedulerCheck logs:
sudo journalctl -u letterboxd-tracker-scheduler -fYour web app is still deployed as before, using Gunicorn and systemd. Example service:
/etc/systemd/system/letterboxd-tracker.service
[Unit]
Description=Gunicorn instance to serve Letterboxd Tracker
After=network.target
[Service]
User=your_user
Group=www-data
WorkingDirectory=/path/to/your/letterboxd-tracker
Environment="PATH=/path/to/your/letterboxd-tracker/venv/bin"
ExecStart=/path/to/your/letterboxd-tracker/venv/bin/gunicorn --workers 3 --bind unix:letterboxd-tracker.sock -m 007 wsgi:app
[Install]
WantedBy=multi-user.target- Scheduler not running:
- Check the systemd service status:
sudo systemctl status letterboxd-tracker-scheduler - Check logs:
sudo journalctl -u letterboxd-tracker-scheduler -f
- Check the systemd service status:
- Web app issues:
- Check the Gunicorn service:
sudo systemctl status letterboxd-tracker - Check logs:
sudo journalctl -u letterboxd-tracker -f
- Check the Gunicorn service:
- Scheduler interval:
- Set
SCHEDULER_INTERVAL_MINUTESin your.envfile.
- Set
- No scraping output:
- Make sure you are running
run_scheduler.pyand not the old Flask-APScheduler setup.
- Make sure you are running
Access the admin dashboard at http://127.0.0.1:5000/admin and log in with your configured credentials.
Features:
- Add new films by Letterboxd URL
- Toggle tracking on/off for individual films
- Manual scraping for immediate updates
- View scheduler status and next run time
- Reorder films in the display list
- Delete films and their associated data
- Access via the navbar:
Compare. - Select two films using the autocomplete inputs (search by title or slug).
- The chart plots both films as scatter series:
- X axis: number of ratings
- Y axis: average rating
- Different rating counts between the two films are fine; each series is plotted independently based on available snapshots.
APIs used by the page:
GET /api/films/search?q=<query>: returns matching films for autocomplete.GET /api/compare?a=<slugA>&b=<slugB>: returns scatter points for both films.
Adding Films:
- Go to the film's page on Letterboxd (e.g.,
https://letterboxd.com/film/28-years-later/) - Copy the URL
- Paste it into the "Add Film" form on the admin dashboard
- The system will automatically extract the film slug and fetch initial data
To test the scheduler manually:
python run_scheduler.pyTo check what films are in the database:
python debug_env.pyUse the provided backup script:
./backup_db.shThis project is open source and available under the MIT License.