Small, configurable Reddit scraping utility built on the public Reddit API.
- Powered by PRAW for authenticated Reddit access.
- Script/app OAuth flow with environment based credentials (read-only by default).
- Query by free-form search strings and/or target subreddits.
- Optional media-only filter with support for image/video downloads.
- Pluggable cache storage (local filesystem or Google Cloud Storage).
- Ledger of crawled posts recorded to CSV or SQLite.
- Create a Reddit App (script type) at https://www.reddit.com/prefs/apps and grab the client id/secret.
- Configure credentials via environment variables or a
.envfile:
REDDIT_CLIENT_ID=your_client_id
REDDIT_CLIENT_SECRET=your_client_secret
REDDIT_USER_AGENT=social-crawler/0.1 by your_username
# Optional: add these to enable script (read/write) flows; omit for read-only access.
REDDIT_USERNAME=reddit_username
REDDIT_PASSWORD=reddit_password- Install dependencies (Python 3.10+ recommended):
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt- Run the scraper with desired filters:
python -m social_crawler.cli \
--query "openai" \
--subreddit "technology" \
--max-posts 25 \
--media-only \
--download-media \
--storage-backend local \
--storage-path cache \
--ledger-mode csv \
--ledger-path data/ledger.csvThe command above searches r/technology for recent posts mentioning "openai", stores post JSON (and media files if present) under cache/, and writes a ledger row for each post at data/ledger.csv.
Note: If you encounter
RuntimeError: praw is required but not installed, install the dependency withpip install praw(or reinstallrequirements.txt) before running the scraper.
Instead of passing multiple flags, provide a JSON file and point the CLI at it with --config:
{
"queries": {
"queries": ["openai", "chatgpt"],
"subreddits": ["technology", "machinelearning"],
"sort": "top",
"time_filter": "week",
"max_posts": 25,
"media_only": true,
"download_media": true
},
"storage": {
"backend": "local",
"local_path": "cache"
},
"ledger": {
"mode": "csv",
"csv_path": "data/ledger.csv"
}
}Run the scraper with:
python -m social_crawler.cli --config config.json- Local (default): caches JSON and media files to a directory you control.
- Google Cloud Storage: pass
--storage-backend gcs --gcs-bucket your-bucket --gcs-prefix optional/prefix. Requiresgoogle-cloud-storagecredentials set via standard environment variables or application default credentials.
--ledger-mode csv --ledger-path <file>: append-only CSV ledger.--ledger-mode sqlite --ledger-path <db>: upsert intoreddit_poststable (primary keypost_id).
- Reddit rate limiting applies; consider throttling invocations or adding sleeps for large crawls.
- When
--media-onlyis set, only posts with Reddit-hosted video/images or direct media links are kept. - The scraper downloads media files only when
--download-mediais on; otherwise it just records the media URL. - For bulk or scheduled usage, wrap the scraper in cron or a workflow manager and point ledger storage to a centralized location.