SeanPedersen/keep-my-social-data

Archive your favorite content from X, Hacker News, Reddit, and GitHub

★ 1Forks 0PythonGitHub ↗Compare
scrapersocial-media-automation

README

Keep My Social Data

Self-hosted browser scraper that archives liked/starred content from X, Hacker News, Reddit, and GitHub.

Features

  • Multi-platform support: X, Hacker News, Reddit (via browser automation), GitHub (via CLI)
  • Web dashboard: FastAPI-based UI for managing accounts and browsing archived items
  • File mirroring: Items saved as markdown files with YAML frontmatter
  • Scheduled scraping: Automatic periodic sync with configurable intervals
  • Encrypted credentials: Passwords encrypted with Fernet before storage

Requirements

  • Python 3.11+
  • uv package manager
  • GitHub CLI (gh) authenticated for GitHub stars

Installation

# Clone and install
git clone <repo-url>
cd keep-my-social-data
uv sync

# Configure data directory
cp .env.example .env
# Edit .env to set your preferred DATA_DIR

Usage

# Start the dev server
uv run kmsd-dev

# Open http://localhost:8001 in your browser

Adding Accounts

  • GitHub: Run gh auth login first, then add via the web UI
  • X/HN/Reddit: Add username and password through the web UI

Data Storage

  • Database: SQLite at $DATA_DIR/kmsd.db
  • Markdown mirror: $DATA_DIR/{platform}/*.md
  • Secrets: $DATA_DIR/secrets.env (auto-generated)

Development

# Run tests
uv run pytest

# Run specific test file
uv run pytest tests/test_github.py

Configuration

Setting Description Default
DATA_DIR Data storage directory ./data
KMSD_BASE_URL Public URL for the app http://localhost:8000
KMSD_SCRAPE_INTERVAL_MINUTES Sync interval 60

License

MIT

Contributors

SeanPedersen

Issues