HemantKumar01/UGC_Overflow

★ 1Forks 0PythonGitHub ↗Compare

Project website ↗

README

UGC Army

Turn any URL into viral AI influencer videos — automatically.

UGC Army is an end-to-end pipeline that scrapes a website, generates viral marketing scripts, creates realistic AI avatar videos, overlays product screenshots with animated captions, and auto-posts to social media. One URL in, ready-to-post UGC out.

Demo

Paste a product URL → pick an AI avatar → get a polished 9:16 UGC video in minutes.

How It Works

URL → Scrape + Screenshot → AI Script → Avatar Video → Captions → Overlay → Post
Step What happens Powered by
1. Scrape Extracts markdown content from the target website Firecrawl API
2. Screenshot Captures the site at multiple scroll positions Playwright
3. Script Generates 3 viral UGC scripts in different styles (mind-blown, discovery, storytime) GPT-4o-mini + vision
4. Avatar Video Creates a talking-head video with an AI avatar reading the script HeyGen API v3
5. Transcribe Gets word-level timestamps from the audio ElevenLabs Scribe v2
6. Captions Adds karaoke-style animated captions PupCaps + FFmpeg
7. Overlay AI decides when to show which screenshot, composites them onto the video GPT-4o-mini + FFmpeg
8. Post Auto-publishes to Twitter/X via browser automation browser-use + GPT-4o

Architecture

┌────────────────────────────────────────────────────┐
│                  Next.js Frontend                   │
│                                                    │
│  Landing Page (/)      App (/app)                  │
│  ┌──────────────┐     ┌─────────────────────────┐  │
│  │  Waitlist     │     │ URL → Avatar → Generate │  │
│  │  (Firebase)   │     │ → Preview → Post        │  │
│  └──────────────┘     └─────────────────────────┘  │
└──────────────────────────┬─────────────────────────┘
                           │ SSE (real-time progress)
                           ▼
┌────────────────────────────────────────────────────┐
│               Python Pipeline                       │
│                                                    │
│  scraper.py → script_generator.py → video_creator  │
│  → transcriber.py → video_composer.py → poster     │
└────────────────────────────────────────────────────┘

Tech Stack

Frontend

  • Next.js 16 (App Router)
  • Tailwind CSS 4
  • Firebase Firestore (waitlist)

Pipeline

  • Python 3.11+
  • OpenAI GPT-4o-mini (script generation + overlay planning)
  • HeyGen API v3 (AI avatar video)
  • ElevenLabs Scribe v2 (speech-to-text with word timestamps)
  • Firecrawl (web scraping)
  • Playwright (screenshots)
  • PupCaps (animated captions)
  • FFmpeg (video composition)
  • browser-use (autonomous social media posting)

Getting Started

Prerequisites

  • Node.js 20+
  • Python 3.11+
  • FFmpeg installed
  • PupCaps installed

1. Clone & Install

git clone https://github.com/HemantKumar01/hermes-buildathon.git
cd hermes-buildathon

# Frontend
npm install

# Pipeline
cd pipeline
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
playwright install chromium

2. Configure Environment

# Frontend — create .env.local with Firebase config
cp .env.local.example .env.local
# Fill in your Firebase project credentials

# Pipeline
cp pipeline/.env.example pipeline/.env
# Fill in API keys

Required API keys:

Key Service Purpose
NEXT_PUBLIC_FIREBASE_* Firebase Waitlist storage
HEYGEN_API_KEY HeyGen Avatar video generation
OPENAI_API_KEY OpenAI Script generation
FIRECRAWL_API_KEY Firecrawl Web scraping
ELEVENLABS_API_KEY ElevenLabs Speech-to-text

3. Run

# Start the web app
npm run dev

# Or run the pipeline directly
cd pipeline
python3 main.py https://your-product.com

# With auto-posting to Twitter
python3 main.py https://your-product.com --post

# Local mode (reuse existing output, skip API calls)
python3 main.py --local

Project Structure

├── src/
│   ├── app/
│   │   ├── page.tsx              # Landing page (waitlist)
│   │   ├── app/page.tsx          # Product app (generation flow)
│   │   └── api/
│   │       ├── generate/route.ts # Streams pipeline progress via SSE
│   │       ├── video/[jobId]/    # Serves generated videos
│   │       └── youtube/route.ts  # YouTube upload endpoint
│   ├── config/
│   │   ├── firebase.ts           # Firebase client config
│   │   └── avatars.ts            # Avatar definitions
│   └── types/index.ts
├── pipeline/
│   ├── main.py                   # Pipeline orchestrator
│   ├── scraper.py                # Firecrawl + Playwright
│   ├── script_generator.py       # GPT-4o script + overlay planning
│   ├── video_creator.py          # HeyGen API v3
│   ├── transcriber.py            # ElevenLabs STT
│   ├── video_composer.py         # FFmpeg composition
│   ├── twitter_poster.py         # browser-use automation
│   └── hermes_skill.py           # Hermes Agent integration
└── .env.local.example

What Makes This Different

  • Fully automated — no manual editing, templating, or scheduling
  • Vision-aware scripts — GPT-4o sees the actual website screenshots and references specific UI elements, making scripts feel authentic rather than generic
  • AI-planned overlays — the model decides when to show which screenshot based on what the avatar is saying at that moment
  • Real-time progress — SSE streaming shows each pipeline step as it happens
  • One-click posting — browser-use autonomously navigates Twitter and posts the finished video

Team

Built for the Hermes Buildathon 2025.

License

MIT

Contributors

HemantKumar01

Issues