Meetpatel006/cvip

★ 0Forks 0PythonGitHub ↗Compare

README

SAM3 Video Background Remover

A full-stack application that performs high-quality, text-prompted instance segmentation and background removal using Meta's SAM3 (Segment Anything Model 3) running on Modal GPUs.

The project contains two main parts:

  1. Frontend: A Next.js web application providing an interactive UI for async video processing and near-live webcam background removal.
  2. Backend: A FastAPI inference service deployed on Modal for heavy GPU processing (handling both asynchronous video files and near-live video frames).

1. Frontend (Next.js application)

The frontend is a modern React web interface built with Next.js, Tailwind CSS, and standard Web APIs.

Features

  • Video Processing Studio (/): Upload a short .mp4 / .mov clip, provide a text prompt (e.g., "person" or "car"), and choose a background strategy (transparent .webm, flat solid color, or an image replacement).
  • Live Camera Feed (/live): Start your webcam and see background removal in near real-time. The frontend smartly decouples frame fetching from rendering: it continuously streams frames to the backend, decodes the SAM3 RLE masks using a custom Javascript decoder, and mathematically masks the webcam on an HTML <canvas> at a smooth ~30/60 FPS.

Setup and Running Locally

cd frontend
# Install dependencies
bun install
# Run the development server
bun run dev

Open http://localhost:3000 to view the tool.


2. Backend (FastAPI + Modal)

The backend handles the heavy lifting, loading the facebook/sam3 model onto a Modal L4 GPU container.

Architecture

  • backend/main.py: The FastAPI ASGI application managing routing, handling file uploads, and tracking async jobs in a modal dictionary.
  • backend/modal_app.py: The Modal App definition. It handles environment setups, volume mounting for large model weights, installing ffmpeg, and passing Huggingface tokens.
  • backend/worker.py & tasks/: The background async workers that actually process multi-frame videos, run the SAM3 inference engine (segment_engine.py), apply bounding boxes via NMS, and handle video extraction/re-encoding with ffmpeg (video.py).

Endpoints

  • POST /process: Upload a video (file) + prompt -> returns a job_id. Processing takes approximately 30-40 seconds per short video.
  • GET /status/{job_id}: Poll background job status until it transitions to completed.
  • GET /download/{job_id}: Download the resulting transparent .webm or composited .mp4.
  • POST /live/frame: High-speed API for single-frame inference. Accepts a frame and returns detected bounding boxes alongside compressed RLE (Run-Length Encoded) masks.

Setup and Deploying

Prerequisites:

  • A Modal account (run modal setup)
  • A Hugging Face account with access granted to the facebook/sam3 repository.
  • A valid HF_TOKEN must be configured in your environment or Modal secrets.

Deploy to Modal:

cd backend
pip install -r requirements.txt
modal deploy modal_app.py

Note: Modal will automatically provision the endpoints and print a public .modal.run URL. You must ensure NEXT_PUBLIC_BG_REMOVER_API or BG_REMOVER_API in your frontend environment variables points to this endpoint.

Contributors

Meetpatel006

Issues