Prompt → LLM-written Manim script → sandboxed render → video. Multi-scene, with per-scene refinement ("make the rectangle blue") that re-renders only that scene.
The Hono API validates a prompt, writes a Project row, enqueues a BullMQ job, and returns 202. A Bun worker plans scenes with Claude, generates Manim code per scene, renders each scene inside a locked-down Docker container (--network=none, CPU/memory caps, timeout), feeds render stderr back to the LLM for up to MAX_FIX_ATTEMPTS self-corrections, then concatenates clips with FFmpeg. The Next.js app polls status and shows per-scene previews with a refine box. LLM-generated code is untrusted — it only ever executes inside the sandbox container.
- Bun 1.2+
- Docker (for Postgres, Redis, and the Manim sandbox)
- FFmpeg + ffprobe on the host (used for concat + duration probing)
bun install
cp .env.example .env # add your ANTHROPIC_API_KEY
docker compose up -d # postgres + redis
docker pull manimcommunity/manim:v0.19.0 # pre-pull the sandbox image
bun run db:migrate # answer the prisma prompts (name: init)Milestone first — prove the pipeline with no servers:
bun run pipeline:test "a blue rectangle appears, then morphs into a circle"
# → prints the path to an MP4 when it worksThen run the whole app:
bun run dev # web :3000 · api :4000 · workerapps/web— Next.js UI: prompt in, scene timeline + previews, refine per sceneapps/api— Hono: validate → persist → enqueue → 202; serves clips/videosapps/worker— BullMQ consumers: generate-video + scene-refine pipelinespackages/pipeline— codegen (plan/generate/fix/refine prompts), Docker sandbox runner, FFmpeg concatpackages/database— Prisma schema + client singletonpackages/queue— queue names, connection, typed job payloadspackages/shared— API DTO types shared by web + api
- Single-host v1: api + worker share
storage/on one machine. The worker shells out todocker, so in production either mount the Docker socket or swapsandbox.service.tsfor Modal/E2B. - No auth yet (deliberate, to reach the magic moment fast). Add Hono JWT middleware + a
Usermodel when you need it. - Later: xfade transitions, voiceover via
manim-voiceover(extendsandbox/Dockerfile), scene reordering, WebSocket progress instead of polling.