A web-based application to transcribe conversations between Jake and Edmund with AI-powered Q&A capabilities.
- ๐๏ธ Audio Upload: Drag-and-drop M4A audio files with progress tracking
- ๐ฃ๏ธ Speaker Diarization: Automatic speaker identification using AssemblyAI
- ๐ Token Tracking: Real-time token counting for transcripts
- ๐ Real-time Status Polling: Auto-updates during transcription processing
- ๐ Session Management: View and manage all transcription sessions
- ๐ฌ Conversation View: Chat-style transcript display with timestamps (Phase 4)
- ๐ Speaker Swap: Swap Jake/Edmund assignment if diarization is incorrect (Phase 4)
- ๐ค AI Q&A: Ask questions about conversations using GPT-4o (Phase 5)
- Next.js 16 with App Router
- React 19
- TypeScript
- Tailwind CSS v4 for styling
- Zustand for state management
- react-dropzone for file uploads
- Node.js with Express
- TypeScript
- SQLite with Prisma ORM
- AssemblyAI SDK (v4.19) for transcription with speaker diarization
- tiktoken for GPT-4o token counting
- Multer for file uploads
- pnpm 10+ (required)
- Node.js 18+
- pnpm 10+
- AssemblyAI API Key (Get one here)
- OpenAI API Key (for Phase 5 Q&A feature)
git clone <repository-url>
cd sparkv0Frontend:
cd frontend
pnpm installBackend:
cd backend
pnpm install
# Generate Prisma Client
pnpm prisma generateCreate a .env file in the root directory:
# Copy from .env.example if available, or create manually
touch .envAdd your API keys to .env:
# Backend API
PORT=3002
# Database
DATABASE_URL="file:./prisma/database.sqlite"
# AI Services
ASSEMBLYAI_API_KEY=your_assemblyai_api_key_here
OPENAI_API_KEY=your_openai_api_key_here # For Phase 5Initialize the Prisma database:
cd backend
pnpm prisma db pushThis creates the SQLite database with the following tables:
sessions- Audio session metadata (status, duration, token count)transcripts- Transcription segments with speaker infoquestions- Q&A history (Phase 5)
cd frontend
pnpm devThe frontend will be available at http://localhost:3001 (or 3000 if available)
cd backend
pnpm devThe backend API will run at http://localhost:3002
- Frontend: Make changes in
frontend/app,frontend/components, etc. - Backend: Make changes in
backend/src - Both servers support hot-reloading
sparkv0/
โโโ frontend/ # Next.js application
โ โโโ app/
โ โ โโโ layout.tsx # Root layout
โ โ โโโ page.tsx # Home page
โ โ โโโ sessions/ # Session pages
โ โโโ components/ # React components
โ โโโ lib/ # Utilities
โ โโโ store/ # Zustand stores
โ โโโ package.json
โ
โโโ backend/ # Express API
โ โโโ src/
โ โ โโโ index.ts # Main server file
โ โ โโโ routes/ # API routes
โ โ โโโ services/ # Business logic
โ โ โ โโโ database.ts
โ โ โ โโโ assemblyai.ts
โ โ โ โโโ openai.ts
โ โ โ โโโ tokens.ts
โ โ โโโ uploads/ # Audio file storage
โ โโโ database.sqlite # SQLite database (auto-created)
โ โโโ package.json
โ
โโโ experiments/ # Experimental code
โโโ PRD.md # Product Requirements Document
โโโ README.md # This file
POST /api/sessions- Upload M4A audio and create sessionGET /api/sessions- List all sessions (ordered by date)GET /api/sessions/:id- Get session details with transcriptGET /api/sessions/:id/status- Get real-time transcription status (polls AssemblyAI)PATCH /api/sessions/:id/speakers- Swap Jake/Edmund assignment (Phase 4)
POST /api/sessions/:id/questions- Ask question about sessionGET /api/sessions/:id/questions- Get Q&A history
GET /api/health- Server status
cd frontend
pnpm build
pnpm startcd backend
pnpm build
pnpm startIf you see Prisma import errors:
cd backend
pnpm prisma generateAfter modifying backend/prisma/schema.prisma:
cd backend
pnpm prisma db push
pnpm prisma generateIf port 3001 or 3002 is already in use:
Frontend: Next.js will automatically use the next available port
Backend: Change the PORT in backend/.env
Install pnpm globally:
npm install -g pnpmOr use corepack (recommended):
corepack enable
corepack prepare pnpm@latest --activateSee PRD.md for detailed feature specifications and implementation phases.
- Initialize Next.js 16 frontend with TypeScript
- Initialize Express backend with TypeScript
- Set up SQLite with Prisma ORM
- Configure project structure
- Implement drag-and-drop component with react-dropzone
- Add M4A file validation
- Create upload endpoint with progress tracking
- Build session list view with Zustand store
- Integrate AssemblyAI SDK with speaker diarization
- Implement async transcription workflow
- Add real-time status polling
- Parse and store utterances in database
- Implement token counting with tiktoken
- Handle transcription errors with retry logic
- Build chat-style transcript UI
- Add timestamp display for each segment
- Implement speaker swap functionality
- Display total token count
- Add navigation from session list
- Integrate OpenAI GPT-4o API
- Build Q&A chat interface
- Implement context management
- Display answers with timestamp references
- Track Q&A token usage
This is a personal project for Jake and Edmund. See Linear for task tracking and project management.
Private project - All rights reserved