jakehemmerle/sparkv0

โ˜… 0Forks 0TypeScriptGitHub โ†—Compare

README

Spark - Conversation Transcription & Analysis

A web-based application to transcribe conversations between Jake and Edmund with AI-powered Q&A capabilities.

Features

โœ… Implemented

  • ๐ŸŽ™๏ธ Audio Upload: Drag-and-drop M4A audio files with progress tracking
  • ๐Ÿ—ฃ๏ธ Speaker Diarization: Automatic speaker identification using AssemblyAI
  • ๐Ÿ“Š Token Tracking: Real-time token counting for transcripts
  • ๐Ÿ”„ Real-time Status Polling: Auto-updates during transcription processing
  • ๐Ÿ“‹ Session Management: View and manage all transcription sessions

๐Ÿšง Planned

  • ๐Ÿ’ฌ Conversation View: Chat-style transcript display with timestamps (Phase 4)
  • ๐Ÿ”„ Speaker Swap: Swap Jake/Edmund assignment if diarization is incorrect (Phase 4)
  • ๐Ÿค– AI Q&A: Ask questions about conversations using GPT-4o (Phase 5)

Tech Stack

Frontend

  • Next.js 16 with App Router
  • React 19
  • TypeScript
  • Tailwind CSS v4 for styling
  • Zustand for state management
  • react-dropzone for file uploads

Backend

  • Node.js with Express
  • TypeScript
  • SQLite with Prisma ORM
  • AssemblyAI SDK (v4.19) for transcription with speaker diarization
  • tiktoken for GPT-4o token counting
  • Multer for file uploads

Package Manager

  • pnpm 10+ (required)

Prerequisites

  • Node.js 18+
  • pnpm 10+
  • AssemblyAI API Key (Get one here)
  • OpenAI API Key (for Phase 5 Q&A feature)

Setup Instructions

1. Clone the Repository

git clone <repository-url>
cd sparkv0

2. Install Dependencies

Frontend:

cd frontend
pnpm install

Backend:

cd backend
pnpm install

# Generate Prisma Client
pnpm prisma generate

3. Environment Configuration

Create a .env file in the root directory:

# Copy from .env.example if available, or create manually
touch .env

Add your API keys to .env:

# Backend API
PORT=3002

# Database
DATABASE_URL="file:./prisma/database.sqlite"

# AI Services
ASSEMBLYAI_API_KEY=your_assemblyai_api_key_here
OPENAI_API_KEY=your_openai_api_key_here  # For Phase 5

4. Database Setup

Initialize the Prisma database:

cd backend
pnpm prisma db push

This creates the SQLite database with the following tables:

  • sessions - Audio session metadata (status, duration, token count)
  • transcripts - Transcription segments with speaker info
  • questions - Q&A history (Phase 5)

Development

Start Frontend (Next.js)

cd frontend
pnpm dev

The frontend will be available at http://localhost:3001 (or 3000 if available)

Start Backend (Express)

cd backend
pnpm dev

The backend API will run at http://localhost:3002

Development Workflow

  1. Frontend: Make changes in frontend/app, frontend/components, etc.
  2. Backend: Make changes in backend/src
  3. Both servers support hot-reloading

Project Structure

sparkv0/
โ”œโ”€โ”€ frontend/                # Next.js application
โ”‚   โ”œโ”€โ”€ app/
โ”‚   โ”‚   โ”œโ”€โ”€ layout.tsx      # Root layout
โ”‚   โ”‚   โ”œโ”€โ”€ page.tsx        # Home page
โ”‚   โ”‚   โ””โ”€โ”€ sessions/       # Session pages
โ”‚   โ”œโ”€โ”€ components/         # React components
โ”‚   โ”œโ”€โ”€ lib/               # Utilities
โ”‚   โ”œโ”€โ”€ store/             # Zustand stores
โ”‚   โ””โ”€โ”€ package.json
โ”‚
โ”œโ”€โ”€ backend/                # Express API
โ”‚   โ”œโ”€โ”€ src/
โ”‚   โ”‚   โ”œโ”€โ”€ index.ts       # Main server file
โ”‚   โ”‚   โ”œโ”€โ”€ routes/        # API routes
โ”‚   โ”‚   โ”œโ”€โ”€ services/      # Business logic
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ database.ts
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ assemblyai.ts
โ”‚   โ”‚   โ”‚   โ”œโ”€โ”€ openai.ts
โ”‚   โ”‚   โ”‚   โ””โ”€โ”€ tokens.ts
โ”‚   โ”‚   โ””โ”€โ”€ uploads/       # Audio file storage
โ”‚   โ”œโ”€โ”€ database.sqlite    # SQLite database (auto-created)
โ”‚   โ””โ”€โ”€ package.json
โ”‚
โ”œโ”€โ”€ experiments/           # Experimental code
โ”œโ”€โ”€ PRD.md                # Product Requirements Document
โ””โ”€โ”€ README.md             # This file

API Endpoints

Sessions

  • POST /api/sessions - Upload M4A audio and create session
  • GET /api/sessions - List all sessions (ordered by date)
  • GET /api/sessions/:id - Get session details with transcript
  • GET /api/sessions/:id/status - Get real-time transcription status (polls AssemblyAI)
  • PATCH /api/sessions/:id/speakers - Swap Jake/Edmund assignment (Phase 4)

Q&A (Phase 5)

  • POST /api/sessions/:id/questions - Ask question about session
  • GET /api/sessions/:id/questions - Get Q&A history

Health Check

  • GET /api/health - Server status

Building for Production

Frontend

cd frontend
pnpm build
pnpm start

Backend

cd backend
pnpm build
pnpm start

Troubleshooting

Prisma Client Not Generated

If you see Prisma import errors:

cd backend
pnpm prisma generate

Database Schema Changes

After modifying backend/prisma/schema.prisma:

cd backend
pnpm prisma db push
pnpm prisma generate

Port Already in Use

If port 3001 or 3002 is already in use:

Frontend: Next.js will automatically use the next available port Backend: Change the PORT in backend/.env

pnpm Not Found

Install pnpm globally:

npm install -g pnpm

Or use corepack (recommended):

corepack enable
corepack prepare pnpm@latest --activate

Development Roadmap

See PRD.md for detailed feature specifications and implementation phases.

Phase 1: Project Setup โœ… Complete

  • Initialize Next.js 16 frontend with TypeScript
  • Initialize Express backend with TypeScript
  • Set up SQLite with Prisma ORM
  • Configure project structure

Phase 2: Audio Upload UI โœ… Complete

  • Implement drag-and-drop component with react-dropzone
  • Add M4A file validation
  • Create upload endpoint with progress tracking
  • Build session list view with Zustand store

Phase 3: Backend & AssemblyAI Integration โœ… Complete

  • Integrate AssemblyAI SDK with speaker diarization
  • Implement async transcription workflow
  • Add real-time status polling
  • Parse and store utterances in database
  • Implement token counting with tiktoken
  • Handle transcription errors with retry logic

Phase 4: Conversation View ๐Ÿšง Next

  • Build chat-style transcript UI
  • Add timestamp display for each segment
  • Implement speaker swap functionality
  • Display total token count
  • Add navigation from session list

Phase 5: Q&A Feature ๐Ÿšง Planned

  • Integrate OpenAI GPT-4o API
  • Build Q&A chat interface
  • Implement context management
  • Display answers with timestamp references
  • Track Q&A token usage

Contributing

This is a personal project for Jake and Edmund. See Linear for task tracking and project management.

License

Private project - All rights reserved

Contributors

jakehemmerle

Issues