A real-time voice-to-topic extraction application that converts speech into hierarchical topic trees using Gemini AI with LLM validation and a beautiful custom visualization interface.
- ๐ค Voice Recording: Real-time voice input with browser MediaRecorder API
- ๐ File Upload: Support for various audio formats (MP3, WAV, M4A, OGG, FLAC, AAC, WebM)
- ๐ฃ๏ธ Speech-to-Text: High-quality transcription using AssemblyAI
- ๐ง AI-Powered Extraction: Uses Google Gemini 2.0-flash to extract hierarchical topics
- โ LLM Validation: Gemini 2.5-flash-preview validates and refines extracted topics
- ๐จ Beautiful Visualization: Custom circular node interface with 23 soft pastel colors
- ๐ฑ Mobile-Friendly: Vertical drill-down interface optimized for all screen sizes
- ๐ฏ Interactive Navigation: Click nodes to explore, select children to drill down
- ๐ Random Colors: Each topic gets a beautiful random soft color for visual variety
Topic Extraction โ Validation โ Visualization
- Gemini 2.0-flash: Initial hierarchical topic extraction from text
- Gemini 2.5-flash-preview: Validates accuracy, completeness, and structure
- Custom Interface: Renders topics in an intuitive drill-down format
Split Layout Design:
- Left Panel (50%): Vertical linear flow with circular topic nodes
- Right Panel (50%): Interactive children selection area
- No External Dependencies: Built from scratch, no React Flow or graph libraries
- Audio Input โ MediaRecorder API or file upload
- Transcription โ AssemblyAI converts speech to text
- Topic Extraction โ Gemini 2.0-flash extracts hierarchical structure
- Validation โ Gemini 2.5-flash-preview validates and corrects
- Visualization โ Custom interface renders interactive topic tree
npm installKey packages:
@langchain/google-genai- Gemini AI integrationassemblyai- Speech-to-text transcriptionlucide-react- Icons for the interface
Create a .env.local file:
# Google Gemini API Key (Required)
NEXT_PUBLIC_GOOGLE_API_KEY=your-google-api-key-here
# AssemblyAI API Key (Required)
NEXT_PUBLIC_ASSEMBLYAI_API_KEY=your-assemblyai-api-key-here- Go to Google AI Studio
- Create a new API key
- Add to
.env.localasNEXT_PUBLIC_GOOGLE_API_KEY
- Go to AssemblyAI
- Sign up and get your API key
- Add to
.env.localasNEXT_PUBLIC_ASSEMBLYAI_API_KEY
npm run devOpen http://localhost:3000 in your browser.
- Text Input: Type directly in the left panel textarea
- Voice Recording: Click the microphone button to record
- File Upload: Click upload button to process audio files
- Start: Process text/audio to see the root topic
- Explore: Click circular nodes to see their children in the right panel
- Drill Down: Click children to add them to the linear path
- Navigate Back: Click any previous node to return to that level
- Circular Nodes: Main topics displayed as colorful circles
- Accuracy Badges: Show AI confidence levels (percentage)
- Children Indicators: Small badges showing number of subtopics
- Connection Lines: Visual links between parent and child topics
- Right Panel: Interactive area for selecting and exploring children
Greens: Emerald, Lime, Green, Teal (various shades)
Blues: Cyan, Sky, Blue, Indigo (various shades)
Purples: Violet, Purple, Fuchsia (various shades)
Pinks: Pink, Rose (multiple tones)
Warm: Red, Orange, Amber, Yellow (soft versions)
Light Variants: Ultra-light versions of core colors
- Size: 128px diameter circles
- Random Colors: Each topic gets a unique soft pastel color
- Hover Effects: Scale animation on interaction
- Selection State: White ring around active node
- Typography: Dark text on light backgrounds for readability
- Focuses on explicitly mentioned topics only
- Creates logical hierarchical structures
- Captures specific values, measurements, and names
- Assigns accuracy scores based on content prominence
- Uses consistent ID naming patterns
- Reviews extraction accuracy against original text
- Checks for completeness and missing topics
- Validates hierarchical relationships
- Adjusts accuracy scores for better precision
- Removes hallucinated or incorrect topics
- 0.9-1.0: Extensively discussed with multiple details
- 0.7-0.9: Clearly mentioned with context
- 0.5-0.7: Mentioned with some detail
- 0.3-0.5: Briefly mentioned or implied