A Voice Activity Projection (VAP) based turn detection model for real-time conversation analysis, inspired by LiveKit's turn detector architecture.
# Clone and setup
git clone <repository-url>
cd vap
python -m venv vap_env
source vap_env/bin/activate # On Windows: vap_env\Scripts\activate
pip install -r requirements.txt
# Run complete pipeline (setup, train baseline, evaluate)
python scripts/run_pipeline.py# After completing Phases 1-2, run optimization locally
python scripts/run_phase3.py# For 10-25x faster training, use RunPod GPU acceleration
python runpod/setup_runpod.py # Setup instructions
# Follow the deployment guide: runpod/DEPLOYMENT_GUIDE.mdโจ NEW: Enhanced Progress Indicators
- Real-time Training Dashboard: Live progress bars, epoch tracking, and performance metrics
- Smart Status Updates: Batch-by-batch progress with loss and accuracy monitoring
- Training Analytics: Epoch summaries, trend analysis, and time estimates
- Performance Tracking: Best validation accuracy monitoring and improvement alerts
๐ NEW: RunPod GPU Acceleration
- 10-25x Faster Training: GPU acceleration vs local CPU
- Cost Effective: $1.20-$4.80 for complete training
- Professional Infrastructure: RTX 4090 with 24GB VRAM
- Easy Deployment: One-click GPU pod deployment
# 1. Setup data
python scripts/setup_phase1.py
# 2. Train baseline model
python scripts/train_baseline.py
# 3. Evaluate baseline
python scripts/evaluate_baseline.py
# 4. Train optimized model (with progress indicators)
python scripts/train_optimized.py
# 5. Evaluate optimized model
python scripts/evaluate_optimized.pyPhase 3 In Progress ๐ง - Model Optimization Underway!
- Model: 128-dim, 2-layer, 4-head transformer (3x capacity increase)
- Training: Enhanced with real VAP labels and audio augmentation
- Performance: Targeting >25% accuracy improvement over baseline
- Architecture: Multi-task learning with 5 prediction heads
- Data: Real audio processing + comprehensive augmentation pipeline
- Status: Phase 3 infrastructure complete, ready for training
The Phase 3 optimization includes a comprehensive training dashboard that provides:
- Live Progress Bars: Visual progress indicators for each epoch and batch
- Performance Metrics: Real-time loss and accuracy tracking
- Epoch Summaries: Detailed statistics after each training epoch
- Trend Analysis: Performance improvement tracking and alerts
- Time Estimates: Accurate predictions for training completion
๐ EPOCH 1/50
===========================================================
Epoch 1: 45%|โโโโโโโโโโโโโโโโโ | 439/975 [02:15<02:45, 3.24batch/s]
๐ EPOCH 1 SUMMARY
--------------------------------------------------
โฑ๏ธ Duration: 135.2s
๐ Train Loss: 2.8476
๐ Val Loss: 2.2341
๐ฏ Val Accuracy: 0.1876
๐ NEW BEST VALIDATION ACCURACY: 0.1876
๐ TRAINING PROGRESS SUMMARY
----------------------------------------
Completed Epochs: 1/50
Progress: 2.0%
Average Epoch Time: 135.2s
Total Training Time: 2.3 minutes
Estimated Time Remaining: 110.2 minutes
Best Validation Accuracy: 0.1876
Recent Trend: โ๏ธ Improving
- Batch Progress: Updates every 50 training batches and 20 validation batches
- Performance Alerts: Immediate notification of new best accuracies
- Trend Detection: Automatic analysis of performance improvements
- Resource Tracking: Memory usage and training efficiency monitoring
- Audio Encoder: Log-Mel spectrogram + CNN downsampling
- Cross-Attention Transformer: Speaker interaction modeling
- Multi-task Heads: VAP patterns, EoT, backchannel, overlap, VAD
- Baseline: 64-dim, 1-layer, 2-heads (Phase 2 - Complete)
- Optimized: 128-dim, 2-layers, 4-heads (Phase 3 - In Progress)
- Standard: 256-dim, 4-layers, 8-heads (Phase 4 - Planned)
vap/
โโโ vap/ # Core package
โ โโโ models/ # Model definitions
โ โโโ tasks/ # Training tasks
โ โโโ data/ # Data processing + augmentation
โ โโโ streaming/ # Real-time inference
โโโ scripts/ # Executable scripts
โ โโโ run_pipeline.py # Complete pipeline runner
โ โโโ train_baseline.py # Baseline training (Phase 2)
โ โโโ train_optimized.py # Optimized training (Phase 3)
โ โโโ evaluate_baseline.py # Baseline evaluation
โ โโโ evaluate_optimized.py # Optimized evaluation
โ โโโ run_phase3.py # Phase 3 optimization pipeline
โ โโโ inference.py # Inference demo
โโโ configs/ # Configuration files
โ โโโ vap_light.yaml # Baseline configuration
โ โโโ vap_optimized.yaml # Optimized configuration (Phase 3)
โโโ data/ # Dataset storage
โโโ checkpoints/ # Trained models
โโโ results/ # Evaluation results
โโโ tests/ # Unit tests
| Script | Purpose | Status |
|---|---|---|
run_pipeline.py |
Complete pipeline runner (Phases 1-2) | โ Ready |
run_phase3.py |
Phase 3 optimization pipeline | โ Ready |
train_baseline.py |
Baseline model training | โ Ready |
train_optimized.py |
Optimized model training | โ Ready |
evaluate_baseline.py |
Baseline performance evaluation | โ Ready |
evaluate_optimized.py |
Optimized performance evaluation | โ Ready |
inference.py |
Inference demo | โ Ready |
setup_phase1.py |
Data preparation | โ Ready |
download_realtime_dataset.py |
LibriSpeech download | โ Ready |
- VAP Pattern Accuracy: 4.73%
- EoT Accuracy: 0.00%
- Backchannel Accuracy: 0.00%
- Overlap Accuracy: 0.00%
- VAD Accuracy: 50.00%
- Overall Accuracy: 10.95%
- Target Improvement: >25% overall accuracy
- Model Capacity: 3x parameter increase (378K โ ~1.5M)
- Architecture: Enhanced transformer (128-dim, 2-layers, 4-heads)
- Training: Real VAP labels + comprehensive augmentation
- Phase 3: Complete optimized training and evaluation
- Phase 4: Performance benchmarking and hyperparameter tuning
- Phase 5: Production deployment and optimization
- Real-time Turn Detection: Live conversation analysis
- Meeting Transcription: Speaker diarization enhancement
- Voice Assistants: Turn-taking behavior modeling
- Research: Conversation dynamics analysis
- Python 3.9+
- PyTorch 2.0+
- PyTorch Lightning
- Audio processing libraries
# Run all tests
python -m pytest tests/
# Run specific test
python -m pytest tests/test_model.py- Implement in
vap/package - Add tests in
tests/ - Update scripts as needed
- Document in README
- VAP Architecture: Voice Activity Projection for turn detection
- LiveKit Turn Detector: Production turn detection system
- LibriSpeech: Open-source speech dataset
- PyTorch Lightning: Training framework
- Fork the repository
- Create feature branch
- Implement changes with tests
- Submit pull request
[Add your license here]
Last Updated: August 16, 2025
Status: Phase 3 In Progress - Model Optimization Underway
Next Milestone: Phase 4 - Performance Benchmarking