A powerful, multi-language voice-to-text transcription tool that automatically detects languages and translates audio content to English. Built with OpenAI's Whisper API and GPT-4 for high-accuracy transcription and translation.
- ๐ Multi-language Support: Automatically detects and transcribes audio in any language
- ๐บ๐ธ English Translation: Translates all detected languages to English
- ๐ต Multiple Audio Formats: Supports both MP3 and OPUS audio files
- ๐ Automatic Conversion: Seamlessly converts OPUS files to MP3 before processing
- ๐ Detailed Analysis: Provides language detection, timestamps, and segment analysis
- ๐พ JSON Export: Saves all transcription results in structured JSON format
- ๐ Secure: Uses environment variables for API key management
- Python 3.11 or higher
- FFmpeg (for audio conversion)
- OpenAI API key
-
Clone or download the project files
-
Install Python dependencies:
pip install -r requirements.txt
-
Install FFmpeg (if not already installed):
- macOS:
brew install ffmpeg - Ubuntu/Debian:
sudo apt install ffmpeg - Windows: Download from ffmpeg.org
- macOS:
-
Set up your API key: Create a
.envfile in the project root:OPENAI_API_KEY=your_openai_api_key_here
voice-to-text-transcriber/
โโโ main.py # Main transcription script
โโโ opus_to_mp3_converter.py # OPUS to MP3 converter utility
โโโ requirements.txt # Python dependencies
โโโ .env # Environment variables (create this)
โโโ README.md # This file
โโโ audio_files/ # Your audio files go here
โโโ hindi-demo-voice.opus
โโโ japanese-text.opus
โโโ other-audio-files.mp3
- Place your audio file in the project directory
- Update the filename in
main.py(line ~200):audio_file = "your-audio-file.opus" # or .mp3
- Run the script:
python3 main.py
- Input:
.opus,.mp3 - Output: Automatically converts OPUS to MP3, then processes for transcription
๐ค Multi-Language Voice to English Transcription
==================================================
๐ต Detected .opus file: hindi-demo-voice.opus
๐ Converting hindi-demo-voice.opus to MP3 format...
โ
Conversion successful: hindi-demo-voice.mp3
๐ Attempting multilingual analysis...
๐ต Analyzing multilingual audio: hindi-demo-voice.mp3
๐ Starting advanced language analysis...
โ
Advanced transcription completed!
๐ Analyzing language segments...
๐ Detected languages: ['hindi']
==================================================
๐ TRANSCRIPTION RESULTS
==================================================
๐ Languages detected: hindi
๐ English Translations:
[HINDI] โ English:
Hello, how are you today? I hope you're doing well.
๐พ Results saved to: transcription_results.json
Create a .env file with:
OPENAI_API_KEY=your_openai_api_key_hereThe script uses:
- Whisper Model:
whisper-1for transcription - GPT Model:
gpt-4o-minifor translation - Temperature: 0.0 for transcription, 0.1 for translation
Contains structured data:
{
"detected_languages": ["hindi"],
"language_segments": {
"hindi": [
{
"start": 0.0,
"end": 3.5,
"text": "เคจเคฎเคธเฅเคคเฅ, เคเฅเคธเฅ เคนเฅ เคเคช?"
}
]
},
"english_translations": {
"hindi": "Hello, how are you?"
},
"full_transcript": "เคจเคฎเคธเฅเคคเฅ, เคเฅเคธเฅ เคนเฅ เคเคช?"
}The opus_to_mp3_converter.py utility provides:
- High-quality MP3 conversion
- Batch processing capabilities
- Configurable quality settings
- Metadata preservation
Usage:
# Single file conversion
python3 opus_to_mp3_converter.py input.opus
# Batch conversion
python3 opus_to_mp3_converter.py input_directory --batch
# Quality options
python3 opus_to_mp3_converter.py input.opus -q high # Best quality
python3 opus_to_mp3_converter.py input.opus -q medium # Good quality
python3 opus_to_mp3_converter.py input.opus -q low # Standard qualityThe main script automatically:
- Detects language changes in audio
- Segments audio by language
- Translates each segment to English
- Provides comprehensive analysis
-
"FFmpeg not found":
- Install FFmpeg:
brew install ffmpeg(macOS) orsudo apt install ffmpeg(Ubuntu)
- Install FFmpeg:
-
"API key not found":
- Check your
.envfile exists and containsOPENAI_API_KEY=your_key
- Check your
-
Audio file not found:
- Verify the filename in
main.pymatches your actual audio file
- Verify the filename in
-
Conversion failed:
- Ensure FFmpeg is properly installed
- Check audio file format and integrity
- โ Audio file not found: Check filename and path
- โ Conversion failed: Verify FFmpeg installation
- โ API key not found: Check
.envfile - โ Transcription failed: Check API key validity and audio quality
- Audio Quality: Higher quality audio = better transcription accuracy
- File Size: Larger files take longer to process
- Language Complexity: Some languages may require more processing time
- API Limits: Be mindful of OpenAI API rate limits and costs
- Never commit your
.envfile to version control - Keep your API key private and secure
- Use environment variables for all sensitive configuration
This project is for educational and personal use. Please respect OpenAI's terms of service and API usage policies.
Feel free to:
- Report bugs and issues
- Suggest new features
- Improve documentation
- Optimize performance
For issues or questions:
- Check the troubleshooting section
- Verify your setup and configuration
- Check OpenAI API status and limits
Happy Transcribing! ๐ตโจ