ilavanyajain/voice-to-text-transcriber

โ˜… 1Forks 0PythonGitHub โ†—Compare

README

๐ŸŽค Voice-to-Text Transcriber

A powerful, multi-language voice-to-text transcription tool that automatically detects languages and translates audio content to English. Built with OpenAI's Whisper API and GPT-4 for high-accuracy transcription and translation.

โœจ Features

  • ๐ŸŒ Multi-language Support: Automatically detects and transcribes audio in any language
  • ๐Ÿ‡บ๐Ÿ‡ธ English Translation: Translates all detected languages to English
  • ๐ŸŽต Multiple Audio Formats: Supports both MP3 and OPUS audio files
  • ๐Ÿ”„ Automatic Conversion: Seamlessly converts OPUS files to MP3 before processing
  • ๐Ÿ“Š Detailed Analysis: Provides language detection, timestamps, and segment analysis
  • ๐Ÿ’พ JSON Export: Saves all transcription results in structured JSON format
  • ๐Ÿ”’ Secure: Uses environment variables for API key management

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.11 or higher
  • FFmpeg (for audio conversion)
  • OpenAI API key

Installation

  1. Clone or download the project files

  2. Install Python dependencies:

    pip install -r requirements.txt
  3. Install FFmpeg (if not already installed):

    • macOS: brew install ffmpeg
    • Ubuntu/Debian: sudo apt install ffmpeg
    • Windows: Download from ffmpeg.org
  4. Set up your API key: Create a .env file in the project root:

    OPENAI_API_KEY=your_openai_api_key_here

๐Ÿ“ Project Structure

voice-to-text-transcriber/
โ”œโ”€โ”€ main.py                          # Main transcription script
โ”œโ”€โ”€ opus_to_mp3_converter.py        # OPUS to MP3 converter utility
โ”œโ”€โ”€ requirements.txt                 # Python dependencies
โ”œโ”€โ”€ .env                            # Environment variables (create this)
โ”œโ”€โ”€ README.md                       # This file
โ””โ”€โ”€ audio_files/                    # Your audio files go here
    โ”œโ”€โ”€ hindi-demo-voice.opus
    โ”œโ”€โ”€ japanese-text.opus
    โ””โ”€โ”€ other-audio-files.mp3

๐ŸŽฏ Usage

Basic Transcription

  1. Place your audio file in the project directory
  2. Update the filename in main.py (line ~200):
    audio_file = "your-audio-file.opus"  # or .mp3
  3. Run the script:
    python3 main.py

Supported Audio Formats

  • Input: .opus, .mp3
  • Output: Automatically converts OPUS to MP3, then processes for transcription

Example Output

๐ŸŽค Multi-Language Voice to English Transcription
==================================================
๐ŸŽต Detected .opus file: hindi-demo-voice.opus
๐Ÿ”„ Converting hindi-demo-voice.opus to MP3 format...
โœ… Conversion successful: hindi-demo-voice.mp3

๐Ÿ”„ Attempting multilingual analysis...
๐ŸŽต Analyzing multilingual audio: hindi-demo-voice.mp3
๐Ÿ”„ Starting advanced language analysis...
โœ… Advanced transcription completed!
๐Ÿ” Analyzing language segments...
๐ŸŒ Detected languages: ['hindi']

==================================================
๐Ÿ“Š TRANSCRIPTION RESULTS
==================================================
๐ŸŒ Languages detected: hindi

๐Ÿ“ English Translations:

[HINDI] โ†’ English:
Hello, how are you today? I hope you're doing well.

๐Ÿ’พ Results saved to: transcription_results.json

๐Ÿ”ง Configuration

Environment Variables

Create a .env file with:

OPENAI_API_KEY=your_openai_api_key_here

API Settings

The script uses:

  • Whisper Model: whisper-1 for transcription
  • GPT Model: gpt-4o-mini for translation
  • Temperature: 0.0 for transcription, 0.1 for translation

๐Ÿ“Š Output Files

transcription_results.json

Contains structured data:

{
  "detected_languages": ["hindi"],
  "language_segments": {
    "hindi": [
      {
        "start": 0.0,
        "end": 3.5,
        "text": "เคจเคฎเคธเฅเคคเฅ‡, เค•เฅˆเคธเฅ‡ เคนเฅ‹ เค†เคช?"
      }
    ]
  },
  "english_translations": {
    "hindi": "Hello, how are you?"
  },
  "full_transcript": "เคจเคฎเคธเฅเคคเฅ‡, เค•เฅˆเคธเฅ‡ เคนเฅ‹ เค†เคช?"
}

๐Ÿ› ๏ธ Advanced Features

OPUS to MP3 Conversion

The opus_to_mp3_converter.py utility provides:

  • High-quality MP3 conversion
  • Batch processing capabilities
  • Configurable quality settings
  • Metadata preservation

Usage:

# Single file conversion
python3 opus_to_mp3_converter.py input.opus

# Batch conversion
python3 opus_to_mp3_converter.py input_directory --batch

# Quality options
python3 opus_to_mp3_converter.py input.opus -q high    # Best quality
python3 opus_to_mp3_converter.py input.opus -q medium  # Good quality
python3 opus_to_mp3_converter.py input.opus -q low     # Standard quality

Multilingual Analysis

The main script automatically:

  • Detects language changes in audio
  • Segments audio by language
  • Translates each segment to English
  • Provides comprehensive analysis

๐Ÿ” Troubleshooting

Common Issues

  1. "FFmpeg not found":

    • Install FFmpeg: brew install ffmpeg (macOS) or sudo apt install ffmpeg (Ubuntu)
  2. "API key not found":

    • Check your .env file exists and contains OPENAI_API_KEY=your_key
  3. Audio file not found:

    • Verify the filename in main.py matches your actual audio file
  4. Conversion failed:

    • Ensure FFmpeg is properly installed
    • Check audio file format and integrity

Error Messages

  • โŒ Audio file not found: Check filename and path
  • โŒ Conversion failed: Verify FFmpeg installation
  • โŒ API key not found: Check .env file
  • โŒ Transcription failed: Check API key validity and audio quality

๐Ÿ“ˆ Performance Tips

  • Audio Quality: Higher quality audio = better transcription accuracy
  • File Size: Larger files take longer to process
  • Language Complexity: Some languages may require more processing time
  • API Limits: Be mindful of OpenAI API rate limits and costs

๐Ÿ” Security

  • Never commit your .env file to version control
  • Keep your API key private and secure
  • Use environment variables for all sensitive configuration

๐Ÿ“ License

This project is for educational and personal use. Please respect OpenAI's terms of service and API usage policies.

๐Ÿค Contributing

Feel free to:

  • Report bugs and issues
  • Suggest new features
  • Improve documentation
  • Optimize performance

๐Ÿ“ž Support

For issues or questions:

  1. Check the troubleshooting section
  2. Verify your setup and configuration
  3. Check OpenAI API status and limits

Happy Transcribing! ๐ŸŽตโœจ

Contributors

ilavanyajain

Issues