π€ OpenAI-compatible TTS API with automatic engine selection
AutoTTS is a modular Text-to-Speech server that automatically chooses the best TTS engine and language based on input text. It provides a fully OpenAI-compatible API that seamlessly integrates with OpenWebUI and other OpenAI-compatible clients.
- π Automatic Engine Selection: Intelligently chooses the best TTS engine based on language and quality
- π Multi-language Support: Automatic language detection with 50+ supported languages
- π Multiple Voices: Support for various voice types (male, female, neutral)
- π¦ Modular Architecture: Easy to add new TTS engines
- π OpenAI Compatible: Drop-in replacement for OpenAI TTS API
- πΎ Smart Caching: Reduces latency with intelligent audio caching
- π High Performance: Async/await architecture for concurrent requests
- OuteTTS: High-quality neural TTS with multilingual support
- Chatterbox TTS: Fast and efficient TTS engine
- More engines coming soon...
# Clone the repository
git clone https://github.com/Wladastic/AutoTTS.git
cd AutoTTS
# Install dependencies
pip install -r requirements.txt
# Copy and configure environment
cp .env.example .env
# Edit .env with your preferences# Start the server
python server.py
# Or with custom options
python server.py --host 0.0.0.0 --port 8000 --log-level infoThe server will be available at http://localhost:8000
# Test the API
python test_client.py --test-all
# Generate a test audio file
python test_client.py --text "Hello from AutoTTS!" --output hello.mp3POST /v1/audio/speech- Generate speech from textGET /v1/models- List available models
GET /health- Health checkGET /v1/voices- List available voicesGET /v1/languages- List supported languagesGET /v1/info- Server information
- In OpenWebUI settings, go to Audio β TTS Settings
- Set the TTS API URL to:
http://your-autotts-server:8000/v1/audio/speech - Choose any supported voice (alloy, echo, fable, onyx, nova, shimmer)
- AutoTTS will automatically detect the language and choose the best engine!
curl -X POST "http://localhost:8000/v1/audio/speech" \
-H "Content-Type: application/json" \
-d '{
"model": "tts-1",
"input": "Hello, this is AutoTTS speaking!",
"voice": "alloy",
"response_format": "mp3",
"speed": 1.0
}' \
--output speech.mp3curl "http://localhost:8000/v1/models"curl "http://localhost:8000/v1/voices"Create a .env file from .env.example:
# Server Configuration
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
# TTS Engine Configuration
ENABLE_OUTETTS=true
ENABLE_CHATTERBOX=true
# Language Detection
AUTO_DETECT_LANGUAGE=true
DEFAULT_LANGUAGE=en
# Cache Configuration
ENABLE_CACHE=true
CACHE_DIR=./cacheβββββββββββββββββββ ββββββββββββββββββββ βββββββββββββββββββ
β FastAPI App β β TTS Manager β β Language β
β β β β β Detector β
β /v1/audio/speechβββββΆβ Engine Selection βββββΆβ β
β /v1/models β β Quality Scoring β β Auto Detection β
β /health β β Fallback Logic β β 50+ Languages β
βββββββββββββββββββ ββββββββββββββββββββ βββββββββββββββββββ
β
βΌ
ββββββββββββββββββββ
β TTS Engines β
β β
β βββββββββββββββ β
β β OuteTTS β β
β βββββββββββββββ β
β βββββββββββββββ β
β β ChatterboxTTSβ β
β βββββββββββββββ β
β βββββββββββββββ β
β β Future β β
β β Engines β β
β βββββββββββββββ β
ββββββββββββββββββββ
- Create a new engine class inheriting from
TTSEngine - Implement the required methods:
initialize(),synthesize(),get_available_voices(), etc. - Register the engine in
tts_manager.py - Update configuration in
config.py
Example:
from engines.base import TTSEngine
class MyTTSEngine(TTSEngine):
def __init__(self, config):
super().__init__("MyTTS", config)
async def initialize(self):
# Initialize your TTS engine
return True
async def synthesize(self, text, voice, language, speed, format):
# Generate speech
return audio_bytes# Install development dependencies
pip install -r requirements.txt
# Run with auto-reload
python server.py --reload
# Run tests
python test_client.py --test-all- Fork the repository
- Create a feature branch
- Add your TTS engine or improvement
- Test thoroughly
- Submit a pull request
MIT License - see LICENSE file for details
- Inspired by openedai-speech
- OuteTTS engine support
- Chatterbox TTS integration
- OpenAI API compatibility
This project is licensed under the GNU Affero General Public License v3.0 (AGPL-3.0).