The world's smartest and most advanced toaster - An AI voice agent with an enthusiastic toaster personality
Toaster 3000 is an AI voice agent that combines Text-to-Speech (TTS), Speech-to-Text (STT), and conversational AI capabilities with a unique toaster-themed personality. The agent is convinced that toasting is the solution to most of life's problems and enthusiastically suggests toasting-related solutions to user queries.
- π€ Multi-Modal Input: Support for text input, push-to-talk, and continuous listening modes
- π Advanced TTS: Sequential audio generation using Kokoro TTS for natural speech output
- π Speech Recognition: High-quality speech-to-text using Faster Whisper
- π€ AI Conversation: Powered by HuggingFace models via smolagents framework
- π Modern UI: Interactive Gradio web interface with custom toaster-themed styling
- π¬ Memory: Maintains conversation context for natural dialogue flow
- β‘ Real-time: Background threading for seamless audio processing
- π§ Configurable: Adjustable reasoning steps and model selection
- π₯ Session-Based: Each user gets an isolated session for concurrent multi-user support
- Python 3.10+ (tested with Python 3.12)
- HuggingFace API Token (Get one here)
- uv package manager (Install uv)
-
Clone the repository
git clone https://github.com/ricklon/toaster-3000.git cd toaster-3000 -
Install dependencies
uv sync --all-extras
-
Set up environment variables
Copy the example file and fill in your token:
cp .env.example .env
Or create a
.envfile in the project root:HUGGINGFACE_API_KEY=your_token_here MODEL_NAME=meta-llama/Llama-3.3-70B-Instruct # Optional GRADIO_SHARE=false # Optional; set true only when you need a public URL
-
Run the application
uv run toaster
-
Open your browser to the provided local URL (typically
http://127.0.0.1:7860)The first run can take longer while TTS and speech recognition models load. Missing
HUGGINGFACE_API_KEYnow fails immediately with a clear startup error.
Simply type your questions in the text input field:
User: "How do I make perfect toast?"
Toaster 3000: "Great question! The key to perfect toast is understanding the golden ratio of heat, time, and bread thickness..."
- Push-to-Talk: Click the microphone button, speak your question, then stop recording
- Continuous Listening: Enable continuous mode for hands-free interaction with automatic speech detection
- "I'm having a bad day" π Suggests making toast to improve mood
- "Help me code a function" π» Provides coding help with toast-related examples
- "What's the weather like?" βοΈ Discusses weather while recommending appropriate toast types
Toaster 3000 uses a modular, session-based architecture that supports concurrent users with isolated state.
βββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββ
β Gradio UI β β Session Mgmt β β Audio Pipeline β
β β’ Text Input ββββββΆβ β’ SessionManager ββββββΆβ β’ TTSService β
β β’ Voice Input β β β’ ToasterSession β β β’ STTService β
β β’ Audio Output β β β’ ChatHistory β β β’ Sequential TTS β
βββββββββββββββββββ ββββββββββ¬ββββββββββ ββββββββββββββββββββ
β
ββββββββββΌββββββββββ
β Shared Runtime β
β β’ Agent (singleton)β
β β’ TTS Model β
β β’ Whisper Model β
ββββββββββ¬ββββββββββ
β
ββββββββββΌββββββββββ
β HuggingFace API β
β β’ InferenceModel β
β β’ smolagents β
ββββββββββββββββββββ
src/toaster_3000/
βββ __init__.py # Package initialization
βββ main.py # Application entry point
βββ app.py # Gradio UI and event handlers
βββ config.py # Immutable configuration dataclass
βββ constants.py # System prompts and defaults
βββ runtime.py # Shared model singleton (ToasterRuntime)
βββ session.py # Per-user session state (ToasterSession, ChatHistoryManager)
βββ session_manager.py # Session lifecycle management
βββ services.py # TTSService and STTService wrappers
βββ theme.py # Gradio theme and CSS
# Format code
uv run black src/ tests/
# Sort imports
uv run isort src/ tests/
# Type checking
uv run mypy src/ tests/
# Linting
uv run flake8 src/ tests/# Run all tests
uv run pytest
# Run specific test file
uv run pytest tests/test_new_architecture.py -v
# Run UI tests
uv run pytest tests/ui/test_toaster_interface.py -v# Build package
uv build
# Install locally in development mode
uv sync --all-extras- OS: Windows 10+, macOS 10.14+, Linux
- Memory: 4GB RAM minimum, 8GB recommended
- Storage: 2GB free space for dependencies
fastrtc- Real-time communication and TTSsmolagents- AI agent frameworkgradio>=5.0.0- Web UI frameworkfaster-whisper- Speech recognitionpython-dotenv- Environment managementsounddevice&soundfile- Audio processingkokoro-onnx>=0.5.0- Kokoro TTS model
See pyproject.toml for complete dependency list.
| Variable | Required | Default | Description |
|---|---|---|---|
HUGGINGFACE_API_KEY |
β | - | Your HuggingFace API token |
MODEL_NAME |
β | meta-llama/Llama-3.3-70B-Instruct |
Model to use for AI responses |
GRADIO_SHARE |
β | false |
Set to true to request a public Gradio share URL |
The application uses tool-capable models via the HuggingFace Inference API
(smolagents CodeAgent). Latest recommended models:
Qwen/Qwen3-Coder-Next(default β latest code-specialized model)Qwen/Qwen3-14Bgoogle/gemma-4-31B-itgoogle/gemma-4-26B-A4B-it(MoE, efficient)mistralai/Mistral-Small-4-119B-2603mistralai/Devstral-Small-2-24B-Instruct-2512meta-llama/Llama-3.3-70B-Instruct
Models switch live at runtime β no restart required.
- TTS Voice:
am_liam(Kokoro TTS) - STT Model:
tiny.en(Faster Whisper) - Audio Quality: 16kHz sample rate
- Response Chunking: 300 character segments for long responses
We welcome contributions! Please see our contributing guidelines and:
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
# Clone your fork
git clone https://github.com/your-username/toaster-3000.git
# Install with dev dependencies
uv sync --all-extras
# Run tests before submitting
uv run pytestThis project is licensed under the MIT License - see the LICENSE file for details.
- HuggingFace - For providing the model inference API
- smolagents - AI agent framework
- Gradio - Web UI framework
- FastRTC - Real-time communication tools
- Faster Whisper - Speech recognition
- CLAUDE.md - Developer guide for Claude Code users
- API Documentation - Detailed API reference
- Deployment Guide - Production deployment instructions
"No module named 'toaster_3000'"
uv sync --all-extras"HUGGINGFACE_API_KEY not found"
- Ensure your
.envfile is in the project root - Check your API key is valid at HuggingFace Settings
Audio issues on Windows
# Install Windows audio dependencies
uv add pyaudio sounddeviceModel loading errors
- Check your internet connection
- Verify the model name is correct
- Ensure your HuggingFace API key has appropriate permissions
For more issues, check our GitHub Issues or create a new one.
Made with β€οΈ and lots of π by ricklon
"Remember, whatever life problems you're facing, toasting something will probably help. That's the Toaster 3000 guarantee!" πβ¨