ricklon/toaster-3000

AI voice agent with toaster personality combining TTS, STT, and conversational AI

β˜… 0Forks 0PythonGitHub β†—Compare

README

🍞 Toaster 3000

The world's smartest and most advanced toaster - An AI voice agent with an enthusiastic toaster personality

Python 3.12+ License: MIT Code style: black

Toaster 3000 is an AI voice agent that combines Text-to-Speech (TTS), Speech-to-Text (STT), and conversational AI capabilities with a unique toaster-themed personality. The agent is convinced that toasting is the solution to most of life's problems and enthusiastically suggests toasting-related solutions to user queries.

✨ Features

  • 🎀 Multi-Modal Input: Support for text input, push-to-talk, and continuous listening modes
  • πŸ”Š Advanced TTS: Sequential audio generation using Kokoro TTS for natural speech output
  • πŸ‘‚ Speech Recognition: High-quality speech-to-text using Faster Whisper
  • πŸ€– AI Conversation: Powered by HuggingFace models via smolagents framework
  • 🌐 Modern UI: Interactive Gradio web interface with custom toaster-themed styling
  • πŸ’¬ Memory: Maintains conversation context for natural dialogue flow
  • ⚑ Real-time: Background threading for seamless audio processing
  • πŸ”§ Configurable: Adjustable reasoning steps and model selection
  • πŸ‘₯ Session-Based: Each user gets an isolated session for concurrent multi-user support

πŸš€ Quick Start

Prerequisites

Installation

  1. Clone the repository

    git clone https://github.com/ricklon/toaster-3000.git
    cd toaster-3000
  2. Install dependencies

    uv sync --all-extras
  3. Set up environment variables

    Copy the example file and fill in your token:

    cp .env.example .env

    Or create a .env file in the project root:

    HUGGINGFACE_API_KEY=your_token_here
    MODEL_NAME=meta-llama/Llama-3.3-70B-Instruct  # Optional
    GRADIO_SHARE=false  # Optional; set true only when you need a public URL
  4. Run the application

    uv run toaster
  5. Open your browser to the provided local URL (typically http://127.0.0.1:7860)

    The first run can take longer while TTS and speech recognition models load. Missing HUGGINGFACE_API_KEY now fails immediately with a clear startup error.

πŸ’¬ Usage Examples

Text Chat

Simply type your questions in the text input field:

User: "How do I make perfect toast?"
Toaster 3000: "Great question! The key to perfect toast is understanding the golden ratio of heat, time, and bread thickness..."

Voice Interaction

  • Push-to-Talk: Click the microphone button, speak your question, then stop recording
  • Continuous Listening: Enable continuous mode for hands-free interaction with automatic speech detection

Example Conversations

  • "I'm having a bad day" 🍞 Suggests making toast to improve mood
  • "Help me code a function" πŸ’» Provides coding help with toast-related examples
  • "What's the weather like?" β˜€οΈ Discusses weather while recommending appropriate toast types

πŸ—οΈ Architecture

Toaster 3000 uses a modular, session-based architecture that supports concurrent users with isolated state.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Gradio UI     β”‚     β”‚   Session Mgmt   β”‚     β”‚  Audio Pipeline  β”‚
β”‚ β€’ Text Input    │────▢│ β€’ SessionManager │────▢│ β€’ TTSService     β”‚
β”‚ β€’ Voice Input   β”‚     β”‚ β€’ ToasterSession β”‚     β”‚ β€’ STTService     β”‚
β”‚ β€’ Audio Output  β”‚     β”‚ β€’ ChatHistory    β”‚     β”‚ β€’ Sequential TTS β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                        β”‚  Shared Runtime   β”‚
                        β”‚ β€’ Agent (singleton)β”‚
                        β”‚ β€’ TTS Model       β”‚
                        β”‚ β€’ Whisper Model   β”‚
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                        β”‚ HuggingFace API  β”‚
                        β”‚ β€’ InferenceModel β”‚
                        β”‚ β€’ smolagents     β”‚
                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Source Structure

src/toaster_3000/
β”œβ”€β”€ __init__.py          # Package initialization
β”œβ”€β”€ main.py              # Application entry point
β”œβ”€β”€ app.py               # Gradio UI and event handlers
β”œβ”€β”€ config.py            # Immutable configuration dataclass
β”œβ”€β”€ constants.py         # System prompts and defaults
β”œβ”€β”€ runtime.py           # Shared model singleton (ToasterRuntime)
β”œβ”€β”€ session.py           # Per-user session state (ToasterSession, ChatHistoryManager)
β”œβ”€β”€ session_manager.py   # Session lifecycle management
β”œβ”€β”€ services.py          # TTSService and STTService wrappers
└── theme.py             # Gradio theme and CSS

πŸ› οΈ Development

Code Quality

# Format code
uv run black src/ tests/

# Sort imports
uv run isort src/ tests/

# Type checking
uv run mypy src/ tests/

# Linting
uv run flake8 src/ tests/

Testing

# Run all tests
uv run pytest

# Run specific test file
uv run pytest tests/test_new_architecture.py -v

# Run UI tests
uv run pytest tests/ui/test_toaster_interface.py -v

Building

# Build package
uv build

# Install locally in development mode
uv sync --all-extras

πŸ“‹ Requirements

System Requirements

  • OS: Windows 10+, macOS 10.14+, Linux
  • Memory: 4GB RAM minimum, 8GB recommended
  • Storage: 2GB free space for dependencies

Python Dependencies

  • fastrtc - Real-time communication and TTS
  • smolagents - AI agent framework
  • gradio>=5.0.0 - Web UI framework
  • faster-whisper - Speech recognition
  • python-dotenv - Environment management
  • sounddevice & soundfile - Audio processing
  • kokoro-onnx>=0.5.0 - Kokoro TTS model

See pyproject.toml for complete dependency list.

βš™οΈ Configuration

Environment Variables

Variable Required Default Description
HUGGINGFACE_API_KEY βœ… - Your HuggingFace API token
MODEL_NAME ❌ meta-llama/Llama-3.3-70B-Instruct Model to use for AI responses
GRADIO_SHARE ❌ false Set to true to request a public Gradio share URL

Model Options

The application uses tool-capable models via the HuggingFace Inference API (smolagents CodeAgent). Latest recommended models:

  • Qwen/Qwen3-Coder-Next (default β€” latest code-specialized model)
  • Qwen/Qwen3-14B
  • google/gemma-4-31B-it
  • google/gemma-4-26B-A4B-it (MoE, efficient)
  • mistralai/Mistral-Small-4-119B-2603
  • mistralai/Devstral-Small-2-24B-Instruct-2512
  • meta-llama/Llama-3.3-70B-Instruct

Models switch live at runtime β€” no restart required.

Audio Settings

  • TTS Voice: am_liam (Kokoro TTS)
  • STT Model: tiny.en (Faster Whisper)
  • Audio Quality: 16kHz sample rate
  • Response Chunking: 300 character segments for long responses

🀝 Contributing

We welcome contributions! Please see our contributing guidelines and:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Development Setup

# Clone your fork
git clone https://github.com/your-username/toaster-3000.git

# Install with dev dependencies
uv sync --all-extras

# Run tests before submitting
uv run pytest

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ™ Acknowledgments

πŸ“š Documentation

πŸ”§ Troubleshooting

Common Issues

"No module named 'toaster_3000'"

uv sync --all-extras

"HUGGINGFACE_API_KEY not found"

Audio issues on Windows

# Install Windows audio dependencies
uv add pyaudio sounddevice

Model loading errors

  • Check your internet connection
  • Verify the model name is correct
  • Ensure your HuggingFace API key has appropriate permissions

For more issues, check our GitHub Issues or create a new one.


Made with ❀️ and lots of 🍞 by ricklon

"Remember, whatever life problems you're facing, toasting something will probably help. That's the Toaster 3000 guarantee!" 🍞✨

Contributors

ricklon

Issues