A command-line text-to-speech tool built with Qwen3-TTS, managed with uv, and exposed via a Click CLI.
This project is designed for:
- Language learning audio generation
- Assignment narration
- Accessibility overlays
- Classroom TTS labs
- Rapid speech prototyping from text files
- ๐ง Qwen3-TTS speech synthesis
- ๐ฃ๏ธ CustomVoice + VoiceDesign models
- ๐๏ธ Style prompting (
--instruct) - ๐ Text or file input
- ๐ป CPU or GPU support
- โก uv-managed reproducible environment
- ๐งฉ Click-based CLI
- ๐ค Smolagents integration for AI agent workflows
qwen_tts_cli/
โโโ main.py
โโโ pyproject.toml
โโโ README.md
โโโ .venv/
git clone <repo-url>
cd qwen_tts_cliOr initialize locally:
uv inituv syncDependencies include:
- qwen-tts
- transformers (pinned)
- numpy (pinned)
- soundfile
- click
uv run -- qwen-tts-cli \
--text "Ciao! Benvenuto all'esercizio." \
--out lesson.wavuv run -- qwen-tts-cli \
--text-file text/airport_lesson.txt \
--language Italian \
--instruct "Voce di insegnante, ritmo lento, pronuncia chiara." \
--out audio/airport.wavtext/airport_lesson.txt
All'aeroporto:
Dov'รจ il ritiro bagagli?
Dove posso prendere un taxi?
Dove si prende lo shuttle per andare alla Stazione Termini?
| Option | Description |
|---|---|
--model |
HF model id |
--language |
Spoken language |
--speaker |
CustomVoice speaker |
--instruct |
Voice style prompt |
--text |
Inline text |
--text-file |
Path to text file |
--out |
Output WAV file |
--device |
auto / cpu / cuda |
| Model | Use Case |
|---|---|
Qwen3-TTS-0.6B-CustomVoice |
Fast classroom demos |
Qwen3-TTS-0.6B-Base |
Voice cloning |
Qwen3-TTS-1.7B-VoiceDesign |
Style-designed voices |
Example:
--instruct "Voce di insegnante, lenta, incoraggiante."Other ideas:
- News anchor
- Tourist guide
- Sci-fi narrator
- ASMR whisper
- Language lab instructor
If running on CUDA (WSL + NVIDIA):
uv pip install \
--index-url https://download.pytorch.org/whl/cu121 \
torch torchvision torchaudioVerify:
uv run -- python -c "import torch; print(torch.cuda.is_available())"Run script directly:
uv run -- python src/qwen_tts_cli/cli.py --text "Test" --out audio/test.wavUse Qwen-TTS as a tool in AI agent workflows with smolagents:
Install smolagents support:
uv sync --extra smolagentsUse as a tool:
from smolagents import CodeAgent, HfApiModel
from qwen_tts_cli.smolagent_tool import QwenTTSTool
tts_tool = QwenTTSTool()
model = HfApiModel(model_id="Qwen/Qwen2.5-Coder-32B-Instruct")
agent = CodeAgent(tools=[tts_tool], model=model, add_base_tools=True)
agent.run("Generate Italian audio saying 'Benvenuto!' in a teacher's voice")See examples/README.md for more details.
- Language listening exercises
- Pronunciation drills
- Assignment narration
- Accessibility audio overlays
- LMS content generation
Ensure pinned version:
uv add "numpy==2.1.3"
uv syncPin version:
uv add "transformers==4.57.3"Use smaller checkpoint:
Qwen3-TTS-0.6B-CustomVoice
See upstream model + repo licenses:
- Qwen3-TTS
- Hugging Face Transformers
- QwenLM
- Hugging Face
- Astral uv
- Click CLI
- Preset voice styles
- Batch folder processing
- Subtitle (.srt) export
- Gradio web UI
- LMS integration tooling
Happy synthesizing ๐๏ธ