ricklon/qwen_tts_cli

A command-line text-to-speech tool built with Qwen3-TTS for educational applications

โ˜… 0Forks 0PythonGitHub โ†—Compare

README

๐ŸŽ™๏ธ Qwen-TTS CLI

A command-line text-to-speech tool built with Qwen3-TTS, managed with uv, and exposed via a Click CLI.

This project is designed for:

  • Language learning audio generation
  • Assignment narration
  • Accessibility overlays
  • Classroom TTS labs
  • Rapid speech prototyping from text files

โœจ Features

  • ๐Ÿง  Qwen3-TTS speech synthesis
  • ๐Ÿ—ฃ๏ธ CustomVoice + VoiceDesign models
  • ๐ŸŽ›๏ธ Style prompting (--instruct)
  • ๐Ÿ“„ Text or file input
  • ๐Ÿ’ป CPU or GPU support
  • โšก uv-managed reproducible environment
  • ๐Ÿงฉ Click-based CLI
  • ๐Ÿค– Smolagents integration for AI agent workflows

๐Ÿ“ฆ Project Structure

qwen_tts_cli/
โ”œโ”€โ”€ main.py
โ”œโ”€โ”€ pyproject.toml
โ”œโ”€โ”€ README.md
โ””โ”€โ”€ .venv/

๐Ÿš€ Quick Start

1๏ธโƒฃ Clone / create project

git clone <repo-url>
cd qwen_tts_cli

Or initialize locally:

uv init

2๏ธโƒฃ Install dependencies

uv sync

Dependencies include:

  • qwen-tts
  • transformers (pinned)
  • numpy (pinned)
  • soundfile
  • click

3๏ธโƒฃ Run the CLI

uv run -- qwen-tts-cli \
  --text "Ciao! Benvenuto all'esercizio." \
  --out lesson.wav

๐Ÿ—ฃ๏ธ Example: Language Assignment Audio

uv run -- qwen-tts-cli \
  --text-file text/airport_lesson.txt \
  --language Italian \
  --instruct "Voce di insegnante, ritmo lento, pronuncia chiara." \
  --out audio/airport.wav

๐Ÿ“„ Example Input File

text/airport_lesson.txt

All'aeroporto:

Dov'รจ il ritiro bagagli?
Dove posso prendere un taxi?
Dove si prende lo shuttle per andare alla Stazione Termini?

๐ŸŽ›๏ธ CLI Options

Option Description
--model HF model id
--language Spoken language
--speaker CustomVoice speaker
--instruct Voice style prompt
--text Inline text
--text-file Path to text file
--out Output WAV file
--device auto / cpu / cuda

๐Ÿง  Model Examples

Model Use Case
Qwen3-TTS-0.6B-CustomVoice Fast classroom demos
Qwen3-TTS-0.6B-Base Voice cloning
Qwen3-TTS-1.7B-VoiceDesign Style-designed voices

๐ŸŽจ Voice Style Prompting

Example:

--instruct "Voce di insegnante, lenta, incoraggiante."

Other ideas:

  • News anchor
  • Tourist guide
  • Sci-fi narrator
  • ASMR whisper
  • Language lab instructor

๐Ÿ–ฅ๏ธ GPU Support (Optional)

If running on CUDA (WSL + NVIDIA):

uv pip install \
  --index-url https://download.pytorch.org/whl/cu121 \
  torch torchvision torchaudio

Verify:

uv run -- python -c "import torch; print(torch.cuda.is_available())"

๐Ÿงช Development

Run script directly:

uv run -- python src/qwen_tts_cli/cli.py --text "Test" --out audio/test.wav

๐Ÿค– AI Agent Integration (Smolagents)

Use Qwen-TTS as a tool in AI agent workflows with smolagents:

Install smolagents support:

uv sync --extra smolagents

Use as a tool:

from smolagents import CodeAgent, HfApiModel
from qwen_tts_cli.smolagent_tool import QwenTTSTool

tts_tool = QwenTTSTool()
model = HfApiModel(model_id="Qwen/Qwen2.5-Coder-32B-Instruct")
agent = CodeAgent(tools=[tts_tool], model=model, add_base_tools=True)

agent.run("Generate Italian audio saying 'Benvenuto!' in a teacher's voice")

See examples/README.md for more details.


๐Ÿ“š Educational Use Cases

  • Language listening exercises
  • Pronunciation drills
  • Assignment narration
  • Accessibility audio overlays
  • LMS content generation

๐Ÿ› ๏ธ Troubleshooting

NumPy import errors

Ensure pinned version:

uv add "numpy==2.1.3"
uv sync

Transformers import issues

Pin version:

uv add "transformers==4.57.3"

Model download slow

Use smaller checkpoint:

Qwen3-TTS-0.6B-CustomVoice

๐Ÿ“œ License

See upstream model + repo licenses:

  • Qwen3-TTS
  • Hugging Face Transformers

๐Ÿ™Œ Credits

  • QwenLM
  • Hugging Face
  • Astral uv
  • Click CLI

๐Ÿ“ฃ Future Enhancements

  • Preset voice styles
  • Batch folder processing
  • Subtitle (.srt) export
  • Gradio web UI
  • LMS integration tooling

Happy synthesizing ๐ŸŽ™๏ธ

Contributors

ricklon

Issues