A fully local, offline first speech-to-text application made for desktop environments running Wayland compositor (DE's like hyprland etc...).
Records audio when a hotkey is pressed, transcribes it using faster-whisper, and automatically types the transcribed text.
This is in beta, since I've been testing only on my system. Please file bug reports if something does not work. Thanks!
Demo done on omarchy running hyprland
demo.mp4
- Support for even faster transcription local methods like nvidia parakeet
- Custom vocabulary support
- Use LLM's like ChatGPT to auto format text before pasting
We'll expand compatibility in the coming days.
- Window manager: Wayland
- Python: 3.8+
- Package manager: UV
uv tool install speechshiftRun test to make sure pipewire, wl clipboard is present. It also downloads the whisper (small - ~80mb) model for transcription.
speechshift --testAdd these lines to your ~/.config/hypr/hyprland.conf:
The recommended default is Super+Shift+R, but you can set it to anything you like
# SpeechShift POC Keybinds
bind = SUPER_SHIFT, R, exec, /path/to/speechshift --toggleand setup speechshift daemon to startup on default by adding these lines to ~/.config/hypr/hyprland.conf
exec-once = bash -c 'source /path/to/.bashrc && /path/to/speechshift --daemon'we'll need to source the right file, which has the ASSEMBLYAI_API_KEY if assembly ai is being used
Then either restart, so that the deamon is automatically run. Or start running the speechshift deamon manually for this session by running
speechshift --deamon- Start recording (Super+Shift+R): You'll see a notification: "๐ค Recording started..."
- Stop recording (Super+Shift+R): Audio is automatically transcribed using faster-whisper or AssemblyAI. Transcribed text is typed into the focused window. Notifications show: "๐ Transcribing audio..." โ "โ Transcribed: [preview]"
SpeechShift can be configured by creating a config.json file in ~/.config/speechshift/. If the file doesn't exist, it will be created with default settings upon first run.
Here's an example configuration to override the default whisper model and language:
{
"transcription": {
"engine": "whisper"
},
"whisper": {
"model": "medium",
"language": "en"
},
"audio": {
"recording_device": null,
"notification_timeout": 3000
}
}to use Assembly AI, make sure to set the ASSEMBLYAI_API_KEY environment variable and set the transcription engine to assemblyai.
Assembly AI is highly recommended since its much better on accuracy & speed.
Keybind (Super+Shift+R)
โ
Main Python Script
โโโ PipeWire Audio Recording (sounddevice)
โโโ AI Transcription (faster-whisper)
โโโ Temporary File Management
โโโ Wayland Text Input (wl-clipboard + hyprctl simulate control+V)
โโโ Smart Notifications (notify-send)
- Keybind Press: Hyprland detects Super+Shift+R press
- Recording Start:
- Python script starts PipeWire audio capture
- Notification: "๐ค Recording started..."
- Audio streams to temporary WAV file in /tmp
- Keybind Release: Hyprland detects key release
- Recording Stop & Transcription:
- Audio capture stops
- Notification: "๐ Transcribing audio..."
- faster-whisper transcribes the audio
- Transcribed text pasted into active window
- Temporary file automatically deleted
- Success notification: "โ Transcribed: [preview]"
- Audio Format: 16-bit WAV, 44.1kHz, mono
- Transcription Model: faster-whisper "base" model (configurable)
- File Handling: Temporary files in
/tmp, auto-cleanup after transcription - Text Insertion: Direct typing via wtype, fallback to clipboard paste
- Notifications: Smart status updates via notify-send
- Error Handling: Graceful fallback with error notifications
-
"sounddevice not available":
# Install manually: pip install --user sounddevice numpy -
"Audio recording failed":
- Check PipeWire is running:
systemctl --user status pipewire - Test microphone:
pw-record --list-targets - Verify permissions: ensure user is in
audiogroup
- Check PipeWire is running:
-
"Hyprland socket not found":
- Ensure running under Hyprland
- Check environment variables:
echo $HYPRLAND_INSTANCE_SIGNATURE
-
"Text insertion not working":
- Verify wtype is installed:
wtype --version - Test manually:
wtype "test" - Check focused window accepts text input
- Verify wtype is installed:
-
"Notifications not showing":
- Test manually:
notify-send "test" "message"
- Test manually:
Enable detailed logging by checking ~/.speechshift.log:
tail -f ~/.speechshift.log