Self-hosted, real-time speech-to-text for the Omi wearable — powered by lightning-whisper-mlx (Apple Silicon GPU) with named speaker identification, direct Omi memory creation, and a live transcript dashboard.
No Omi subscription required. Runs on Mac Apple Silicon, Raspberry Pi 5, or any Linux/CUDA machine.
Omi's cloud STT costs money and sends your audio to third-party servers. This project replaces it entirely — your audio never leaves your machine, speaker names are resolved locally, and conversations are pushed directly to your Omi memory via the official API, bypassing the 2-minute conversation processing delay.
| Feature | |
|---|---|
| 🎙️ | Real-time WebSocket STT — streams from Omi pendant with ~2–3s lag |
| 👤 | Named speaker identification — "Rahul" instead of SPEAKER_00 |
| ⚡ | Fast speaker ID — whole-utterance embed (~0.03s) vs Sortformer diarization (~1–3s) |
| 🧠 | Direct Omi memory creation — POSTs transcript to Omi API after 30s silence, no 2-min wait |
| 🚫 | TV/movie voice filter — blocks ambient entertainment audio from being saved as memories |
| 📱 | Live transcript UI — real-time dashboard at /ui/live with speaker colour coding |
| 🔊 | Voice enrollment UI — listen to unknown clips, assign names in one click |
| 🏋️ | Benchmark tool — per-stage RTF measurement, hardware comparison |
| 🔔 | Push notifications — ntfy.sh alerts for pipeline events |
sequenceDiagram
participant P as Omi Pendant
participant CF as Cloudflare Tunnel
participant S as omi-whisperx
participant O as Omi Cloud API
P->>CF: PCM16 audio frames (WebSocket)
CF->>S: wss://your-domain.com/live
Note over S: VAD Gate — flush on 0.5s silence
S->>S: lightning-whisper-mlx transcribe (~2s, Apple GPU)
S->>S: Fast speaker ID — whole-utterance embed (~0.03s)
S-->>P: {"segments": [...]} → shown in app
Note over S: 30s silence timer
S->>O: POST /conversations/from-segments
O-->>S: {"id": "conv-uuid"}
Note over O: AI processing: title, action items, memory
Live path latency (~2–3s end-to-end). The WebSocket path is tuned for low lag: per-utterance speaker ID uses a single whole-utterance embedding (
~0.03s) instead of full Sortformer diarization + per-speaker embedding (~1–3s) — pendant utterances are near-always single-speaker. Word-level alignment is skipped (SKIP_LIVE_ALIGN), and pinning the language (WHISPER_DEFAULT_LANG=hi) skips Whisper's per-utterance autodetect for a further ~30%. SetLIVE_FAST_SPEAKER_ID=falseto restore in-utterance Sortformer diarization (multi-speaker per utterance, higher lag).
flowchart LR
A[Audio chunk] --> B[lightning-whisper-mlx<br/>transcribe]
B --> X{LIVE_FAST_SPEAKER_ID?}
X -->|true · live| C[Resemblyzer<br/>embed whole utterance<br/>~0.03s]
X -->|false| Y[Sortformer diarize<br/>+ per-speaker embed<br/>~1–3s]
Y --> C
C --> D{Cosine similarity<br/>vs enrolled profiles}
D -->|above threshold| E[Named speaker<br/>Rahul Sharma]
D -->|below threshold| F{Blocked voice?}
F -->|Yes| G[Discard<br/>silently]
F -->|No| H[UNKNOWN<br/>Save clip for review]
H --> I[ui/speakers<br/>Assign name]
I --> J[Enroll<br/>average embedding]
sequenceDiagram
participant App as Omi iOS App
participant S as omi-whisperx
participant OB as Omi Backend
Note over App,OB: Standard path (built-in STT)
App->>OB: Audio via Omi backend WS
OB->>OB: Deepgram STT
OB-->>App: Segments (real-time)
Note over OB: Wait conversation_timeout (120s)
OB->>OB: LLM process → memory created
Note over App,OB: omi-whisperx path
App->>S: Audio via custom STT WS
S-->>App: Segments (real-time, ~3s lag)
Note over S: 30s silence debounce
S->>OB: POST /v1/dev/user/conversations/from-segments
OB->>OB: LLM process → memory created
Note over S,OB: Total delay: ~35s vs ~120s
- Python 3.12
ffmpeg+libsndfile1- HuggingFace token — accept pyannote/speaker-diarization-3.1 license
- Omi API key — from Omi Developer Settings
git clone https://github.com/Rahulsharma0810/omi-whisperx
cd omi-whisperx
cp .env.example .env
# Edit .env — set HF_TOKEN and OMI_API_KEY at minimum./start.shCreates ~/.venvs/whisperx, installs deps, launches on :8080. First start downloads models (~2–4 GB).
Omi pendant needs HTTPS. Use Cloudflare Tunnel:
cloudflared tunnel run --url http://localhost:8080In Omi → Settings → Developer → Cloud Provider > Custom:
| Field | Value |
|---|---|
| WebSocket URL | wss://your-domain.com/live |
| Sample Rate | 16000 |
| Language | en (or leave blank for auto-detect) |
Enable VAD Gate in Omi app settings for best performance (strips silence before sending).
| URL | Description |
|---|---|
/ui/live |
Live transcript — real-time segments with speaker colours, lag metrics |
/ui |
Pipeline monitor — SSE stream of every inference stage |
/ui/speakers |
Speaker manager — listen to unknown clips, assign names, enroll, block |
/benchmark |
Benchmark — per-stage RTF, hardware comparison |
/health |
System status — model, devices, speaker count, config |
Segments stream in real time with speaker name, colour coding, and lag metrics:
audio@+2.3s → sent@+5.1s | lag=2.8s (queue=0.0s proc=2.8s)
Three ways to teach the system who's speaking:
Every unknown voice is saved as a short clip (max 5 per unique voice). Go to /ui/speakers, play each clip, type a name, click Confirm. Done.
curl -X POST http://localhost:8080/speakers \
-F "name=Alice" \
-F "file=@alice_voice.m4a"Open /ui/speakers → click the mic icon next to any name → speak for 5s → auto-enrolled.
Unknown voices from TV/movies appear in clips. Click Block on any clip → that voice is permanently ignored and never recorded again. Uses embedding similarity so it works even if the audio quality changes.
All settings via environment variables. See .env.example for the full list.
| Variable | Default | Description |
|---|---|---|
WHISPER_MODEL |
small |
tiny base small medium large-v2 |
WHISPER_BATCH_SIZE |
16 |
Lower to 4 on Raspberry Pi 5 |
WHISPER_DEFAULT_LANG |
— | Pin language when client sends none (e.g. hi). Skips per-utterance autodetect (~30% faster); empty = autodetect |
HF_TOKEN |
— | Required — HuggingFace token |
| Variable | Default | Description |
|---|---|---|
FAST_SPEAKER |
true |
true = resemblyzer (~0.1s), false = pyannote (~15s) |
SPEAKER_THRESHOLD |
0.85 |
Cosine similarity cutoff. Higher = stricter matching |
PROFILES_DIR |
~/.omi/speakers |
Speaker embedding storage |
RECORDINGS_MAX_AGE_DAYS |
7 |
Auto-expire unknown clips after N days |
| Variable | Default | Description |
|---|---|---|
TRUST_CLIENT_VAD |
true |
Trust Omi VAD Gate — skip server-side VAD, flush on 0.5s frame gap |
SKIP_LIVE_ALIGN |
true |
Skip word-level alignment in WS path (saves ~2s, not needed for Omi) |
LIVE_FAST_SPEAKER_ID |
true |
true = single whole-utterance embed (~0.03s); false = Sortformer diarize + per-speaker embed (~1–3s, multi-speaker per utterance) |
MAX_QUEUE_AGE |
30 |
Drop queued utterances older than N seconds |
| Variable | Default | Description |
|---|---|---|
OMI_API_KEY |
— | Omi developer API key — enables direct conversation creation |
OMI_USER_NAME |
— | Your name as enrolled speaker — marks your segments is_user=true |
OMI_CONV_DEBOUNCE |
30 |
Seconds of silence before POSTing conversation to Omi API |
OMI_API_BASE |
https://api.omi.me/v1/dev/user |
Omi API base URL |
| Variable | Default | Description |
|---|---|---|
CONTENT_FILTER |
false |
Enable two-tier entertainment filter |
NLI_ENABLED |
false |
Tier 1: zero-shot NLI classifier |
NLI_THRESHOLD |
0.85 |
Confidence cutoff — below escalates to Ollama |
OLLAMA_ENABLED |
false |
Tier 2: Ollama LLM fallback |
OLLAMA_URL |
http://localhost:11434 |
Ollama server |
OLLAMA_MODEL |
deepseek-v3.1:671b-cloud |
Ollama model name |
Real-time STT for Omi pendant.
Query params: language, uid, sample_rate (default 16000), codec (default opus)
Omi sends: Binary PCM16 frames (640 bytes = 20ms at 16kHz) + JSON control messages ({"type": "CloseStream"})
Server sends:
{
"segments": [
{
"text": "Hello world",
"speaker": "Rahul Sharma",
"start": 0.0,
"end": 2.4
}
]
}HTTP transcription endpoint (Omi Transcript Provider mode).
curl -X POST http://localhost:8080/inference \
-F "[email protected]" \
-F "language=en" \
-F "response_format=verbose_json"Response:
{
"segments": [{"start": 0.0, "end": 2.4, "text": "Hello", "speaker": "Rahul"}],
"text": "Hello",
"language": "en",
"duration": 5.2
}| Method | Path | Description |
|---|---|---|
GET |
/speakers |
List enrolled speakers |
POST |
/speakers |
Enroll — name + file form fields |
PATCH |
/speakers/{name} |
Rename |
DELETE |
/speakers/{name} |
Delete profile |
GET |
/speakers/recordings |
List unknown clips with similarity scores |
POST |
/speakers/recordings/{id}/assign |
Assign name → enroll |
POST |
/speakers/recordings/{id}/block |
Block voice permanently |
DELETE |
/speakers/recordings/{id} |
Discard clip |
POST |
/speakers/recordings/purge |
Delete clips that now match enrolled speakers |
GET/DELETE |
/speakers/blocked |
View / clear blocked voices |
| Event | When | Key fields |
|---|---|---|
ws_connected |
Pendant connects | uid, lang |
ws_audio |
First audio frame | frame_bytes |
ws_transcript |
Utterance processed | segments, lag |
ws_disconnected |
Pendant disconnects | frames |
audio_received |
HTTP inference start | chunk_id, size_kb |
transcribed |
After WhisperX | chunk_id, lang, segments |
transcript |
Inference complete | chunk_id, duration, segments |
sudo apt-get install -y python3.12 python3.12-venv ffmpeg libsndfile1
python3.12 -m venv ~/.venvs/whisperx
source ~/.venvs/whisperx/bin/activate
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu
pip install -r requirements.txt
# Recommended .env for RPi5
WHISPER_MODEL=base
WHISPER_BATCH_SIZE=4
FAST_SPEAKER=true
TRUST_CLIENT_VAD=true
SKIP_LIVE_ALIGN=trueExpect ~2–3s per utterance with base model on RPi5.
| Hardware | Whisper device | Diarization | Notes |
|---|---|---|---|
| Mac Apple Silicon | CPU + int8 | MPS | Best dev setup |
| Linux + CUDA | CUDA + float16 | CUDA | Fastest inference |
| Raspberry Pi 5 | CPU + int8 | CPU | Use base/tiny model |
| Linux CPU-only | CPU + int8 | CPU | Use base/tiny model |
CTranslate2 (Whisper backend) and resemblyzer (speaker encoder) do not support MPS — CPU is used intentionally.
# Quick (no HF_TOKEN needed)
python benchmark.py --no-diarization --no-embedding --language en --trials 3
# Full pipeline on real audio
python benchmark.py audio.wav --trials 3 --output json --output-file results_mac.json
# Compare two machines
python benchmark.py --compare results_mac.json results_rpi5.jsonOr use the interactive UI at /benchmark.
# Install
cp com.rvs.whisperX.plist ~/Library/LaunchAgents/
launchctl load ~/Library/LaunchAgents/com.rvs.whisperX.plist
# Restart after config changes
launchctl unload ~/Library/LaunchAgents/com.rvs.whisperX.plist
launchctl load ~/Library/LaunchAgents/com.rvs.whisperX.plist
# Logs
tail -f ~/omi-whisperx/server.logPRs and issues welcome.
Do not:
- Add MPS to WhisperX calls — CTranslate2 doesn't support it
- Add MPS to
VoiceEncoder— resemblyzer doesn't support it - Add pip deps without checking
aarch64wheels exist (must run on RPi5) - Make
classify_content()synchronous — Ollama call is async HTTP - Call NLI pipeline directly from async code without
asyncio.to_thread
MIT