jaredh159/dictator

transcription helper

β˜… 0Forks 0SwiftGitHub β†—Compare

README

πŸ’‚ dictator

macos menu bar app for voice dictation. records your voice, transcribes with openai whisper, cleans up the text with gpt-4o-mini, and copies to clipboard.

installation

  1. clone the repo
  2. open Dictator.xcodeproj in xcode
  3. select your signing team in Signing & Capabilities (your personal team works fine)
  4. build and run

setup

create two config files in ~/.config/dictator/:

dictator.secrets.toml (your api key):

openai_api_key = "sk-your-key-here"

dictator.personalities.toml (your personalities):

[[personality]]
name = "Claude"
hotkey = "cmd-shift-k"
prompt = """
You are a light text editor. You receive raw transcriptions from a seasoned
software engineer that were converted from audio to text. Your job is to do
a very light cleanup...
"""

[[personality]]
name = "Slack"
hotkey = "cmd-shift-s"
prompt = """
Another cleanup prompt here...
"""

each personality has:

  • name: displayed in the window while recording
  • hotkey: unique hotkey to trigger this personality (e.g. "cmd-shift-k", "ctrl-opt-d")
  • prompt: the system prompt sent to gpt-4o-mini for text cleanup

for dotfiles users: the personalities file can be stowed from ~/.dotfiles/dictator/.config/dictator/dictator.personalities.toml

usage

  • press a personality's hotkey to summon the window and start recording
  • press the same hotkey again to stop recording
  • transcription happens automatically
  • cleaned text is copied to your clipboard
  • window dismisses when done

the app runs in your menu bar. doesn't steal focus from your current app.

voice command mode

an always-on, on-device (free, no api) speech listener that matches short phrases and runs shell commands β€” built for driving tmux hands-free. configured in ~/.config/dictator/dictator.commands.toml:

always_on = true                # start listening at app launch
toggle_hotkey = "cmd-shift-m"   # hotkey to toggle listening on/off

sleep_phrases = ["go to sleep"] # voice-pause: only "wake up" matches after this
wake_phrases = ["wake up"]

[dictation]
start_phrases = ["record", "start recording"]
stop_phrases = ["stop recording", "finish recording"]
spokenly_shortcut = "cmd-shift-o"  # spokenly's global toggle shortcut

[[command]]
phrases = ["left", "pane left"]
run = "tmux select-pane -L"

# capture commands: words spoken after the phrase are passed as $VOICE_ARG
[[command]]
phrases = ["open session"]
capture = true
run = "~/.local/scripts/tmux-sessionizer.sh \"$VOICE_ARG\""

phrases match against the end of what you say (filler before a command is fine), longest phrase wins. commands run via zsh -c with homebrew on the PATH. before each command, dictator resolves the focused tmux client and exports $TMUX_VOICE_CLIENT (tty) and $TMUX_VOICE_SESSION so commands can target the window you're looking at. capture commands fire ~1s after you stop talking, with the trailing words joined into $VOICE_ARG.

send_phrases (e.g. "send it") work like stop phrases, but once spokenly's transcript lands (detected via the clipboard), submit_command runs β€” dictate and submit to an agent in one breath.

dictation handoff (spokenly)

say a start phrase ("record") and dictator taps spokenly's global shortcut to start a spokenly recording, then ignores everything you say except the stop phrases. say "stop recording" and it taps the shortcut again β€” spokenly inserts the text at your cursor and dictator goes back to matching commands.

two things to know:

  • requires accessibility permission (System Settings β†’ Privacy & Security β†’ Accessibility β†’ Dictator) to post the synthetic keystroke.
  • spokenly will transcribe the trailing stop phrase into your text. strip it by adding a regex word replacement in spokenly: original (?i)[\s,.!?]*(?:stop|finish) recording[\s,.!?]*$ β†’ replacement: (empty).

while dictator's own personality recording is active, command matching is suspended so your prompt speech can't trigger commands. note: if you start spokenly with its keyboard shortcut directly (bypassing "record"), dictator has no way to know β€” prefer starting dictation by voice while command mode is on, or say "go to sleep" first.

menu bar icon shows the state: waveform = listening, record = dictation handoff, moon = asleep, pause = suspended, plain mic = off.

requirements

  • macos 15+
  • openai api key with access to whisper and gpt-4o-mini

microphone permissions

the app will request microphone access on first launch. if you don't see a prompt or recording isn't working:

  1. open System Settings β†’ Privacy & Security β†’ Microphone
  2. find Dictator in the list and enable it

if you previously denied permission, macos won't prompt againβ€”you'll need to enable it manually in System Settings.

Contributors

jaredh159

Issues