legb78/hex_windows

Local voice dictation for Windows — what Hex is on macOS. Hold a hotkey, speak, release: the text lands at your cursor. Parakeet TDT v3, offline, ~0.2 s.

★ 5Forks 0C#GitHub ↗Compare

Project website ↗

asrcsharpdictationdotnetlocal-firstofflineonnxruntimeparakeetprivacypush-to-talksherpa-onnxspeech-recognitionspeech-to-textvoice-typingwhisper-alternativewindows

README

HexWin

Local voice dictation for Windows. Hold right Shift, speak, release — the text lands at your cursor.

Everything runs on your machine. No cloud service, no subscription, no network connection needed once the model is downloaded.

Hex, but for Windows

Hex is a hold-to-talk dictation app for macOS, and a good one. It has never run on Windows, and the question comes up often enough that this exists to answer it.

HexWin is not a port. The macOS original is Swift on CoreAudio; this is C# on .NET 9, WASAPI and a Win32 keyboard hook, written from scratch for the platform. What the two share is the shape of the thing and the engine underneath. If you came looking for a Windows alternative to Hex, Wispr Flow or SuperWhisper: this is free, open source, offline, and asks for no account.

legb78.github.io/hex_windows — the short version, if you would rather see it than read this.

Architecture

How fast it is

Measured on a Dell Pro Max 16 (Core Ultra 7 255H), on CPU:

Speech length Wait after release
0.6 s 0.05 s
1.6 s 0.08 s
5 s 0.19 s
40 s 1.58 s

The engine is Parakeet TDT v3 by NVIDIA — the same one Hex uses on macOS. It identifies the spoken language on its own among 25 European languages, so the language you dictate in is not a setting, and cannot be set wrong. The one language setting, further down, is for HexWin's own windows and menus.

Whisper was tried first here, and dropped after measurement. It is an autoregressive encoder-decoder, which imposes a fixed cost of about 1.4 s per transcription regardless of how short the audio is — and on a one-sentence dictation, the normal case, that cost is paid in full. Parakeet is a transducer and has no such bottleneck. Same recordings, Parakeet on CPU against Whisper on GPU: 2.29 s to 0.19 s on a five-second clip.

Install

If you would rather not compile anything

Download the archive from the latest release, unzip it, then in PowerShell:

.\get-model.ps1     # about 480 MB, once
.\HexWin.exe

Nothing else to install — not even .NET, which is bundled inside the executable.

HexWin has no installer, so the first start is where it asks, once, whether to put a shortcut on the desktop. That shortcut starts HexWin; double-clicked while HexWin is already running, it opens the settings window instead. It can be added or removed later from the settings, on the general page.

The shortcut points at the executable where it stood when the shortcut was made: move the HexWin folder and it breaks. Start HexWin from its new place, open the settings and save with the shortcut switch on: the shortcut is written anew on every save, for the new place.

Windows will warn you the first time. The executable is not code-signed, so SmartScreen shows "Windows protected your PC". Click More info, then Run anyway. That is expected for an unsigned binary downloaded from the internet; the alternative is a paid signing certificate.

Some antivirus products flag the app too, because it installs a global keyboard hook. That hook is what makes push-to-talk work at all: HexWin has to see your hotkey whichever window has focus. It looks at nothing else and stores no keystrokes. The code responsible is KeyboardHook.cs — short, and worth reading if you would rather check than trust.

If you compile it yourself

You need the .NET 9 SDK (https://dotnet.microsoft.com/download).

.\scripts\get-model.ps1     # downloads the model
dotnet build -c Release
.\src\HexWin\bin\x64\Release\net9.0-windows\HexWin.exe

Using it

The tray icon shows the current state:

Colour State
Grey Loading the model, a few seconds
Blue Ready
Red Recording
Orange Transcribing
Crossed grey Model not found

Hold right Shift, speak, release. The text arrives at your cursor.

For long dictations, turn on sentence by sentence in the settings (off by default): each pause in your speech then closes a piece, transcribed and inserted while you keep talking, instead of everything arriving at release. A dictation cancelled halfway — another key pressed while the hotkey is held — keeps the sentences already inserted.

A circle appears at the top of the screen while the application is busy, in the same colours as the tray icon — red while recording, orange while transcribing, nothing at rest — and a short tone marks each end of the recording. The tray is the wrong place to look while you are watching your own text; this is the same signal, where the eye already is. Both are set by feedback below, and either can be turned off on its own.

Right-clicking the icon opens a menu whose first entry, in bold, is the settings window; the others open the settings file and the log folder, and offer a start with Windows toggle. Worth enabling: the app does not come back on its own after a reboot otherwise. Double-clicking the icon opens the settings window directly, and so does launching HexWin.exe — or its desktop shortcut — while HexWin is already running.

The same menu carries the two cues as independent switches — the circle during dictation and the tone at each end. Both take effect on the next dictation, with no restart, and are written back to settings.json as the feedback value below.

The settings window

Every setting below except provider, which has only one possible value, in four pages — without having to know a single key name. Dictation holds the hotkey, the minimum and maximum recording durations, paste or type, and sentence by sentence with its pause; cues, the circle and the tone, and the circle's colour, size, opacity and top margin; engine, the model folder, the threads and the idle unload; general, start with Windows, the desktop shortcut, the dictation log, the display language, and two buttons that open settings.json and the log folder. It follows the Windows light or dark theme, and takes the Windows 11 frame (round corners, Mica title bar) where the system offers it.

  • The hotkey is captured, not typed: click the button beside it, hold the keys you want, release. The window names the side of each key — Right Shift, not just Shift — and warns when a single key costs you something for typing. While it listens, the whole keyboard goes to the capture; Escape or a click elsewhere hands it back.
  • Save writes only what changed, one line per setting, so the comments in settings.json survive. A setting missing from an older file is added.
  • What can change on the fly does: the hotkey, the insertion mode, sentence by sentence and its pause, the circle and the tones apply on the next dictation. The engine settings, the recording durations and the log are read at startup; the window says which ones wait for a restart.
  • Cancel changes nothing, in the file or in the running app.
  • It speaks French or English, following the Windows display language unless language says otherwise: the window, the tray menu and its tooltip, the balloons and dialogs, the key names — Maj droite or Right Shift. Not the dictation, whose language Parakeet works out on its own, and not the log.

Settings

Everything lives in settings.json, next to the executable. The settings window edits it for you; editing it by hand still works. Comments are allowed, and an invalid value falls back to its default rather than preventing startup.

Setting What it does
hotkey Keys to hold. Fn cannot be used: it is handled by the keyboard controller and emits no code Windows can see.
insertion Paste (clipboard, instant) or Type (simulated keystrokes, for apps that ignore pasting).
feedback What marks the start and end of a recording: Both (default), Visual (circle only), Sound (tones only), None.
feedbackColor auto to follow the tray icon colours, or #RRGGBB to force one.
feedbackSize Diameter of the circle in pixels, 16 to 512.
feedbackOpacity Opacity of the circle, 20 to 255.
feedbackTopMargin Pixels between the circle and the top of the screen.
modelPath Folder holding the Parakeet model, relative to the executable.
provider cpu, the only one available: the published native libraries are built for CPU only.
threads Threads allocated to decoding.
minRecordingMilliseconds Below this, the keypress is treated as accidental.
maxRecordingSeconds Stops recording if the key stays held.
segmentation true inserts a long dictation sentence by sentence, while you keep talking. false (default) inserts everything at release.
pauseMilliseconds With segmentation on, a pause this long closes a piece of the dictation: 700 (default), 100 ms to 5 s in the window. Kept while segmentation is off. Detected on the sound level: in a noisy room no pause is seen and the text simply arrives at release.
unloadAfterMinutes Frees the model after this long without dictating, reclaiming about 1 GB. 0 keeps it resident. Reloading starts when you press the hotkey, so it overlaps with you speaking.
logEnabled true (default) logs every dictation — duration, captured level, characters produced. See below.
language Language of HexWin's own interface: auto follows the Windows display language, fr and en force one. The dictation language is not affected, nor the log.

When something goes wrong

Five diagnostic modes. Run the first four in this order — each isolates one layer, which is how you find out where the fault actually is instead of guessing. The fifth stands alone: it needs neither model nor microphone.

# 1. Is the microphone picking anything up?
.\HexWin.exe --record test.wav

# 2. Does the engine transcribe that file?
.\HexWin.exe --transcribe test.wav

# 3. Does the hotkey fire?
.\HexWin.exe --watch-hotkey

# 4. Does insertion reach the target window?
.\HexWin.exe --inject "some text"

# 5. Do the start and end cues show and sound?
.\HexWin.exe --test-feedback

Start with the first. A muted microphone, or one blocked by the privacy settings, produces a perfectly valid file of the right duration that is completely silent — and the engine then invents a sentence from it. Without the level reading you would go looking for the fault in the transcription, which is the wrong end entirely.

The log lives in %LOCALAPPDATA%\HexWin and records, for every dictation, the duration, the captured level and the number of characters produced. Those three numbers together tell you whether the microphone heard you, whether the engine understood you, and whether the text made it out.

Known limits

  • The right Shift no longer shifts while HexWin runs: the hook swallows the key, so Windows never sees it. Use the left one for capitals. It behaves normally again once HexWin is stopped, and briefly while the watchdog reinstalls the hook after a long silence. Pick another key in the settings window to get it back.
  • AltGr is captured as two keys on layouts that have it, French included: Windows sends a left Ctrl along with the right Alt, and the capture records both. Pick another key, or edit hotkey by hand to ["RightAlt"].
  • Elevated windows: a non-elevated app cannot send keystrokes to a window running as administrator (Windows UIPI isolation). Run HexWin elevated if you need to dictate into one.
  • Antivirus: a global keyboard hook is a flagged pattern. An exclusion on the executable may be needed.
  • Unsigned binary: see the SmartScreen note above.

Development

dotnet build -c Release     # no warnings tolerated
dotnet test                 # unit tests

dotnet test runs everything. The integration tests load the real engine, so they skip themselves with a message when the model has not been downloaded yet — a fresh clone gives you a green run and a count of what was skipped, rather than failures that say nothing about the code.

Download the model and they run for real. CI excludes them up front, since a runner has no reason to spend minutes discovering they would skip.

How the code is organised

The architecture deliberately separates two layers, for testability (see the diagram above; editable source in docs/architecture.excalidraw):

  • The Windows shells (KeyboardHook, AudioRecorder, ParakeetEngine, TextInjector, RecordingOverlay, CueTones) wire up system APIs and decide nothing. They cannot be tested automatically — a CI runner has no microphone and no interactive session, and Windows marks program-generated keystrokes as injected, which the hook ignores on purpose. They are verified by hand.
  • The pure logic (ChordDetector, RecordingGuards, TranscriptCleaner, AppSettings, DictationCoordinator, IdlePolicy, FeedbackPolicy) holds every decision and is tested without Windows. This is where the expensive bugs live: keyboard auto-repeat, keys released out of order, a key left stuck after a session lock, a second dictation triggered while one is still transcribing.

If you add a decision, it belongs in the pure layer.

Contributing

Two long-lived branches:

Branch Role
main Published releases. Only ever advances from develop.
develop Integration. Work branches are merged here.

One branch per change, created from develop and merged back into it: feat/, fix/, chore/, ci/, docs/. Messages follow Conventional Commits, squash-merged.

git checkout develop
git pull
git checkout -b feat/my-topic
gh pr create --base develop

Full details in CONTRIBUTING.md.

Publishing a release

.\scripts\publish.ps1        # builds the self-contained executable locally

Directory.Build.props holds <Version> and is the single source of truth. Merging into main publishes that version, and does nothing if the tag already exists — so releasing means bumping the number in a pull request.

Licence

HexWin is released under the Apache 2.0 licence.

The recognition engine is Parakeet TDT 0.6B v3 by NVIDIA, under CC-BY-4.0: commercial use permitted, attribution required. The other components (sherpa-onnx, ONNX Runtime, NAudio, .NET) are Apache 2.0 or MIT.

Attributions in full: NOTICE.

Contributors

legb78dependabot[bot]weldhammadi

Issues