- Global hotkeys for voice recording
- Local Whisper transcription with fallback to OpenAI
- Multi-language support
- Screenshot OCR
- Offline processing using local Whisper models
- Clone the repository:
git clone <repository-url>
cd tool- Set up Whisper models:
# Create models directory
mkdir models
# Download Whisper models (choose one or more):
# Recommended base or small models for best performance
curl -L "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.bin" -o models/model.bin
# OR for smaller size:
curl -L "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin" -o models/model.bin- (Optional) Set up OpenAI API key for fallback:
export OPENAI_API_KEY="your-api-key-here"-
(Optional) Set up LM Studio for local AI processing:
- Download and install LM Studio
- Load a model in LM Studio
- Start the local server (default: http://localhost:1234)
-
Build the project:
swift build- Run the application:
swift run| Hotkey | Function | Description |
|---|---|---|
โโฅโM |
LM Studio Assistant | Record audio โ AI assistant response with screen context |
โโฅโN |
LM Studio Translator | Record audio โ translate Russian to English |
โโฅโV |
OpenAI Assistant | Record audio โ OpenAI assistant response with screen context |
โโฅโB |
Screenshot OCR | Select screen area โ extract text to clipboard |
- Press and hold any audio hotkey to start recording
- Speak into the microphone
- Release hotkey to stop recording and process
- Result appears in tooltip and copies to clipboard
- Press
โโฅโBfor screenshot OCR - Select area on screen with cursor
- Text is extracted and copied to clipboard
- The app automatically uses the first
.binmodel found in themodels/directory - Recommended models for different use cases:
ggml-base.bin(~140MB) - good balance of speed/qualityggml-small.bin(~460MB) - better qualityggml-large-v3-turbo.bin(~1.5GB) - best quality, faster than large-v3ggml-tiny.bin(~39MB) - fastest, lower quality
- Download from lmstudio.ai
- Load any compatible model (e.g., Llama, Mistral, etc.)
- Start local server in LM Studio
- Provides local AI processing for enhanced features
- Get API key from platform.openai.com
- Set environment variable:
export OPENAI_API_KEY="sk-..."- Used automatically if local Whisper fails
On first launch, macOS will request permissions:
- Microphone access - for voice recording
- Accessibility features - for global hotkeys