IlyaPomaskin/tool

โ˜… 0Forks 0SwiftGitHub โ†—Compare

README

Tool

๐Ÿš€ Features

  • Global hotkeys for voice recording
  • Local Whisper transcription with fallback to OpenAI
  • Multi-language support
  • Screenshot OCR
  • Offline processing using local Whisper models

โš™๏ธ Installation

  1. Clone the repository:
git clone <repository-url>
cd tool
  1. Set up Whisper models:
# Create models directory
mkdir models

# Download Whisper models (choose one or more):
# Recommended base or small models for best performance
curl -L "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-base.bin" -o models/model.bin
# OR for smaller size:
curl -L "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin" -o models/model.bin
  1. (Optional) Set up OpenAI API key for fallback:
export OPENAI_API_KEY="your-api-key-here"
  1. (Optional) Set up LM Studio for local AI processing:

  2. Build the project:

swift build
  1. Run the application:
swift run

๐ŸŽฏ Usage

Global Hotkeys

Hotkey Function Description
โŒƒโŒฅโŒ˜M LM Studio Assistant Record audio โ†’ AI assistant response with screen context
โŒƒโŒฅโŒ˜N LM Studio Translator Record audio โ†’ translate Russian to English
โŒƒโŒฅโŒ˜V OpenAI Assistant Record audio โ†’ OpenAI assistant response with screen context
โŒƒโŒฅโŒ˜B Screenshot OCR Select screen area โ†’ extract text to clipboard

How to Use

Audio Processing

  1. Press and hold any audio hotkey to start recording
  2. Speak into the microphone
  3. Release hotkey to stop recording and process
  4. Result appears in tooltip and copies to clipboard

Screenshot OCR

  1. Press โŒƒโŒฅโŒ˜B for screenshot OCR
  2. Select area on screen with cursor
  3. Text is extracted and copied to clipboard

๐Ÿ”ง Configuration

Whisper Models

  • The app automatically uses the first .bin model found in the models/ directory
  • Recommended models for different use cases:
    • ggml-base.bin (~140MB) - good balance of speed/quality
    • ggml-small.bin (~460MB) - better quality
    • ggml-large-v3-turbo.bin (~1.5GB) - best quality, faster than large-v3
    • ggml-tiny.bin (~39MB) - fastest, lower quality

LM Studio (Local AI)

  1. Download from lmstudio.ai
  2. Load any compatible model (e.g., Llama, Mistral, etc.)
  3. Start local server in LM Studio
  4. Provides local AI processing for enhanced features

OpenAI API (Fallback)

  1. Get API key from platform.openai.com
  2. Set environment variable:
export OPENAI_API_KEY="sk-..."
  1. Used automatically if local Whisper fails

Permissions

On first launch, macOS will request permissions:

  • Microphone access - for voice recording
  • Accessibility features - for global hotkeys

Contributors

IlyaPomaskin

Issues