Hands-free audio programming assistant for Meta Ray-Ban smart glasses. Snap a photo of code on a screen, whiteboard, or paper, and hear the explanation spoken directly into your ears.
Reading code on physical whiteboards, presentation slides, or printed handouts can be difficult for developers with visual impairments or when working hands-free.
MetaHelper turns Meta Ray-Ban smart glasses into an audio coding companion. When you capture a photo of code, the app reads the syntax verbatim, identifies syntax errors or logic bugs using Gemini Vision AI, and speaks a clear explanation directly through the open-ear glasses speakers.
Backend deployment: Self-host the Dockerized Spring Boot service with Coolify, or run the included Compose definition on any Docker host.
When you take a photo with your glasses, the photo syncs to your phone. MetaHelper's mobile companion app detects the new image, sends it to the cloud vision engine, and speaks the solution through the glasses. Double-tapping the glasses stem replays the audio.
sequenceDiagram
actor User as User (glasses)
participant GW as Android ยท GalleryWatcher / iOS ยท PhotosObserver
participant GM as Android/iOS ยท GlassesManager
participant API as Android/iOS ยท ApiClient
participant BE as Backend ยท Spring Boot (Java)
participant V as vision.py ยท Gemini
participant T as Azure Speech ยท Java SDK
participant A as audio.py ยท pydub / ffmpeg
participant AP as Android/iOS ยท AudioPlayer
User->>GW: Take photo of a coding problem
GW->>GM: New gallery photo detected (MediaStore / Photos framework)
GM->>API: Read image bytes
API->>BE: multipart POST /process-image (file)
BE->>V: Read & solve the problem (gemini-3-pro-preview)
V-->>BE: Solution text (verbatim + narrative)
BE->>T: Synthesize speech (en-US-GuyNeural)
T-->>BE: MP3 audio
BE->>A: Scale playback gain
A-->>BE: Quieted MP3
BE-->>API: 200 ยท audio/mpeg (MP3 bytes)
API->>AP: Hand off audio
AP-->>User: Speak the solution
User->>AP: Double-tap glasses to replay
Capture note:Photo capture currently works throughgallery polling โ the glasses take the photo through Meta AI natively and the app reads it from the phone gallery (
READ_MEDIA_IMAGESon Android, Photos framework on iOS). The Meta Wearables SDK's direct-capture path (StreamSessionon Android,MWDATon iOS) is stubbed/in-progress and is the intended future approach.
MetaHelper/
โโโ backend/ Java 26 ยท Spring Boot 4.1.1 โ vision โ TTS โ audio pipeline
โ โโโ src/main/java/com/metahelper/
โ โโโ controller/ImageController.java GET / and POST /process-image
โ โโโ service/ Gemini, Azure Speech, and ffmpeg pipeline
โโโ shared/ Kotlin Multiplatform โ shared business logic
โ โโโ src/
โ โโโ commonMain/kotlin/com/metahelper/shared/
โ โ โโโ GlassesManager.kt Core flow coordinator
โ โ โโโ ApiClient.kt Platform-agnostic HTTP client
โ โ โโโ GalleryWatcher.kt expect/actual for photo detection
โ โ โโโ AudioPlayer.kt expect/actual for audio playback
โ โ โโโ VolumeController.kt expect/actual for volume control
โ โ โโโ WearablesConnectionMonitor.kt expect/actual for SDK connection
โ โโโ androidMain/... Android implementations (MediaStore, MediaPlayer, etc.)
โ โโโ iosMain/... iOS implementations (Photos, AVFoundation, MWDAT)
โโโ android/ Kotlin ยท Jetpack Compose โ Meta Wearables SDK client
โ โโโ app/src/main/kotlin/com/metahelper/app/
โ โโโ GalleryWatcher.kt Detects new glasses photos via MediaStore
โ โโโ GlassesManager.kt Reads photo bytes, drives the flow
โ โโโ ApiClient.kt multipart POST /process-image
โ โโโ AudioPlayer.kt Plays the returned MP3 (double-tap to replay)
โโโ iosApp/ iOS ยท Compose Multiplatform โ shared UI + iOS platform code
โโโ assets/ Shared brand assets (banner, logo) referenced by the README
Requires Java 26, Gradle 9.7.1, and ffmpeg (used for audio export). Azure Speech is used for text-to-speech.
cd backend
./gradlew bootRunCopy backend/.env.exampletobackend/.env and fill in your values:
| Variable | Required | Default | Purpose |
|---|---|---|---|
GOOGLE_API_KEY |
yes | โ | Google Gemini API key (create one) |
AZURE_SPEECH_KEY |
yes | โ | Azure Speech resource key |
AZURE_SPEECH_REGION |
yes | โ | Azure Speech resource region, such as westus2 |
AZURE_SPEECH_VOICE |
no | en-US-GuyNeural |
Azure neural voice |
AUDIO_AMPLITUDE_MULTIPLIER |
no | 0.1 |
Playback gain (0.0โ1.0); lower keeps audio from overpowering the glasses' speakers |
| Method | Route | Body | Returns |
|---|---|---|---|
GET |
/ |
โ | JSON health check |
POST |
/process-image |
multipart form, fieldfile(image) |
audio/mpeg MP3 bytes |
cd backend
./gradlew test(Tests live in backend/src/test/.)
The Mac checkout is source-only. Android lint, tests, APK packaging, and emulator checks run in the x86_64 GitHub Actions Android job. The job uses Java 21, API 34, and an x86_64 emulator, then publishes a seven-day meta-helper-debug-apk artifact.
Android builds are not run on the Mac or the ARM64 verifier. Use the workflow artifact for manual device testing.
Meta Wearables SDK access (required).The app depends on the Meta Wearables SDK (com.meta.wearable:mwdat-core/mwdat-camera 0.3.0), which is published toGitHub Packagesat https://maven.pkg.github.com/facebook/meta-wearables-dat-android. GitHub Packages requires authentication even for read access, so you must supply aGitHub Personal Access Token with the read:packages scope or Gradle cannot resolve the SDK and the build will fail.
The workflow supplies the existing GH_PACKAGES_TOKEN repository secret as GITHUB_TOKEN. Do not create a Mac local.properties file or place the token in the source checkout.
Point the app's ApiClient at your backend. Android accepts the URL through
the metahelper.backend.url Gradle property; iOS reads MetaHelperBackendURL
from Info.plist. See the self-hosted deployment guide.
The backend ships with a Dockerfile (Java 26, ffmpeg baked in):
docker build -t metahelper-backend ./backend
docker run -p 8080:8080 --env-file backend/.env metahelper-backendFor the hosted deployment path, use the Coolify instructions. The backend image is published to GHCR after the main CI workflow succeeds, and Coolify can redeploy it through a repository webhook.
MetaHelper is released under the MIT License. See LICENSE for the full text.