joewired/claude-code-local

Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server, 65 tok/s Qwen 3.5 122B, Llama 3.3 70B, Gemma 4 31B. Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.

โ˜… 0Forks 0GitHub โ†—Compare

Project website โ†—

README

๐Ÿง โšก Claude Code Local โ€” The Lineup

Three local AI brains. Four modes. One MacBook. Zero cloud.
Pick your fighter and run Claude Code 100% on-device.
๐Ÿ“ Now with DeepSeek V4 Flash ยท 1M-token context ยท via Antirez's ds4 engine.

GitHub stars GitHub forks 3 Models 4 Modes Qwen 3.5 speed Claude Code task time 100% Local Hands-Free Voice Ambient Computing MIT Join the NiceDreamzApps Discord

Built by Matt Macosko in Arcata, CA. Started with a chicken problem. Still figuring it out.

๐ŸŽฌ Demo ยท ๐ŸฅŠ Lineup ยท ๐ŸŽฎ Modes ยท ๐Ÿš€ Quick Start ยท ๐Ÿ“Š Benchmarks ยท ๐Ÿ”’ Safety ยท ๐ŸŽค Voice ยท ๐Ÿงฉ The Stack ยท ๐Ÿ›ฃ๏ธ Roadmap ยท ๐Ÿค Contribute


๐ŸŽฌ WATCH THE DEMO โ€” AirGap AI

A real NDA. Llama 3.3 70B. Wi-Fi physically OFF. lsof running live.
Watch a 70-billion-parameter model audit a confidential legal document, on-device, with the receipts on screen.

AirGap AI โ€” Wi-Fi OFF NDA Demo

Watch on YouTube ย  Subscribe

Built for lawyers, accountants, doctors, therapists, contractors โ€” anyone handling other people's private stuff.

Book a call


๐Ÿ–ฅ๏ธ Don't want to DIY? Get this stack on a Mac mini, ready to plug in.

The AirGap Box ships a pre-configured Mac mini to your office with this stack, a 31B-parameter language model, and three working agents already installed.
One-time price. No subscription. Founding-customer pricing for the first 5 buyers.

AirGap Box waitlist


๐ŸŒŒ THE REMATCH โ€” 4 AI Engines Build Northern Lights, 3 Fully Local

Same prompt. Four engines. One MacBook.
The new local challenger โ€” Qwen3.6 27B โ€” painted the best aurora, and never touched the internet.

The Rematch โ€” 4 AI engines build northern lights on one MacBook, 3 fully local

Watch on YouTube ย  Subscribe

Qwen3.6 27B: 5,262 tok / 163s ยท DeepSeek V4 Flash: 3,879 tok / 115s ยท Cloud Claude: 110s ยท Gemma 31B: 2,001 tok / 83s โ€” 3 of 4 ran fully offline.

๐Ÿ HEXAGON SHOOTOUT โ€” Free AI vs $100/mo Claude Code

Three AIs. One laptop. Same prompt. Live counters.
Watch Gemma 31B local, Llama 70B local, and Claude cloud race the same HTML physics prompt on a MacBook.

Hexagon Shootout โ€” 3 AIs, 1 laptop, same prompt, live

Watch on YouTube ย  Subscribe

Gemma 31B: 56s ยท Claude cloud: 22s ยท Llama 70B: 2:17 โ€” two of three ran with zero cloud calls.


๐ŸŽค Also on the channel โ€” NarrateClaude (Hands-Free Ambient AI)

Speak to Claude Code, hear replies in a cloned voice โ€” 100% on-device. 2:31.

NarrateClaude Hands-Free Ambient AI Demo

Watch ย  Subscribe


๐Ÿ  New โ€” My Mac mini at home is the AI. I just talk to it from any browser.

Open any browser on any phone โ€” chat with the Mac mini at home, hear it reply in your own cloned voice. 0:50.

My Mac mini at home is the AI โ€” browser-anywhere demo

Watch ย  Subscribe


๐Ÿงฉ This repo is the BRAIN of a 4-part local-first ambient-computing stack

Brain (here) ยท ๐ŸŽค Ears+Mouth ยท ๐ŸŒ Hands ยท ๐Ÿ“ฑ Phone. Each repo stands alone; together they take Claude Code off the keyboard and off the screen. Jump to the stack diagram โ†’

๐Ÿ–ฅ๏ธ More of my open-source software: nicedreamzwholesale.com/software


๐ŸฅŠ The Lineup โ€” Pick Your Fighter

We started with one model. Now we ship a roster. Same MLX server, same Anthropic API, swap one env var and you swap the brain โ€” plus the brand-new ds4 engine for DeepSeek V4 Flash slotted in via its own native Metal runtime.

๐ŸŸข Gemma 4 31B ๐Ÿ”ต Qwen 3.5 122B ๐Ÿณ DeepSeek V4 Flash โญ
Nickname The Quick One The Beast The 1M-Context Whale
Build 4-bit IT abliterated 4-bit MoE (A10B) 2-bit asymmetric (ds4 GGUF)
Speed ~15 tok/s 65 tok/s ๐Ÿš€ ~32 tok/s
Params 31 B dense 122 B / 10 B active 284 B / 37 B active
Context 128 K 256 K 1 M tokens
RAM ~18 GB ~75 GB ~81 GB
Disk 18 GB 65 GB 81 GB (+ disk KV cache)
Best at Daily coding, fits 64 GB Mac Max throughput, active sparsity Long context, agentic loops
Engine MLX Native MLX Native antirez/ds4
Launcher Gemma 4 Code.command Claude Local.command DeepSeek V4 Flash.app
Min RAM to run 32 GB 96 GB 128 GB

๐Ÿ’ก Fun fact: Qwen wins raw speed because it's an MoE โ€” only 10B of 122B params activate per token. DeepSeek V4 Flash is even bigger (284B) but only ~37B active per token, and it ships with on-disk KV cache so a 25k-token Claude Code system prompt prefills exactly once, ever.

๐Ÿณ New: DeepSeek V4 Flash via ds4

We tested it the day Antirez (the Redis guy) shipped ds4. Local DeepSeek beat cloud Claude on wall-clock time on the same MacBook, same prompt.

Three-way local AI comparison โ€” DeepSeek V4 Flash vs Cloud Claude vs Gemma 4 31B
โ–ถ Watch on YouTube โ€” DeepSeek V4 Flash vs Cloud Claude vs Gemma 4 31B
same prompt ยท three completely different auroras ยท one MacBook

๐Ÿง  Engine antirez/ds4 โ€” pure C + Metal kernels, ~few thousand lines
๐Ÿค— Weights antirez/deepseek-v4-gguf (q2: 81 GB, q4: 153 GB)
๐Ÿ“ฆ Server wrapper ~/.local/bin/ds4-server-up (boots on demand)
๐Ÿš€ Claude Code wrapper ~/.local/bin/claude-ds4 (drop-in replacement for claude)
๐Ÿ“ Context 1 M tokens; 200 K is sane for most agent runs
๐Ÿ’พ Disk KV cache Persists across restarts โ€” first prefill is the only one that ever happens

โญ Our Own MLX Abliterated Uploads

The models in this lineup aren't from generic mirrors โ€” we package and upload our own abliterated MLX builds to HuggingFace so anyone running this repo can pull them with one command. Browse the full set at huggingface.co/divinetribe (also showcased at nicedreamzwholesale.com/software/huggingface/).

# Llama 3.3 70B โ€” full-precision feel
MLX_MODEL=divinetribe/Llama-3.3-70B-Instruct-abliterated-8bit-mlx \
  bash scripts/start-mlx-server.sh

# Gemma 4 31B โ€” fast daily driver
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh

# Hermes 4 14B โ€” sweet spot for 16/32 GB Macs (NEW ยท May 2026)
MLX_MODEL=divinetribe/Hermes-4-14B-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh
Model Quant Disk Params Context Best for
Llama-3.3-70B-Instruct-abliterated-8bit-mlx 8-bit, g64 ~75 GB 71 B dense 128 K Hardest reasoning on 96 GB+ Macs
gemma-4-31b-it-abliterated-4bit-mlx 4-bit, g64 ~17 GB 31 B dense 128 K Daily coding on a 32 GB+ Mac
Hermes-4-14B-abliterated-4bit-mlx 4-bit, g64 ~8 GB 14 B dense (Qwen3 base) 40 K 16 GB Macs, instruction-following, tool use

Abliteration sources: huihui-ai (Llama, Gemma) and Babsie (Hermes). MLX conversion + quantization by us โ€” chosen to preserve quality over minimal footprint. See what abliteration means.

โš ๏ธ Use it responsibly. "Abliterated" suppresses the model's built-in refusal direction so it doesn't refuse benign-but-edgy requests. It is not a general capability upgrade, and you remain bound by each upstream license (Llama 3.3, Gemma, Hermes/Qwen3).


๐ŸŽฎ The Modes

Four ways to run the lineup. Each one is a double-clickable launcher in launchers/.

Mode What it does Launcher
๐Ÿค– Code Run Claude Code with a local model โ€” same UX, no API key Claude Local.command, Gemma 4 Code.command, Llama 70B.command
๐ŸŒ Browser Local AI controls real Brave browser via Chrome DevTools Browser Agent.command
๐ŸŽค Hands-Free Voice Speak in, hear replies in your cloned voice โ€” full loop, 100% on-device Narrative Gemma.command + NarrateClaude
๐Ÿ“ฑ Phone iMessage in โ†’ text/image/video out, full pipeline ~/.claude/imessage-*.sh

๐Ÿค” What Is This?

Your MacBook has a powerful GPU built right into the chip. This project uses that GPU to run massive AI models โ€” the same kind that power ChatGPT and Claude โ€” entirely on your computer.

๐Ÿšซ No internet needed ๐Ÿ’ฐ No monthly subscription ๐Ÿ”’ No one sees your code or data โœ… Full Claude Code experience โ€” write code, edit files, manage projects, control your browser, or run a full hands-free voice session where you speak every question and hear every reply in your own cloned voice (both directions on-device)

         ๐Ÿ“ฑ You (Mac or Phone)
          โ”‚
     ๐Ÿค– Claude Code           โ† the AI coding tool you know
          โ”‚
     โšก MLX Native Server      โ† our server (~1000 lines of Python)
          โ”‚
     ๐ŸฅŠ Pick your fighter     โ† Gemma 4 31B ยท Llama 3.3 70B ยท Qwen 3.5 122B
          โ”‚
     ๐Ÿ–ฅ๏ธ  Apple Silicon GPU    โ† your M-series chip does all the work

๐Ÿ”’ Safety + How the Data Flows

This is the part we're proudest of. Your code never leaves your Mac. Not for a model call. Not for telemetry. Not for "anonymous analytics". Not ever.

๐Ÿ›ก๏ธ The Data-Flow Diagram

   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   โ”‚                    ๐Ÿ–ฅ๏ธ  YOUR MACBOOK                          โ”‚
   โ”‚                                                             โ”‚
   โ”‚   ๐Ÿ“ Your code         โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
   โ”‚       โ”‚                โ”‚  ๐Ÿค– Claude Code     โ”‚               โ”‚
   โ”‚       โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ถโ”‚  (CLI on your Mac)  โ”‚               โ”‚
   โ”‚                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
   โ”‚                                 โ”‚  HTTP localhost:4000       โ”‚
   โ”‚                                 โ–ผ                            โ”‚
   โ”‚                        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
   โ”‚                        โ”‚  โšก MLX Server      โ”‚               โ”‚
   โ”‚                        โ”‚  (Python, ours)    โ”‚               โ”‚
   โ”‚                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
   โ”‚                                 โ”‚  Metal API                 โ”‚
   โ”‚                                 โ–ผ                            โ”‚
   โ”‚                        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
   โ”‚                        โ”‚  ๐Ÿง  Local model     โ”‚               โ”‚
   โ”‚                        โ”‚  (GemmaยทLlamaยทQwen)โ”‚               โ”‚
   โ”‚                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
   โ”‚                                 โ”‚                            โ”‚
   โ”‚                                 โ–ผ                            โ”‚
   โ”‚                        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”               โ”‚
   โ”‚                        โ”‚  ๐Ÿ–ฅ๏ธ  Apple GPU      โ”‚               โ”‚
   โ”‚                        โ”‚  (unified memory)  โ”‚               โ”‚
   โ”‚                        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
   โ”‚                                                             โ”‚
   โ”‚             ๐Ÿšซ ZERO outbound network calls                  โ”‚
   โ”‚             ๐Ÿšซ ZERO telemetry                               โ”‚
   โ”‚             ๐Ÿšซ ZERO phone-home                              โ”‚
   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                   โ”‚
                   โœ—  โ†  Nothing from *our* code crosses this line.
                   โ”‚
   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   โ”‚                    โ˜๏ธ  THE INTERNET                          โ”‚
   โ”‚                  (your code never goes here)                 โ”‚
   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ” What We Audited (Every Component)

Component Source Outbound calls Verdict
server.py (ours) We wrote it line by line 0 โœ… Safe
browser agent (separate repo) nicedreamzapp/browser-agent โ€” we wrote it 0 (talks to localhost CDP only) โœ… Safe
mlx-lm Apple ML team 0 โœ… Safe
MLX framework Apple 0 โœ… Safe
Model weights HuggingFace verified mlx-community repos 0 at runtime โœ… Safe
iMessage scripts Pure shell + AppleScript localhost only (Studio Record port 17494) โœ… Safe
Claude Code CLI Anthropic (closed-source binary) 0 with our launchers โ€” lsof-verified, only localhost:4000 โœ… Safe

โœ… Verified offline (as of v0.1.0). Claude Code 2.1's own binary previously reached out to api.anthropic.com on startup for telemetry, statsig feature flags, marketplace auto-install, and the autoupdater โ€” even with ANTHROPIC_BASE_URL set. PR #32 (thanks @tadrianonet) plugs all four channels via documented Anthropic env vars, and the new launchers set them automatically:

CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
DISABLE_AUTOUPDATER=1
CLAUDE_CODE_DISABLE_OFFICIAL_MARKETPLACE_AUTOINSTALL=1
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1

Run lsof -p $(pgrep -f claude) while a session is active โ€” you'll see only localhost:4000. Your prompts, code, and completions never leave the machine. Our code (server.py, launchers, scripts) has always made zero outbound connections; the Claude Code CLI now matches.

๐Ÿšซ What We Ripped Out

โš ๏ธ We removed LiteLLM after supply-chain attack concerns. Every dependency was re-audited from scratch. If a package had unexplained network calls, it didn't ship.

โœ… What This Means in Practice

Scenario Cloud Claude This Repo
Working with NDA / proprietary code โŒ Risky โœ… Air-gapped (lsof-verified)
Coding on a plane (no wifi) โŒ Doesn't work โœ… Works
Running on a kill-switch firewall โŒ Blocked โœ… Works
Healthcare / legal / finance review โš ๏ธ Compliance burden โœ… Stays on-device
Worry about training-data leakage โš ๏ธ Trust required โœ… Mathematically impossible

๐Ÿ”’ The math is simple: if there are no outbound HTTP calls, your data cannot leak. We grep'd every file for requests, urllib, urlopen, httpx, socket.connect โ€” the only network calls in the entire codebase are to localhost. Run lsof -i -P while it's running. You'll see nothing leaving your Mac.


๐Ÿ“Š Benchmarks

Three generations of optimization. Each one got faster.

โšก Speed Comparison

Generation Approach Speed
๐ŸŒ Gen 1 Ollama 30 tok/s
๐Ÿƒ Gen 2 llama.cpp 41 tok/s
๐Ÿš€ Gen 3 MLX Native (ours) 65 tok/s

โฑ๏ธ Real-World Claude Code Task

How long to ask Claude Code to write a function:

Setup Time
๐Ÿ˜ด Ollama + Proxy 133 s
๐Ÿ˜ llama.cpp + Proxy 133 s
๐Ÿ”ฅ MLX Native (no proxy) 17.6 s

7.5ร— faster โšก โ€” one change (killing the proxy) produced the entire delta. ~1000 lines of Python, no C++ fork, no generic inference backend.

๐ŸฅŠ Lineup Comparison

Model tok/s RAM Best For
๐ŸŸข Gemma 4 31B Abliterated ~15 ~18 GB Daily coding on a 64 GB Mac
๐ŸŸ  Llama 3.3 70B Abliterated ~7 ~70 GB Hardest reasoning, full precision
๐Ÿ”ต Qwen 3.5 122B-A10B 65 ~75 GB Maximum throughput, MoE sparsity

Qwen 122B numbers are measured on M5 Max 128 GB. Gemma and Llama are observed real-world approximations. Full benchmarks for all three pending โ€” see BENCHMARKS.md.

โ˜๏ธ vs Cloud APIs

๐Ÿ–ฅ๏ธ Our Local Setup โ˜๏ธ Claude Sonnet โ˜๏ธ Claude Opus
Speed 65 tok/s ~80 tok/s ~40 tok/s
Monthly cost $0 ๐ŸŽ‰ $20-100+ $20-100+
Privacy 100% local ๐Ÿ”’ Cloud Cloud
Works offline Yes โœˆ๏ธ No No
Data leaves your Mac Never Always Always

๐Ÿ’ก Our local setup beats cloud Opus on raw speed (65 vs 40 tok/s) at $0/month.


๐Ÿ”ง Tool-Call Reliability (v2 โ€” March 2026)

Local models don't format tool calls perfectly. They want to call a tool but mix XML and JSON syntax. Claude Code sees no valid tool call, re-prompts, and the model does it again. The result: infinite loops where the AI says "let me do that" but never actually does anything.

We fixed this. Here's what was happening and what we did about it.

๐Ÿ› The Problem

The model was generating garbled tool calls like this:

<tool_call>
<function=Bash><parameter=command>rm -rf /tmp/old</parameter></function>
</tool_call>

Instead of the correct JSON format Claude Code expects:

<tool_call>
{"name": "Bash", "arguments": {"command": "rm -rf /tmp/old"}}
</tool_call>

The JSON parser choked, Claude Code saw no tool call, re-prompted the model, and the model garbled it the exact same way again โ€” creating an infinite loop.

โœ… The Fix (4 changes to server.py)

Change What Why
KV Cache 4-bit โ†’ 8-bit, quantization starts at token 1024 Model retains conversation context instead of "forgetting" earlier messages
Temperature 0.7 โ†’ 0.2 Less randomness = more consistent tool formatting
Garbled Recovery New recover_garbled_tool_json() function Catches XML-in-JSON hybrids, <function=X><parameter=Y> inside <tool_call> tags, and infers tool names from parameter keys
Retry Logic Up to 2 retries when tool intent is detected but parsing fails Re-prompts with explicit formatting instructions before giving up

๐Ÿงช Test Results

We built an automated test suite (scripts/test_mlx_server.py) that sends real Anthropic API requests to the server simulating multi-step tasks โ€” the exact kind that were failing before.

Test Suite: 14 tests per run
โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
  โœ… Simple Bash commands
  โœ… Directory creation (mkdir -p)
  โœ… File reading (Read tool)
  โœ… Complex Bash with pipes
  โœ… File editing (Edit tool with find/replace)
  โœ… Multi-tool sequences (Glob โ†’ Read)
  โœ… 5 rapid-fire sequential commands
  โœ… Multi-step calendar scenario (create โ†’ delete โ†’ verify)

Results: 98/98 tests passed across 7 consecutive runs. Zero failures.

The multi-step calendar scenario โ€” create 12 month folders, delete all but September, verify โ€” was the exact task that triggered infinite loops before the fix. Now it passes every time.

# Run the test suite yourself:
python3 scripts/test_mlx_server.py

โš™๏ธ Tuning

You can override defaults with environment variables:

Variable Default What It Does
MLX_MODEL divinetribe/gemma-4-31b-it-abliterated-4bit-mlx Pick which fighter to load
MLX_KV_BITS 8 KV cache quantization bits (4 saves memory, 8 improves coherence)
MLX_KV_QUANT_START 1024 Token position where KV quantization begins
MLX_TOOL_RETRIES 2 Max retries when a garbled tool call is detected
MLX_MAX_TOKENS 8192 Max output tokens per response
MLX_SUPPRESS_THINKING 1 Pre-fill an empty thinking block so Gemma 4 skips its reasoning chain entirely. Saves ~1 min/request. Set to 0 if you want the model to reason before responding.

๐Ÿ“ฑ Control From Your Phone โ€” Full Media Pipeline

You don't have to be at your Mac to use this. Text a command, get back a full video.

๐Ÿ“ฑ Your iPhone                    ๐Ÿ’ป Your Mac
     โ”‚                                โ”‚
     โ”‚โ”€โ”€ "find me an article  โ”€โ”€โ”€โ”€โ”€โ”€>โ”‚โ”€โ”€ imessage-receive.sh reads it
     โ”‚    and send me a video"        โ”‚โ”€โ”€ local model plans the task
     โ”‚                                โ”‚โ”€โ”€ Brave browser finds the article
     โ”‚                                โ”‚โ”€โ”€ speak narrates in your voice
     โ”‚                                โ”‚โ”€โ”€ Studio Record captures it all
     โ”‚                                โ”‚โ”€โ”€ build_production_video.py edits it
     โ”‚<โ”€โ”€ ๐ŸŽฅ video in iMessage โ”€โ”€โ”€โ”€โ”€โ”€โ”‚โ”€โ”€ imessage-send-video.sh ships it
     โ”‚                                โ”‚
   ๐Ÿ›‹๏ธ  From your couch            ๐Ÿ–ฅ๏ธ  At your desk

Everything works โ€” text, images, and video:

Command What happens You get
"summarize this article" Local model reads + replies ๐Ÿ’ฌ Text
"send me a screenshot of X" Claude screenshots ๐Ÿ“ธ Image in iMessage
"screen record you doing Y" Records + sends ๐ŸŽฅ Video in iMessage
"make me a produced video" Full edit pipeline ๐ŸŽฌ Title card + subs

Full pipeline repo: nicedreamzapp/claude-screen-to-phone โ†’ Clone it, run setup.sh, fill in your phone number. Works with this local AI stack or Claude cloud.

We built this before Anthropic shipped their Dispatch feature. Same concept, but ours uses iMessage, runs on your local model, and can send back media โ€” not just text.

๐Ÿ’ก Pro tip: Anthropic's Dispatch doesn't read your CLAUDE.md. Mention it in your message or it'll miss your custom setup. Our iMessage system doesn't have this problem.


๐Ÿ’ก How We Got Here

Most people trying to run Claude Code locally hit the same wall:

Claude Code speaks Anthropic API. Local models speak OpenAI API. Different languages. ๐Ÿคท

So everyone builds a proxy to translate between them. That proxy adds latency, complexity, and breaks things.

We took a different approach:

๐ŸŒ What everyone else does ๐Ÿš€ What we did
Claude Code โ†’ Proxy โ†’ Ollama โ†’ Model Claude Code โ†’ Our Server โ†’ Model
3 processes, 2 API translations 1 process, 0 translations
133 seconds per task 17.6 seconds per task

๐ŸŽฏ That one change โ€” eliminating the proxy โ€” made it 7.5x faster.


๐Ÿ’ป What You Need

Your Mac RAM What You Can Run
M1/M2/M3/M4 (base) 8-16 GB ๐ŸŸก Small models (4B)
M1/M2/M3/M4 Pro 18-36 GB ๐ŸŸ  Gemma 4 31B (tight)
M2/M3/M4/M5 Max 64-128 GB ๐ŸŸข Gemma 4 31B + ๐Ÿ”ต Qwen 3.5 122B
M2/M3/M4 Ultra 128-192 GB ๐Ÿ”ต Multiple large models, all three fighters

Also need:

  • ๐Ÿ Python 3.12+ (for MLX)
  • ๐Ÿค– Claude Code (npm install -g @anthropic-ai/claude-code)

๐Ÿš€ Quick Start (3 Commands)

git clone https://github.com/nicedreamzapp/claude-code-local
cd claude-code-local
bash setup.sh

setup.sh auto-detects your RAM, picks a model from the lineup, downloads it, installs the MLX server, and creates a Claude Local.command launcher on your Desktop.

Then double-click Claude Local.command. You're coding locally.

๐Ÿ› If the launcher asks you to sign in to a Claude account: your claude CLI is too old. The launchers pass --bare to force local-only API-key auth, but older versions of the CLI don't support that flag and fall through to the Anthropic login prompt. Fix:

npm install -g @anthropic-ai/claude-code
claude --version   # should print a recent version

๐Ÿ› ๏ธ Note for contributors / hackers: setup.sh installs the server as a symlink at ~/.local/mlx-native-server/server.py pointing back at this repo's proxy/server.py. Edit the file in the repo, restart the MLX server, done โ€” no re-running setup.sh, no copying, no silent drift between "what I committed" and "what's actually running." There is one source of truth for the server, and it's proxy/server.py in the repo.

Or do it manually

# 1. Set up the MLX virtualenv
python3.12 -m venv ~/.local/mlx-server
~/.local/mlx-server/bin/pip install mlx-lm

# 2. Pick a fighter and download (one time, ~18-75 GB)
bash scripts/download-and-import.sh gemma   # or 'llama' or 'qwen'

# 3. Start the server
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh

# 4. Launch Claude Code
ANTHROPIC_BASE_URL=http://localhost:4000 \
ANTHROPIC_API_KEY=sk-local \
claude --model claude-sonnet-4-6

๐Ÿ’ก Or just double-click a launcher in launchers/. They do all of this automatically.


๐Ÿ”ง How It Works

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚              Your MacBook (M5 Max)               โ”‚
โ”‚                                                  โ”‚
โ”‚  ๐Ÿ“ You type โ”€โ”€> ๐Ÿค– Claude Code                  โ”‚
โ”‚                      โ”‚                           โ”‚
โ”‚                      โ–ผ                           โ”‚
โ”‚                 โšก MLX Server (port 4000)        โ”‚
โ”‚                      โ”‚                           โ”‚
โ”‚                      โ–ผ                           โ”‚
โ”‚                 ๐ŸฅŠ Local model โ”€โ”€> ๐Ÿ–ฅ๏ธ  GPU        โ”‚
โ”‚                 (GemmaยทLlamaยทQwen)               โ”‚
โ”‚                      โ”‚                           โ”‚
โ”‚                      โ–ผ                           โ”‚
โ”‚  ๐Ÿ“ Answer <โ”€โ”€โ”€ โœจ Clean response                โ”‚
โ”‚                                                  โ”‚
โ”‚         ๐Ÿ”’ Nothing leaves this box. Ever.        โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

The server (proxy/server.py) is one file, ~1000 lines. It does six things:

  1. ๐Ÿ“ฆ Loads the model โ€” Apple's MLX framework, native Metal GPU, unified memory. Handles Gemma's RotatingKVCache quirk automatically so sliding-window models don't crash on the first request.
  2. ๐Ÿ”Œ Speaks Anthropic API โ€” Claude Code thinks it's talking to Anthropic's cloud. It's not.
  3. ๐Ÿ”ง Translates tool use โ€” Three different tool-call formats in and out: Gemma 4 native (<|tool_call>call:Name{...}<tool_call|>), Llama 3.3 raw JSON ({"type":"function",...}), and HuggingFace <tool_call> JSON (Qwen and others). All converted โ†” Anthropic tool_use blocks, with garbled-output recovery for small models.
  4. ๐Ÿงน Cleans the output โ€” Local models think out loud in <think> / <|channel>thought tags, emit stop markers (<turn|>, <|python_tag|>), and sometimes drop in reasoning preamble. A real-time ThinkingFilter strips thinking blocks token-by-token during generation โ€” before they accumulate in the buffer โ€” then clean_response handles the rest.
  5. โšก Reuses prompt caches across requests โ€” so Claude Code's 4K-token system prompt doesn't get re-prefilled on every turn. Huge speedup for short questions.
  6. ๐ŸŽฏ Code mode โ€” auto-detects Claude Code coding sessions (any of Bash/Read/Edit/Write/Grep/Glob in the tools list), swaps Claude Code's ~10K-token harness prompt for a slim ~150-token one tuned for local models, and strips verbose tool descriptions down to name + parameter types. In practice: 35 tools with full descriptions = ~5 600 prompt tokens; after code mode, ~200 tokens โ€” a 28ร— reduction that cuts prefill time from ~60 s to ~2 s on Gemma 4 31B. Also stops models from refusing with "I am not able to execute this task."

๐Ÿ”Œ MCP Servers โ€” Claude Code's plugin ecosystem, 100% local

The only way to run Claude Code's full MCP plugin ecosystem 100% local on Apple Silicon.

Claude Code talks to the world through MCP servers โ€” Anthropic's plugin protocol. There's a fast-growing ecosystem of them: filesystem, GitHub, Postgres, Slack, web search, Apple Notes, Notion, Chrome DevTools, and a couple hundred more. They're how Claude Code reads your files, browses the web, queries your databases, controls your browser.

Most local-LLM proxies break MCP. They strip the tool definitions, mangle the tool_use blocks, or refuse to forward the streaming format Claude Code expects. So even if you swap in a "Claude alternative," your plugins stop working.

claude-code-local doesn't break MCP. The proxy passes tool definitions through to your local model and translates the model's tool_use blocks back into Anthropic's format โ€” across all three model families (Gemma 4 native, Llama 3.3 raw JSON, Qwen <tool_call> JSON), with garbled-output recovery for small models. From Claude Code's perspective it's talking to Anthropic. From your MCP server's perspective, the same Claude Code is calling it. Nothing in the middle changes.

How to plug in a server

Wire MCP servers up the normal Claude Code way (~/.claude.json or per-project .mcp.json). Make sure your ANTHROPIC_BASE_URL is pointed at the local proxy, then add the server. Three quick examples:

1. Filesystem โ€” let the local model read/write a folder

# Anthropic's reference filesystem MCP server
claude mcp add filesystem -- npx -y @modelcontextprotocol/server-filesystem ~/projects

Now you can launch Claude Code (any of the launchers in launchers/) and ask it to "summarize every README in ~/projects" โ€” it'll call the filesystem MCP server, which streams files back to your local Gemma/Qwen, which writes the summary. Zero cloud round-trips.

2. GitHub โ€” issues, PRs, code search, all local

claude mcp add github --env GITHUB_TOKEN=$GITHUB_TOKEN -- npx -y @modelcontextprotocol/server-github

Now your local model can read GitHub issues, draft PRs, search code across repos. The model still runs on your Mac; only the GitHub API calls leave the building (which is fine โ€” that's GitHub's data, not yours).

3. Web search โ€” for when the local model needs fresh info

# Brave Search MCP (free tier, no PII)
claude mcp add brave-search --env BRAVE_API_KEY=$BRAVE_API_KEY -- npx -y @modelcontextprotocol/server-brave-search

Now your local Gemma can answer "what's the latest version of MLX?" without hallucinating.

MLX_BROWSER_MODE โ€” optimized for chrome-devtools MCP

Claude Code's chrome-devtools MCP integration sends a 30+ tool list and a 10K-token system prompt to every request. That's fine for cloud Claude. It's death for a local model.

Set MLX_BROWSER_MODE=1 when starting the proxy and it auto-detects Claude Code MCP browser sessions (by looking for mcp__chrome-devtools__* tool registrations), strips the bloat, and keeps only the 9 essential browser-control tools. Same browser automation, ~99% fewer tokens to chew through.

MLX_BROWSER_MODE=1 ./scripts/start-mlx-server.sh

Direct clients (anything that brings its own system prompt + tools) are passed through untouched โ€” only Claude Code MCP sessions get the optimization.

What this unlocks

Honestly the whole MCP ecosystem becomes available to you with no compromise. Every tool the cloud-Claude-Code-using developer has โ€” filesystem, GitHub, web search, browser automation, database access, calendar, anything someone has shipped an MCP server for โ€” works the same against your local Gemma or Qwen. The 200+ tool universe is yours, just running on your machine instead of someone else's.


๐ŸŒ Browser Agent

A standalone browser agent that controls your real Brave browser via Chrome DevTools Protocol โ€” powered entirely by local AI. No Claude Code wrapper needed.

๐Ÿงญ The browser agent lives in its own repo: nicedreamzapp/browser-agent. It's not bundled inside this repo. The Browser Agent.command launcher here points at the installed location (~/.local/browser-agent/agent.py) that you get from cloning the browser-agent repo separately. Keeping it in its own project keeps both repos focused and stops "edit the wrong file" drift between a vendored copy and the real source of truth.

         ๐Ÿ“ Your task
          โ”‚
     ๐Ÿค– agent.py              โ† autonomous browser agent (separate repo)
          โ”‚
     โšก MLX Server             โ† local AI decides what to do
     (Gemma ยท Llama ยท Qwen)
          โ”‚
     ๐ŸŒ Brave (CDP port 9222) โ† clicks, types, navigates your real browser
          โ”‚
     ๐Ÿ“Š Context Meter          โ† shows memory usage so you know its limits

Context memory pipeline โ€” the agent doesn't forget what it's doing:

๐ŸŒ Old Behavior ๐Ÿš€ New Pipeline
Memory Hard drop after 5 steps Smart trim at 60% of 32K budget
When trimming Deletes old steps entirely Compresses into summary
Original task Lost after step 6+ Re-injected every cycle
Visibility None โ€” flying blind Color-coded context meter
Response tokens 1,024 2,048

The context meter shows green/yellow/red after each step:

  Step 5 snapshot() 2.2s
         โ†’ [101] heading "The Best Coffee Cake Recipe"...
  [Context: 18% โ–ˆโ–ˆโ–ˆโ–ˆโ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘โ–‘ 6K/32K tokens]    โ† green = plenty of room

๐Ÿ’ก Double-click Browser Agent.command to launch. It starts the MLX server, opens Brave with remote debugging, and drops you into the agent.


๐ŸŽค Hands-Free Voice Mode โ€” The Whole Loop On-Device

Talk to your Mac. It talks back in your own cloned voice. Nothing touches the internet in either direction.

This is the feature I'm proudest of in the whole stack, and the one I haven't seen anyone else demo publicly. Most "AI voice" demos use cloud STT (Whisper API, Deepgram, Google Cloud Speech) and cloud TTS (ElevenLabs cloud, OpenAI, Azure) โ€” so your voice hits someone else's server before you see a word of transcript, and every reply makes another cloud round-trip back as audio. This doesn't. Both sides of the loop run fully on your Mac, end to end.

The full voice loop

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     YOUR MACBOOK (M-series)                     โ”‚
โ”‚                                                                 โ”‚
โ”‚    ๐ŸŽ™๏ธ  Your voice                                               โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐ŸŽง listen  (custom Swift binary)                             โ”‚
โ”‚       โ€ข Apple SFSpeechRecognizer โ€” on-device engine             โ”‚
โ”‚       โ€ข Continuous listening, stability-based utterance end     โ”‚
โ”‚       โ€ข Auto-pauses during playback to stop feedback loops      โ”‚
โ”‚       โ€ข Wedge-detection watchdog, preventive 10-min recycle     โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐Ÿ“ฌ dispatch  (bash watchdog + router)                        โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    โŒจ๏ธ  inject  (AppleScript โ†’ target Terminal window by id)     โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐Ÿค– claude  (narration persona loaded from CLAUDE.md)         โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    โšก MLX Server โ†’ ๐ŸฅŠ Gemma 4 31B  (local, 4-bit, ~15 tok/s)    โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐Ÿ”Š ~/.local/bin/speak  "naturally phrased reply"             โ”‚
โ”‚       โ€ข Pocket TTS with your own cloned voice                   โ”‚
โ”‚       โ€ข Or any TTS that takes text + plays audio                โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐ŸŽต afplay  (listen pauses itself during this so the          โ”‚
โ”‚                model's own voice doesn't feed back in)          โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ–ผ                                                       โ”‚
โ”‚    ๐Ÿ‘‚ You hear it                                               โ”‚
โ”‚         โ”‚                                                       โ”‚
โ”‚         โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ and you keep talking                   โ”‚
โ”‚                                                                 โ”‚
โ”‚           ๐Ÿ”’ Your voice never leaves this box. Ever.            โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

What makes this actually work

  • ๐ŸŽ™๏ธ Speech-in โ€” a compiled Swift binary wraps Apple's SFSpeechRecognizer (the same on-device engine that powers macOS Dictation) in a continuous listening loop rather than the usual Fn-Fn toggle. End of utterance is detected via partial-result stability: if the transcribed text stops changing for 2.5 seconds, the recognizer finalizes that sentence. That's way more robust than silence/RMS heuristics against background noise, fans, or music.
  • ๐Ÿ”Š Speech-out โ€” a CLI at ~/.local/bin/speak wraps Pocket TTS driving a cloned copy of Matt's own voice. Any TTS that accepts a string and plays audio slots in โ€” macOS say, Piper, local ElevenLabs, your choice.
  • ๐Ÿ” Feedback-loop prevention โ€” the listener auto-pauses while afplay is running, so the TTS output of one turn never gets picked up as input for the next. No "the model talking to itself" loops.
  • ๐Ÿง  Speak-every-turn is enforced via system prompt โ€” NarrativeGemma/CLAUDE.md is loaded as the narration persona. It tells Gemma to narrate every tool call, every reasoning step, every result, before it writes the text reply. You're never staring at a silent terminal wondering if it's thinking.
  • ๐Ÿ›ก๏ธ Real production hardening โ€” 10-minute preventive process recycle (dodges a known SFSpeech daemon wedge), queue-backlog detection with a non-zero exit code when the listener is stuck. Not a demo script โ€” a tool that has to run unattended for hours.

Why it matters

"Voice-controlled AI" is everywhere right now, but under the hood almost every public demo is a cloud pipeline wearing a local-looking coat. If the network drops, the demo dies. If your client's laptop blocks outbound connections, the demo dies. If you're on a plane, in a Faraday cage, or debugging on a disconnected-by-policy machine, the demo dies.

This setup doesn't die. Apple's on-device speech engine is a fully local model that already ships with the OS, and accessing it via SFSpeechRecognizer is a first-class macOS API โ€” it's just that almost nobody wraps it in a continuous-listen daemon with production hardening and plumbs it to a local LLM with a cloned-voice reply stream. Now there's one.

How to wire it up

๐Ÿ› ๏ธ The listening stack lives in its own repo. The Listen.swift binary, the dictation / dispatch / inject scripts, and the narrative-claude.sh launcher are a sibling project: nicedreamzapp/NarrateClaude. Same design as the browser agent: one repo per focused tool, so edits don't drift between a vendored copy and the real source of truth.

The two halves of the loop, and where each half lives

๐Ÿ—ฃ๏ธ The speak-and-think half (this repo, claude-code-local):

  • launchers/Narrative Gemma.command โ€” boots the MLX server with Gemma 4 31B and injects the narration persona via MLX_APPEND_SYSTEM_PROMPT_FILE so Gemma narrates every turn
  • NarrativeGemma/CLAUDE.md โ€” the narration persona itself (opt-in, sanitized, generic)
  • ~/.local/bin/speak โ€” your chosen TTS CLI (Matt uses Pocket TTS with a cloned voice; say "$@" works as a three-line stub if you don't have a fancier setup)

๐ŸŽง The listen-and-inject half (NarrateClaude, sibling repo):

  • A compiled Swift binary wrapping Apple's SFSpeechRecognizer in continuous-listen mode with stability-based end-of-utterance detection and wedge-recovery
  • A bash dispatch pipeline that respawns the listener, watches the target Terminal window, and tears everything down cleanly when you close the session
  • An AppleScript injector that writes transcribed utterances straight into the bound Terminal tab by window ID
  • A narrative-claude.sh one-click launcher that opens the Terminal, starts Claude Code, captures the window ID, and starts the listener

Running the full hands-free loop

# 1. Install this repo (claude-code-local) โ€” gives you the MLX server + Narrative launcher
git clone https://github.com/nicedreamzapp/claude-code-local.git "$HOME/Desktop/Local AI Setup"
cd "$HOME/Desktop/Local AI Setup" && bash setup.sh

# 2. Install the sibling NarrateClaude repo โ€” gives you the listening pipeline
git clone https://github.com/nicedreamzapp/NarrateClaude.git ~/NarrateClaude
cd ~/NarrateClaude && chmod +x dictation/bin/* narrative-claude.sh
./dictation/bin/dictation setup   # compiles the Swift listener + grants permissions

# 3. Launch the full loop
bash ~/NarrateClaude/narrative-claude.sh

๐Ÿ’ก Double-click Narrative Gemma.command from this repo to run the model-and-speak side standalone (keyboard in, voice out โ€” useful when you don't want to be on mic). Run narrative-claude.sh from the NarrateClaude repo to launch the full hands-free loop (voice in, voice out, no keyboard at all).


โœˆ๏ธ When To Use This

Situation Use This? Why
On a plane โœ… Full AI coding, no internet needed
Sensitive client code โœ… Nothing leaves your machine
Don't want API fees โœ… $0/month forever
Want fastest possible โ˜๏ธ Cloud Sonnet is still slightly faster
Need Claude-level reasoning โ˜๏ธ Local models are good, not Claude-level
Controlling from phone โœ… iMessage pipeline works offline
Healthcare / legal / finance review โœ… 100% on-device, audit-friendly

๐Ÿ“ What's In This Repo

๐Ÿ“ฆ claude-code-local/
 โ”œโ”€โ”€ โšก proxy/
 โ”‚   โ””โ”€โ”€ server.py              โ† MLX Native Anthropic Server with tool-call recovery (~1000 lines)
 โ”œโ”€โ”€ ๐Ÿš€ launchers/
 โ”‚   โ”œโ”€โ”€ Claude Local.command    โ† Default fighter โ€” Claude Code + local model
 โ”‚   โ”œโ”€โ”€ Gemma 4 Code.command    โ† ๐ŸŸข THE QUICK ONE
 โ”‚   โ”œโ”€โ”€ Llama 70B.command       โ† ๐ŸŸ  THE WISE ONE
 โ”‚   โ”œโ”€โ”€ Browser Agent.command   โ† ๐ŸŒ Autonomous Brave browser control
 โ”‚   โ”œโ”€โ”€ Narrative Gemma.command โ† ๐ŸŽญ Auto-narration mode
 โ”‚   โ””โ”€โ”€ lib/claude-local-common.sh โ† Shared: model-aware restart, local-cache resolver, health-wait
 โ”œโ”€โ”€ ๐ŸŽญ NarrativeGemma/
 โ”‚   โ””โ”€โ”€ CLAUDE.md              โ† Narration persona (sanitized, generic, opt-in)
 โ”œโ”€โ”€ ๐Ÿ› ๏ธ  scripts/
 โ”‚   โ”œโ”€โ”€ download-and-import.sh โ† Download a fighter (`gemma` / `llama` / `qwen`)
 โ”‚   โ”œโ”€โ”€ persistent-download.sh โ† Auto-retry downloader for big models
 โ”‚   โ”œโ”€โ”€ start-mlx-server.sh    โ† Server start helper
 โ”‚   โ”œโ”€โ”€ test_mlx_server.py     โ† Tool-call reliability test suite
 โ”‚   โ””โ”€โ”€ upload-mlx-quant.sh    โ† Publish your own MLX-quantized uploads to HF
 โ”œโ”€โ”€ ๐Ÿ“Š docs/
 โ”‚   โ”œโ”€โ”€ BENCHMARKS.md          โ† Detailed speed comparisons
 โ”‚   โ””โ”€โ”€ TWITTER-THREAD.md      โ† Social media content
 โ”œโ”€โ”€ ๐Ÿ“ฑ IMESSAGE_MEDIA_PIPELINE.md โ† Phone control + media sending docs
 โ””โ”€โ”€ setup.sh                    โ† One-command installer

๐Ÿ›ค๏ธ The Journey

We didn't start here. We went through three generations in one night:

Gen What We Tried Speed ๐Ÿ’ก What We Learned
1๏ธโƒฃ Ollama + custom proxy 30 tok/s Ollama works but Claude Code can't talk to it directly
2๏ธโƒฃ llama.cpp TurboQuant + proxy 41 tok/s TurboQuant compresses KV cache 4.9x, but the proxy is the bottleneck
3๏ธโƒฃ MLX native server 65 tok/s Kill the proxy. Speak Anthropic API directly. 7.5x faster.
4๏ธโƒฃ The lineup 65 / 15 / 7 tok/s Three brains, one server. Same MLX, same Anthropic API โ€” swap one env var to change the fighter.

๐ŸŽฏ Each generation taught us something. Killing the proxy made it fast. Adding the lineup made it flexible.


๐Ÿงฉ The Complete Local-First Stack

claude-code-local is the brain โ€” MLX Anthropic server, launcher lineup, tool-call translation. It pairs with three sibling repos to form a local-first ambient computing stack that never sends a keystroke, a voice clip, or a page load to the cloud. Each repo stands alone.

๐Ÿค– claude-code-local โ€” Brain (you are here)

MLX + Gemma 31B / Llama 70B / Qwen 122B ยท Anthropic API server ยท tool-call parsing ยท prompt cache. Zero cloud, 65 tok/s on Apple Silicon.

๐ŸŽค NarrateClaude โ€” Ears + Mouth

Talk to Claude, hear replies in your cloned voice โ€” both directions on-device. Fully hands-free loop using Apple SFSpeech + cloned-voice TTS.

๐ŸŒ browser-agent โ€” Hands

Drives a real Brave browser via Chrome DevTools Protocol. Handles iframes, Shadow DOM, ProseMirror.

๐Ÿ“ฑ claude-screen-to-phone โ€” Remote

Turns your iPhone into a full Claude Code terminal. Text any command โ€” git, shell, file edits, deploys, browser tasks โ€” and get back whatever Claude produces (text, screenshots, screen recordings, produced videos) right in Messages. Works over iMessage โ€” no bots, no third-party apps, no cloud relay.

Pair any combination. All four = ambient computing on one Mac, nothing in the cloud.

๐Ÿชด Why this matters โ€” the ambient-computing angle

The real goal isn't "a faster Claude Code" โ€” it's getting off screens and mice. Hunched-over-screen computing is breaking our bodies: carpal tunnel, curved spines, $1500 ergonomic chairs bought to patch the damage the rest of the desk is doing. That era is ending. These three repos are pieces of what comes next โ€” computing that's around you instead of in front of you. Screens become optional, typing becomes optional, sitting still becomes optional, but your data and your voice never leave your house.

๐Ÿ‘‰ For the full manifesto, see the "Why I Built This โ€” Ambient Computing Starts Here" section in the NarrateClaude README. That's where the philosophy lives; the repos are just the first implementations.


๐Ÿ›ฃ๏ธ What's Next

We ship fast and in public. Rough direction for the next few weeks โ€” if any of these excite you, hit Watch (top-right of the repo) to get the release ping.

  • ๐ŸŸก Full Qwen 3.5 122B benchmark suite โ€” reliability, tool-call pass rate, cold-start vs warm, long-context behavior vs Gemma
  • ๐ŸŸก Fully-local Whisper fallback โ€” drop-in alternative to the Apple SFSpeechRecognizer path for older Macs and non-English voices
  • ๐ŸŸก One-click DMG installer โ€” double-click-to-run setup for folks who just want Claude Code + local AI without a terminal
  • ๐ŸŸก MLX_MODEL=<hf-url> โ€” point at any HuggingFace repo and have the lineup auto-register a new fighter
  • ๐ŸŸก More fighters โ€” open to PRs adding launchers for DeepSeek, Mistral, Phi, anything MLX-compatible

๐Ÿ’ก Want something that's not on this list? Open an issue โ†’. Every serious request gets read and usually replied to within 24h.


๐Ÿค Contributing & Ideas

A lot has changed since this repo was one night of "can I run Claude Code on Ollama." It's now a full local-AI stack: a ~1000-line MLX-native Anthropic server, prompt-cache reuse, Gemma / Llama / Qwen native tool-call parsing, code mode (auto-strips Claude Code's 10K-token harness prompt for local models), the browser agent, narration mode, an iMessage pipeline, model-aware launcher restart, and โ€” the piece I think is the biggest deal โ€” a fully on-device hands-free voice loop (Apple SFSpeechRecognizer + cloned-voice TTS) that lives in the sibling NarrateClaude project. Way past what "The Journey" table above covers.

I built this because it solves my workflow end to end. Coding on planes, sensitive client work, drafting from my phone, handing off to local models when I don't want cloud latency or cloud bills, and (the thing I come back to most) running actual coding sessions hands-free โ€” speak a request, listen to Gemma narrate the plan, hear it confirm the result, keep talking. No keyboard, no screen-watching. The whole loop is in-place today. I'd love to hear how others could use it.

If you have ideas, bug reports, a new launcher for a model I don't run, a better code-mode prompt, or a workflow this doesn't cover โ€” open an issue or a PR. I read them all. Especially interested in hearing from:

  • ๐Ÿง  People on older Apple Silicon (M1 / M2, 16โ€“36 GB) who know which models actually fit and still do useful coding work
  • ๐ŸŽค Anyone who wants to stress-test the hands-free voice loop on different hardware, different TTS voices, or different dictation accents โ€” we're currently running it on one M5 Max with one cloned voice
  • ๐Ÿ”Š TTS recipes beyond Pocket TTS โ€” Piper, local ElevenLabs, MLX-TTS, Kyutai Moshi, or anything else that slots cleanly into ~/.local/bin/speak
  • ๐Ÿ”Œ Folks with workflows this doesn't touch yet โ€” what would you want from a local Claude Code?
  • ๐Ÿ› Anyone who runs into edge cases I'll never hit on an M5 Max with 128 GB

Small PRs welcome, huge PRs welcome, issues with no PR welcome. The whole point is that it's yours to bend.


๐Ÿ™ Credits

Built on the shoulders of giants:

Project What It Does By
๐Ÿค– Claude Code AI coding agent Anthropic
๐ŸŽ MLX Apple Silicon ML framework Apple
๐Ÿ“ฆ mlx-lm Model loading + inference Apple
๐ŸŸข Gemma The 31B fighter (base weights) Google DeepMind
โญ Gemma 4 31B Abliterated 4-bit MLX Our own MLX-packed abliterated upload โ€” THE QUICK ONE in the lineup divinetribe (us)
๐ŸŸ  Llama The 70B fighter (base weights) Meta
โญ Llama 3.3 70B Abliterated 8-bit MLX Our own MLX-packed abliterated upload โ€” THE WISE ONE in the lineup divinetribe (us)
๐Ÿ”ง huihui-ai Original abliteration of Llama 3.3 70B Instruct huihui-ai
๐Ÿ“– Abliteration explained The technique we built on Maxime Labonne
๐Ÿ”ต Qwen 3.5 The 122B fighter Alibaba
โšก TurboQuant KV cache compression research Google Research

Tested on Apple M5 Max with 128 GB unified memory.


๐Ÿ‘‹ Who built this

Built by Matt Macosko in Arcata, CA. All of this is part of the Nice Dreamz LLC umbrella โ€” the consulting + open-source side of what I do day-to-day.

If this repo is useful to you, here's the rest of the work:

๐Ÿ”’ AirGap AI Local AI partnership for firms that can't put client work in someone else's cloud โ€” law, accounting, medical, therapy. I install on-device AI inside your office, train your team, and keep your stack on the current best open-source model as the field evolves (signed USB w/ checksums, or in-person install โ€” your call). One firm at a time. Book a 15-min call.
๐Ÿ–ฅ๏ธ Nice Dreamz Software The rest of the open-source lineup โ€” NarrateClaude, the browser agent, CemaniHomesteadRobot, VisionBuilder, and more.
๐ŸŒฟ Divine Tribe The hardware side โ€” Core XL, V5, Ruby Twist. 13 years of building physical products.
๐Ÿ“ฐ Marijuana Union Community + news site. Where the long-form writing lives.
๐ŸŒฑ Tribe Seed Bank Seeds marketplace.

Find me:

X LinkedIn YouTube GitHub Instagram


๐Ÿ›Ÿ Related: claude-failover

If you'd rather keep using Claude as primary but want a local backstop for when your Max plan limit pinches or Anthropic has an outage, see the sibling project:

๐Ÿ‘‰ github.com/nicedreamzapp/claude-failover โ€” one-command flip between Claude and a local mlx-lm model. Lazy-loads, zero RAM cost when not in failover. Different angle from this repo: Claude stays default, local kicks in only when you flip the switch.


๐Ÿ’ฌ Community

A Discord for builders running, contributing to, or hacking on claude-code-local, NarrateClaude, and browser-agent. Share what you're building, ask questions, swap MLX tips. Quiet, builder-tone, no bots.

Join the NiceDreamzApps Discord

๐Ÿ‘‰ discord.gg/ZdSqgAxUW


๐Ÿ“œ MIT License โ€” Use it however you want.

โญ Star this repo if it helped you! โญ

GitHub stars GitHub forks

Contributors

nicedreamzappasdmomentkevbarns0xshugotripathiprateek

Issues