- Install Karabiner and set up Halmak layout
- Install nvm
- Install Node 24:
nvm install 24 - Install TypeScript:
npm i -g typescript - Install brew:
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" - Install zed
- Install Haskell via GHCUp:
curl --proto '=https' --tlsv1.2 -sSf https://get-ghcup.haskell.org | sh - Install GitHub CLI:
brew install gh - Install GitHub Copilot:
npm i -g @github/copilot - Install Cocktail:
npm i -g @camunda8/cli@alpha - Install Pi:
npm install -g --ignore-scripts @earendil-works/pi-coding-agent - Install Pi LSP:
pi install npm:pi-lsp-extension - Install Pi GitHub:
pi install npm:pi-gh-cli - Install Pi llama.cpp:
pi install npm:pi-llama-cpp - Install Pi Context-Mode:
npm install -g context-mode && pi install npm:context-mode - Install Pi Memory:
pi install npm:@pi-unipi/memory - Install Zed Biome: "Open zed: extensions and search for Biome"
- Install Zed Haskell: "Open zed: extensions and search for Haskell"
- Install llama.cpp:
brew install llama.cpp - Get the Gemma 4 model (from here):
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M --port 1234 -ngl 99 -c 32768 -np 1 --jinja -ctk q8_0 -ctv q8_0 - Configure pi for llama.cpp — edit
~/.pi/agent/settings.jsonadd:"llamaServerUrl": "http://127.0.0.1:1234" - Generate SSH Key for GitHub
- Install OpenJDK:
brew install openjdk - Install GhostTTY
- Install Orbstack:
brew install orbstack - Install uv (Python):
curl -LsSf https://astral.sh/uv/install.sh | sh - Install asdf:
brew install asdf - Install asdf .NET plugin:
asdf plugin add dotnet - Install .NET 8:
asdf install dotnet 8.0.421 - Install .NET 9:
asdf install dotnet 9.0.31430: Install .NET 10:asdf install dotnet latest - Install Clojure:
brew install clojure/tools/clojure - Install Zed Clojure: "Open zed: extensions and search for Clojure"
- Install Little Coder:
npm i -g little-coder - Set nano as git editor:
git config --global core.editor "nano" - Set up git identity:
git config --global --edit - Install bun:
npm i -g bun - Install qmd for Pi memory search:
npm i -g @tobi/qmd - Install MacMLX, and download Qwen model
- Install omlx:
brew install omlx --with-grammar - Model for omlx config in ~/.config/little-coder/models.json:
{
"providers": {
"omlx": {
"api": "openai-completions",
"baseUrl": "http://127.0.0.1:8000/v1",
"apiKey": "IGNORED",
"models": [
{
"id": "Qwen3-32B-4bit",
"name": "Qwen3.6-35B-A3B (local omlx, 150K)",
"reasoning": true,
"input": ["text"],
"contextWindow": 150000,
"maxTokens": 4096,
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 }
}
]
}
}
}Serve several local GGUF models behind one OpenAI-compatible endpoint using
llama.cpp's built-in router mode. This exposes llama.cpp's native model
management (/models, /models/load, /models/sse) that the pi-llama-cpp
extension uses to browse/load/switch models — plain llama-swap does not
implement those endpoints (pi gets HTTP 404), so use the router for pi.
./start-llm-router.sh # binds :: (IPv4+IPv6) on :8888- Models are defined in
llama-router.ini(currentlyqwen3.8andminicpm5-2b); the[*]section holds shared defaults. Local models usemodel = /path/to.gguf, remote ones usehf = user/repo:quant. - The router also auto-lists any other GGUFs in your Hugging Face cache.
- Env overrides:
LLAMA_ROUTER_PORT,LLAMA_ROUTER_HOST,LLAMA_ROUTER_MODELS_MAX(default2= both resident;1= hot-swap). - Clients pick the model via the OpenAI
modelfield, e.g.qwen3.8orminicpm5-2b. pi: setllamaServerUrltohttp://<host>:8888; little-coder: setbaseUrltohttp://<host>:8888/v1. - LAN clients that resolve
<host>.localto an IPv6 link-local address are covered by the dual-stack (::) bind; if a client still fails, use the IPv4 literal (e.g.http://192.168.0.141:8888).
start-llm-qwen-3.8.sh remains a standalone single-model launcher with KV-cache
save/restore (takes an optional port arg); the router does not use it.