My personal llama-swap configuration.
This is meant to replace my Ollama configuration, so I use the same port (11434) for serving.
> system_profiler SPDisplaysDataType
Graphics/Displays:
Apple M4 Pro:
Chipset Model: Apple M4 Pro
Type: GPU
Bus: Built-In
Total Number of Cores: 20
Vendor: Apple (0x106b)
Metal Support: Metal 4
-
Install
llama.cppwith homebrewbrew install llama.cpp
-
Install
llama-swapwith homebrewbrew tap mostlygeek/llama-swap brew install llama-swap
Start the llama-swap server on port 11434
chmod a+x start.sh
./start.shList available and running models (only available while server is running.)
chmod a+x list.sh
./list.shStop the server
chmod a+x stop.sh
./stop.shI've also added the following to my .zshrc file to start the server from anywhere
# Llama-swap (update the path to where you cloned the repo)
LLAMA_SWAP_DIR="$HOME/path/to/llama-swap-config"
alias llm-start="$LLAMA_SWAP_DIR/start.sh"
alias llm-stop="$LLAMA_SWAP_DIR/stop.sh"
alias llm-list="$LLAMA_SWAP_DIR/list.sh"
alias llm-config="code $LLAMA_SWAP_DIR/config.yaml"Which enables you to run llm-start to start the server, llm-list to list available models, and llm-stop to stop the server.
-
Create a configuration file for
llama-swap, e.g.models: gpt-oss-20b: cmd: llama-server --port ${PORT} -hf openai/gpt-oss-20b
If you downloaded a model, use the path to the
.gguffile or model folder insteadmodels: model1: cmd: llama-server --port ${PORT} --model /path/to/model.gguf
-
Run your model command once to test/download your model. Specify an open port in your test
llama-server --port 1111 -hf openai/gpt-oss-20b
-
Run
llama-swap --config path/to/config.yaml --listen localhost:8080
LiteLLM proxy sits on top of llama-swap (port 11434) and adds cloud model routing (Claude via Vertex AI).
Ensure you have logged in to your Google Cloud project and the following environment variables are set:
export ANTHROPIC_VERTEX_PROJECT_ID=
export OPENAI_API_KEY=# Install globally
uv tool install 'litellm[cli,proxy,google]'
# Install in a virtual environment
uv pip install 'litellm[cli,proxy,google]'This will add the lite, litellm, and litellm-proxy CLIs to your machine.
Start the proxy (no Docker needed):
chmod a+x litellm-start.sh
./litellm-start.shOr run directly:
uv run litellm --config litellm_config.yaml --port 4000Your litellm_config.yaml defines:
qwen3.6-35b-coder→ Routes to your local llama-swap atlocalhost:11434claude-opus-4-6→ Routes to Vertex AI (Claude)gpt-5.6-sol→ Routes to OpenAI's GPT-5.6 model via API keysmart-router→ Complexity-based auto-routing between local and cloud models
Claude models with extended thinking return "thinking blocks" in their responses. In multi-turn conversations, these blocks can end up back in the request history with their content stripped (by the client or proxy for token savings). The Anthropic API rejects empty thinking blocks with:
messages.N.content.0.thinking: each thinking block must contain thinking
This is especially common when the smart router switches between local and cloud models mid-conversation. The strip_thinking.py callback hooks into LiteLLM's async_pre_call_hook to remove only empty thinking blocks from the conversation history before sending the request — valid thinking blocks with content are preserved.
Once running on port 4000, use it as your OpenAI-compatible endpoint:
# Test with curl
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer llm-swap-secret-key" \
-d '{
"model": "smart-router",
"messages": [{"role": "user", "content": "Hello!"}]
}'The lite CLI requires a master key. The --skip-verify flag goes after the subcommand:
# Run Claude Code through the proxy
lite --api-key llm-swap-secret-key claude --skip-verify --model claude-opus-4-6 "Your prompt here"
# Run interactive chat
lite --api-key llm-swap-secret-key chat --skip-verify --model smart-router
# Use a specific local model
lite --api-key llm-swap-secret-key chat --skip-verify --model qwen3.6-35b-coderAdd these to your ~/.zshrc:
# LiteLLM (update the path to where you cloned the repo)
LLAMA_SWAP_DIR="$HOME/path/to/llama-swap-config"
export LITELLM_PROXY_API_KEY="llm-swap-secret-key"
alias litellm-start="$LLAMA_SWAP_DIR/litellm-start.sh"
alias litellm-stop="$LLAMA_SWAP_DIR/litellm-stop.sh"
alias lite-chat='lite --api-key $LITELLM_PROXY_API_KEY chat --skip-verify'
alias lite-claude='export ANTHROPIC_MODEL='smart-router'; unset CLAUDE_CODE_USE_VERTEX; unset ANTHROPIC_SMALL_FAST_MODEL; lite --api-key $LITELLM_PROXY_API_KEY claude --skip-verify'Then you can use Claude Code with the smart-router:
lite-claude --model smart-router