taagarwa-rh/llama-swap-config

★ 2Forks 0ShellGitHub ↗Compare

README

Llama Swap Config

My personal llama-swap configuration. This is meant to replace my Ollama configuration, so I use the same port (11434) for serving.

System Info

> system_profiler SPDisplaysDataType

Graphics/Displays:

    Apple M4 Pro:

      Chipset Model: Apple M4 Pro
      Type: GPU
      Bus: Built-In
      Total Number of Cores: 20
      Vendor: Apple (0x106b)
      Metal Support: Metal 4

Setup

  1. Install llama.cpp with homebrew

    brew install llama.cpp
  2. Install llama-swap with homebrew

    brew tap mostlygeek/llama-swap
    brew install llama-swap

Usage

Start the llama-swap server on port 11434

chmod a+x start.sh
./start.sh

List available and running models (only available while server is running.)

chmod a+x list.sh
./list.sh

Stop the server

chmod a+x stop.sh
./stop.sh

I've also added the following to my .zshrc file to start the server from anywhere

# Llama-swap (update the path to where you cloned the repo)
LLAMA_SWAP_DIR="$HOME/path/to/llama-swap-config"
alias llm-start="$LLAMA_SWAP_DIR/start.sh"
alias llm-stop="$LLAMA_SWAP_DIR/stop.sh"
alias llm-list="$LLAMA_SWAP_DIR/list.sh"
alias llm-config="code $LLAMA_SWAP_DIR/config.yaml"

Which enables you to run llm-start to start the server, llm-list to list available models, and llm-stop to stop the server.

Run your own models

  1. Create a configuration file for llama-swap, e.g.

    models:
      gpt-oss-20b:
        cmd: llama-server --port ${PORT} -hf openai/gpt-oss-20b

    If you downloaded a model, use the path to the .gguf file or model folder instead

    models:
      model1:
        cmd: llama-server --port ${PORT} --model /path/to/model.gguf
  2. Run your model command once to test/download your model. Specify an open port in your test

    llama-server --port 1111 -hf openai/gpt-oss-20b
  3. Run

    llama-swap --config path/to/config.yaml --listen localhost:8080

LiteLLM Setup

LiteLLM proxy sits on top of llama-swap (port 11434) and adds cloud model routing (Claude via Vertex AI).

Prerequisites

Ensure you have logged in to your Google Cloud project and the following environment variables are set:

export ANTHROPIC_VERTEX_PROJECT_ID=
export OPENAI_API_KEY=

Installation

# Install globally
uv tool install 'litellm[cli,proxy,google]'

# Install in a virtual environment
uv pip install 'litellm[cli,proxy,google]'

This will add the lite, litellm, and litellm-proxy CLIs to your machine.

Running the Proxy

Start the proxy (no Docker needed):

chmod a+x litellm-start.sh
./litellm-start.sh

Or run directly:

uv run litellm --config litellm_config.yaml --port 4000

Configuration

Your litellm_config.yaml defines:

  • qwen3.6-35b-coder → Routes to your local llama-swap at localhost:11434
  • claude-opus-4-6 → Routes to Vertex AI (Claude)
  • gpt-5.6-sol → Routes to OpenAI's GPT-5.6 model via API key
  • smart-router → Complexity-based auto-routing between local and cloud models

Thinking Block Cleanup (strip_thinking.py)

Claude models with extended thinking return "thinking blocks" in their responses. In multi-turn conversations, these blocks can end up back in the request history with their content stripped (by the client or proxy for token savings). The Anthropic API rejects empty thinking blocks with:

messages.N.content.0.thinking: each thinking block must contain thinking

This is especially common when the smart router switches between local and cloud models mid-conversation. The strip_thinking.py callback hooks into LiteLLM's async_pre_call_hook to remove only empty thinking blocks from the conversation history before sending the request — valid thinking blocks with content are preserved.

Using the Proxy

Once running on port 4000, use it as your OpenAI-compatible endpoint:

# Test with curl
curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer llm-swap-secret-key" \
  -d '{
    "model": "smart-router",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Using the lite CLI

The lite CLI requires a master key. The --skip-verify flag goes after the subcommand:

# Run Claude Code through the proxy
lite --api-key llm-swap-secret-key claude --skip-verify --model claude-opus-4-6 "Your prompt here"

# Run interactive chat
lite --api-key llm-swap-secret-key chat --skip-verify --model smart-router

# Use a specific local model
lite --api-key llm-swap-secret-key chat --skip-verify --model qwen3.6-35b-coder

Optional: Add Aliases

Add these to your ~/.zshrc:

# LiteLLM (update the path to where you cloned the repo)
LLAMA_SWAP_DIR="$HOME/path/to/llama-swap-config"
export LITELLM_PROXY_API_KEY="llm-swap-secret-key"
alias litellm-start="$LLAMA_SWAP_DIR/litellm-start.sh"
alias litellm-stop="$LLAMA_SWAP_DIR/litellm-stop.sh"
alias lite-chat='lite --api-key $LITELLM_PROXY_API_KEY chat --skip-verify'
alias lite-claude='export ANTHROPIC_MODEL='smart-router'; unset CLAUDE_CODE_USE_VERTEX; unset ANTHROPIC_SMALL_FAST_MODEL; lite --api-key $LITELLM_PROXY_API_KEY claude --skip-verify'

Then you can use Claude Code with the smart-router:

lite-claude --model smart-router

Contributors

taagarwa-rh

Issues