Run OpenAI-compatible LLM servers locally with Tum and MLX.
- Python 3.10+
- macOS with Apple Silicon for MLX model execution
uvfor local development commands
Install from PyPI:
pip install tumThen run:
tum --helpList bundled model IDs:
tum modelsFilter by name:
tum models --query Qwen --limit 10Show small or big model catalogs:
tum models --size small --query Qwen
tum models --size big --query LlamaUse --all to print every match, or --pager for long output.
tum mlx \
--model mlx-community/Qwen2.5-0.5B-Instruct-4bit \
--prompt "What is the capital of France?" \
--max-tokens 128tum serve \
--model mlx-community/Qwen2.5-0.5B-Instruct-4bit \
--host 127.0.0.1 \
--port 8080Then call the chat completions endpoint:
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Qwen2.5-0.5B-Instruct-4bit",
"messages": [{"role": "user", "content": "Say hello"}]
}'Stream tokens as server-sent events:
curl -N http://127.0.0.1:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Qwen2.5-0.5B-Instruct-4bit",
"messages": [{"role": "user", "content": "Say hello"}],
"stream": true
}'Supported endpoints include:
GET /healthGET /v1/modelsPOST /v1/chat/completionsPOST /chat/completionsPOST /v1/completions
Run tests:
uv run pytestRun lint:
uv run ruff check tum tests