ericjlake/llama.cpp
LLM inference in C/C++
LLM inference in C/C++
Swift API for MLX
LLMs and VLMs with MLX Swift
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, + iOS iPhone app.