TheTom/vllm-swift
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
Working on LLM inference systems, KV cache compression, and kernel-level optimizations (TurboQuant).
vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon
The agent that grows with you
LLM inference in C/C++
I'm crazy and trying to make a ForScan OBD reader work on my mac.
Driverless NVIDIA Pascal (GTX 1060) compute from macOS Apple Silicon over Thunderbolt eGPU
Opt-in menu-bar controller for safely using any Mac as a GitHub Actions burst runner
A high-throughput and memory-efficient inference and serving engine for LLMs
SSBM GALE01 decomp progress fork (doldecomp/melee). Provide your own main.dol.
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
A decompilation of Super Smash Bros Melee brought to you by a bunch of clever folks.
FFAI — F*cking Fast AI. Dual-engine inference: Swift (Apple/iPhone, native Metal) + Rust (cross-platform: CUDA/Vulkan/ROCm). Shared kernels via metaltile.
Deterministic State Recovery for AI Coding Agents
MLX: An array framework for Apple silicon
Offline Spurgeon study companion — pastoral counsel, sermon prep, and Sword & Trowel sermon grading. Gemma-4-12B fine-tune served on TurboQuant llama.cpp.
A Rust-embedded DSL for writing Apple Metal GPU kernels. Write tile-level algorithms in Rust, get optimized Metal Shading Language out.
Private backup of TQ+ Metal port work on antirez/ds4. Branch tom/turbo3-kv-cuda (turbo2/3/4 + Wave M3 inline-dequant + half-tile + CUDA stubs + h8 Flash WIP).
DeepSeek 4 Flash local inference engine for Metal and CUDA
Pure Rust Inference Engine
Open long-context inference stack: retrieval + open weights, no closed parts. pip install longctx.
Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across llama.cpp, vLLM, MLX, and vllm-swift.
Homebrew tap for TheTom projects (vllm-swift)
TurboQuant KV cache compression for llama.cpp — HIP/ROCm port for AMD RDNA3 (gfx1100)
Swift API for MLX
LLMs and VLMs with MLX Swift
Community maintained hardware plugin for vLLM on Apple Silicon
StemVille real-time data plotting experiment