Tom Turney

@TheTom · User

GitHub profile ↗ · Compare

Working on LLM inference systems, KV cache compression, and kernel-level optimizations (TurboQuant).

Texas603 followers46 repositories

Repositories

TheTom/vllm-swift

vLLM Metal plugin powered by mlx-swift — high-performance LLM inference on Apple Silicon

★ 274PythonForks 19

TheTom/pascal-egpu

Driverless NVIDIA Pascal (GTX 1060) compute from macOS Apple Silicon over Thunderbolt eGPU

★ 5PythonForks 3

TheTom/vllm-turboquant

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 7PythonForks 1

TheTom/melee

SSBM GALE01 decomp progress fork (doldecomp/melee). Provide your own main.dol.

★ 0CForks 0

TheTom/model-serving-minefield

Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.

★ 0Forks 0

TheTom/melee-decomp

A decompilation of Super Smash Bros Melee brought to you by a bunch of clever folks.

★ 0CForks 0

TheTom/ffai

FFAI — F*cking Fast AI. Dual-engine inference: Swift (Apple/iPhone, native Metal) + Rust (cross-platform: CUDA/Vulkan/ROCm). Shared kernels via metaltile.

★ 2SwiftForks 0

TheTom/momento

Deterministic State Recovery for AI Coding Agents

★ 4PythonForks 1

TheTom/mlx

MLX: An array framework for Apple silicon

★ 5Forks 0

TheTom/pastors-pocket-spurgeon

Offline Spurgeon study companion — pastoral counsel, sermon prep, and Sword & Trowel sermon grading. Gemma-4-12B fine-tune served on TurboQuant llama.cpp.

★ 3PythonForks 0

TheTom/metaltile

A Rust-embedded DSL for writing Apple Metal GPU kernels. Write tile-level algorithms in Rust, get optimized Metal Shading Language out.

★ 1MetalForks 0

TheTom/ds4-turboquant_plus

Private backup of TQ+ Metal port work on antirez/ds4. Branch tom/turbo3-kv-cuda (turbo2/3/4 + Wave M3 inline-dequant + half-tile + CUDA stubs + h8 Flash WIP).

★ 1CForks 0

TheTom/ds4

DeepSeek 4 Flash local inference engine for Metal and CUDA

★ 0Forks 0

TheTom/longctx

Open long-context inference stack: retrieval + open weights, no closed parts. pip install longctx.

★ 7PythonForks 2

TheTom/tqkit

Unified toolkit for benchmarking and integrating TurboQuant+ KV-cache compression across llama.cpp, vLLM, MLX, and vllm-swift.

★ 1PythonForks 1

TheTom/vllm-metal

Community maintained hardware plugin for vLLM on Apple Silicon

★ 1Forks 0

TheTom/cs161

StemVille real-time data plotting experiment

★ 2JavaScriptForks 0