ashhart/TensorFold
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
Independent AI systems architect building local-first AI tools and governed inference infrastructure.
Fast, exact LLM decoding on Apple Silicon (MLX) behind an OpenAI-compatible endpoint
Metal Cuda Direct Memory Access
Homebrew tap for TensorFold: brew install ashhart/tensorfold/tensorfold
SparkPilot puts your NVIDIA DGX Spark to work, free forever.
Speeds up TTFT
OpenBot - A reimplementation of GrokBot
A small script to stop your DGX Spark OOMing
Use computer use with tools like OMP or OpenCode.
We gave AI models telepathy.
Decision engine
Local AI Chrome Plugin for AI Image Detection.
MLX: An array framework for Apple silicon
A lightweight runtime that helps a small local model reason, discover, and complete work like a frontier model.
Use Codex with local models.
Two models, one workspace: an OMP plugin for shared tasks and Agent Hub conversations.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Qwen 3.8 Flash Next 4bit MTP running at just under 100tks decode
Benchmarking Alex Ziskind's Silicone Exchange with DGX Sparks
AshHart
LLMs and VLMs with MLX Swift
Swift API for MLX
Private Inference Network on Idle Macs
DeepSeekV4-Flash-0731-MXFP4-MLX Apple Silicone