yohann-bearzi/omlx
oMLX fork: GLM-5.2 JANGTQ loading + steel-tiled TurboQuant MoE kernel (3.8x prefill on 2-bit experts)
oMLX fork: GLM-5.2 JANGTQ loading + steel-tiled TurboQuant MoE kernel (3.8x prefill on 2-bit experts)
jang-tools fork: GLM-5.2 fused sparse-MLA prefill + Hadamard/gather kernel fixes
Donkey — ANE drafter: world model predicting latent-space sequences for speculative decode. Package: donkey-mlx.
Standalone MLX extension: block_fp8 (DeepSeek-V3/MiMo-style E4M3) quantized matmul and MoE kernels. Builds against stock upstream MLX.
Mestra — continual per-expert recompressor for FP8 MoE weights on Apple Silicon. Change the form, keep the soul.
MLX: An array framework for Apple silicon
Own your AI. The native macOS harness for AI agents -- any model, persistent memory, autonomous execution, cryptographic identity. Built in Swift. Fully offline. Open source.
vMLX Swift Engine.
vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!