christopherowen/spark3-vllm-ds41f
Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric.
Reproducible Docker/vLLM deployment, tuning, and benchmarks for DeepSeek V4.1 Flash on a switchless three-node DGX Spark fabric.
Userland fan control for NVIDIA DGX Spark, with additive RPM floors, a temperature-based curve, Secure Boot signing, and DKMS support.
Power telemetry and NVIDIA-bounded power limits for NVIDIA DGX Spark, via the firmware's SPBM page, with Secure Boot signing and DKMS support.
🚀 Run Codex App UI Anywhere: Linux, Windows, or Termux on Android 🚀
Tile primitives for speedy kernels
Docker configuration for running VLLM on dual DGX Sparks
llama-benchy - llama-bench style benchmarking tool for all backends
CUDA Templates and Python DSLs for High-Performance Linear Algebra
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine for LLMs