This repository is the Bazel builder and deployment recipe site for local-inference-lab/vLLM on NVIDIA DGX Spark. It packages that fork and its source-built dependencies into an Open Container Initiative (OCI) image. It is not the vLLM fork itself, an upstream vllm-project/vllm image, or a general-purpose builder for arbitrary vLLM releases.
The local build path targets DGX Spark. Bazel compiles PyTorch and CUDA extensions from source, then assembles their wheels into layered Python runtimes.
The implemented image is a multiarchitecture index targeting Linux ARM64 and Linux x86-64, CUDA 13.4.1, and Python 3.12. CI builds both architectures and runs the image contract for each. The ARM64 manifest carries the production two-node deployment evidence; the x86-64 manifest is verified by its CI contract plus a single-GPU serving smoke on an RTX 5090 (see Latest image publication). Publication requires exactly the two manifests together: the publisher and the recipe resolver reject an index missing either architecture.
The image is built from pinned fork revisions rather than from upstream branches, so this table records which upstream changes the current lock carries. Regenerate it with scripts/vllmb12x-included-changes.py.
| Component | Change | Included as |
|---|---|---|
vLLM 0a6739845a32 |
Native MXFP8 MTP through the ModelOpt B12X backend (no upstream pull request) | Merged into the pinned revision (4 commits) |
vLLM 0a6739845a32 |
Partial port of vLLM PR 779: Qwen HC ownership and checkpoint coalescing (no upstream pull request) | Merged into the pinned revision (9 commits) |
vLLM 0a6739845a32 |
local-inference-lab/vllm#800 Bounded shared-memory broadcast waits | Merged into the pinned revision (1 commits) |
vLLM 0a6739845a32 |
Rewrite flash_attn.cute imports to vllm.vllm_flash_attn.cute (no upstream pull request) | Applied at build time by third_party/vllm_flash_attn_cute_namespace.patch |
B12X 965e748f74fe |
Native block-scaled MXFP8 W8A8 MoE execution (no upstream pull request) | Merged into the pinned revision (3 commits) |
B12X 965e748f74fe |
PLE checkpoint exports, prepared launcher closures and collective entry barriers (no upstream pull request) | Merged into the pinned revision (11 commits) |
Start with Building Locally on DGX Spark. The current build requires GCC 15, compatible userspace, and persistent writable standard ccache storage. CI replaces this with its durable cache volume. Stock DGX OS is not automatically compatible. A full local Spark build has not yet been verified.
With the prerequisites satisfied, use the Justfile helpers:
just analyze
just build
just test
just loadanalyze resolves the graph without compiling. build produces the image, test checks its structure, and load imports it into Docker without starting a container. Cold source builds can take hours.
The public deployment recipe site provides the tested DeepSeek configuration, generated Kubernetes and Docker instructions, and the guided cluster setup path.
| Need | Document |
|---|---|
| Learn the build graph | Inspect Your First Build |
| Prepare a Spark build environment | Build Locally on DGX Spark |
| Build, test, and load | Build and Test an Image |
| Measure a native edit | Measure Incremental Builds |
| Use model, rendezvous, and health helpers | Use the Kubernetes Image Helpers |
| Run a pod as a pinned non-root uid | Run the Image as a Non-Root User |
| Look up targets and options | Build Configuration |
| Configure the Mooncake runtime addition | Mooncake Transfer Engine |
| Understand build boundaries | Build and Cache Design |
NativeLink is our optional CI execution backend, not a prerequisite for local builds. CI access and deployment configuration are private and are not distributed here.
Repository code is licensed under Apache License 2.0, consistent with existing SPDX headers. Upstream sources, Python packages, CUDA redistributables, and base-image contents retain their respective licences and redistribution requirements.