R4ZZ3/flux2-webgpu

★ 0Forks 0TypeScriptGitHub ↗Compare

README

FLUX.2 [klein] 4B — WebGPU

Text-to-image in the browser on WebGPU. A typed prompt produces a matching image, with no ONNX and no transformers.js — the tokenizer, both models, every kernel and the sampler are written here against raw navigator.gpu.

A typed prompt and the image it produced

Prompt: "a photograph of a mountain lake at sunrise" — 448x256 in 73 s on an NVIDIA Lovelace GPU.

The pipeline

BPE tokenizer  ->  chat template  ->  Qwen3-4B text encoder
    (vocab.json + merges.txt)          (q4, 27 layers, grouped-query, causal)
                                              |
                                   taps at layers 9 / 18 / 27
                                   concatenated -> [128, 7680]
                                              v
                     FLUX.2 [klein] DiT  (q8, 5 double + 20 single blocks)
                                              |
                                   rectified-flow Euler sampler
                                     (one forward per step, no CFG)
                                              v
                              VAE decoder  ->  canvas

Weights come from gandamu/flux2-klein-4b-qwen3-4b-webgpu (Apache-2.0) and are fetched by HTTP range request, so only the tensors actually needed are downloaded. No weights are redistributed here — this repository is the runtime only, and the URLs are pinned to a commit so an upstream re-upload cannot silently change what a cached client loads.

Licence and attribution

This runtime is Apache-2.0 (see LICENSE). It is an independent implementation; the model weights it loads are the work of others:

  • FLUX.2 [klein] — Black Forest Labs.
  • Qwen3-4B — Alibaba Cloud.
  • The flux2web-1 / qwen3web-1 repack — gandamu, Apache-2.0. The formats are undocumented upstream and were reverse-engineered from the shipped manifests; see PLAN_AND_PROGRESS.md in the platform repo for what was established and how.

Check the upstream model licences before using this for anything beyond experimentation — they govern the weights, not this code.

Running it

npm install
npm run dev      # then open the printed URL

In the page: Load model (~3.6 GB) then Load text encoder (~1.7 GB), type a prompt, Generate.

The first load takes about 20 minutes — the CDN serves at roughly 5 MB/s. After that the weights are cached with the Cache API and a warm load is ~43 s. "Clear cache" frees the ~5.3 GB.

Below the generator is a diagnostics panel: adapter limits, a ranged-read check, the q8 quantization analysis, all 20 WGSL kernels checked against CPU references, and the per-phase numeric gates.

Development

npm test          # 267 offline tests
npm run test:live # network-backed checks against the published bundle
npm run build     # tsc -b && vite build
npm run gen:shaders   # inline shaders/*.wgsl into src/shaders.generated.ts

Shaders live in shaders/ and are inlined into the bundle at build time, so the runtime never fetches them.

How it is verified

Every phase is gated against fixtures the bundle ships for itself:

Phase Gate Result
VAE decoder check.vlat -> check.vdec cos 0.999999
DiT forward check.z -> check.vel cos 0.9886 (threshold 0.999)
Text encoder vs prompts.0.cond pos-0 cos 0.9998
All kernels vs CPU reference 20/20 at cos 1.000000000

Two of those are honest shortfalls rather than passes. The DiT sits at 0.9886 against a 0.999 threshold, and the encoder's later positions at ~0.98; both are consistent with q8/q4 quantization error accumulating across 25-27 blocks, and neither prevents correct images — but neither is proven.

The encoder's aggregate cosine is deliberately not quoted as the headline number. It reads 0.994 for the right prompt but 0.991 for the empty string, because at T = 128 a short prompt leaves ~100 shared padding positions. The per-position probe is the meaningful one: position 0 is the same token for every prompt and attends only to itself, so it exercises the whole stack independent of the prompt.

Layout

shaders/          20 WGSL kernels
src/
  bpe.ts          byte-level BPE tokenizer
  manifest.ts     flux2web-1 parser      qwen3Manifest.ts  qwen3web-1 parser
  loader.ts       ranged shard reads     cachedSource.ts   Cache API persistence
  dequant.ts      q8 decode              dequantQ4.ts      q4 decode
  ops/            model graphs + CPU reference implementations
  gpu/            WebGPU context and the per-model backends
  demo/           the app, the gates and the kernel checks

Each model's graph is written once, over a backend interface, so the CPU reference and the GPU runtime execute the same description of the architecture.

Issues