Text-to-image in the browser on WebGPU. A typed prompt produces a matching
image, with no ONNX and no transformers.js — the tokenizer, both models, every
kernel and the sampler are written here against raw navigator.gpu.
Prompt: "a photograph of a mountain lake at sunrise" — 448x256 in 73 s on an NVIDIA Lovelace GPU.
BPE tokenizer -> chat template -> Qwen3-4B text encoder
(vocab.json + merges.txt) (q4, 27 layers, grouped-query, causal)
|
taps at layers 9 / 18 / 27
concatenated -> [128, 7680]
v
FLUX.2 [klein] DiT (q8, 5 double + 20 single blocks)
|
rectified-flow Euler sampler
(one forward per step, no CFG)
v
VAE decoder -> canvas
Weights come from
gandamu/flux2-klein-4b-qwen3-4b-webgpu
(Apache-2.0) and are fetched by HTTP range request, so only the tensors
actually needed are downloaded. No weights are redistributed here — this
repository is the runtime only, and the URLs are pinned to a commit so an
upstream re-upload cannot silently change what a cached client loads.
This runtime is Apache-2.0 (see LICENSE). It is an independent
implementation; the model weights it loads are the work of others:
- FLUX.2 [klein] — Black Forest Labs.
- Qwen3-4B — Alibaba Cloud.
- The
flux2web-1/qwen3web-1repack — gandamu, Apache-2.0. The formats are undocumented upstream and were reverse-engineered from the shipped manifests; seePLAN_AND_PROGRESS.mdin the platform repo for what was established and how.
Check the upstream model licences before using this for anything beyond experimentation — they govern the weights, not this code.
npm install
npm run dev # then open the printed URLIn the page: Load model (~3.6 GB) then Load text encoder (~1.7 GB), type a prompt, Generate.
The first load takes about 20 minutes — the CDN serves at roughly 5 MB/s. After that the weights are cached with the Cache API and a warm load is ~43 s. "Clear cache" frees the ~5.3 GB.
Below the generator is a diagnostics panel: adapter limits, a ranged-read check, the q8 quantization analysis, all 20 WGSL kernels checked against CPU references, and the per-phase numeric gates.
npm test # 267 offline tests
npm run test:live # network-backed checks against the published bundle
npm run build # tsc -b && vite build
npm run gen:shaders # inline shaders/*.wgsl into src/shaders.generated.tsShaders live in shaders/ and are inlined into the bundle at build time, so the
runtime never fetches them.
Every phase is gated against fixtures the bundle ships for itself:
| Phase | Gate | Result |
|---|---|---|
| VAE decoder | check.vlat -> check.vdec |
cos 0.999999 |
| DiT forward | check.z -> check.vel |
cos 0.9886 (threshold 0.999) |
| Text encoder | vs prompts.0.cond |
pos-0 cos 0.9998 |
| All kernels | vs CPU reference | 20/20 at cos 1.000000000 |
Two of those are honest shortfalls rather than passes. The DiT sits at 0.9886 against a 0.999 threshold, and the encoder's later positions at ~0.98; both are consistent with q8/q4 quantization error accumulating across 25-27 blocks, and neither prevents correct images — but neither is proven.
The encoder's aggregate cosine is deliberately not quoted as the headline number. It reads 0.994 for the right prompt but 0.991 for the empty string, because at T = 128 a short prompt leaves ~100 shared padding positions. The per-position probe is the meaningful one: position 0 is the same token for every prompt and attends only to itself, so it exercises the whole stack independent of the prompt.
shaders/ 20 WGSL kernels
src/
bpe.ts byte-level BPE tokenizer
manifest.ts flux2web-1 parser qwen3Manifest.ts qwen3web-1 parser
loader.ts ranged shard reads cachedSource.ts Cache API persistence
dequant.ts q8 decode dequantQ4.ts q4 decode
ops/ model graphs + CPU reference implementations
gpu/ WebGPU context and the per-model backends
demo/ the app, the gates and the kernel checks
Each model's graph is written once, over a backend interface, so the CPU reference and the GPU runtime execute the same description of the architecture.
