This repository is a small source snapshot for our KernelGen Challenge 9 Task 3 implementation, whose required entry point is:
def silu_dot_fwd_bwd_quant_fuse(
x,
grad_y,
grad_input_q,
grad_input_s,
y_q_t,
y_s_t,
group_size=128,
):
...Official challenge page:
https://kernelgen.flagos.io/challenge/9?lang=zh&tab=readme
The implementation recomputes the SwiGLU forward value, computes the backward gradients, and writes two INT8-quantized products:
grad_input_q/grad_input_s: row-wise 128-channel quantization of[d_gate, d_up].y_q_t/y_s_t: transposed 128-token-group quantization ofy = silu(gate) * up.
The main solution.py file is exported from our stable submitted source:
| Item | Value |
|---|---|
| Submission | sub_e8229df252bc |
| Source commit | b13093a |
| Platform pass count | 7/7 |
| Average speedup | about 10.46x |
| Date context | June 2026 offline competition run |
Per-platform scores recorded in notes/submission-log.md:
| Platform | Speedup |
|---|---|
| Ascend | 1.14 |
| Hygon / AMD | 11.18 |
| MetaX | 10.33 |
| MooreThreads | 8.86 |
| NVIDIA | 13.57 |
| T-Head | 14.36 |
| TianShu | 13.76 |
solution.py: the competition implementation.reference.py: a PyTorch reference used for local sanity checks.tests/: local correctness tests over the official benchmark shapes.notes/: submission log and optimization notes preserved for context.pyproject.tomlanduv.lock: local Python environment metadata.
The solution.py snapshot is fixed to the submitted commit above. The reference/test helper files are included so that the repository can be read and checked locally.
The optimization is built around a few simple constraints:
- Use
128x128tiles because both quantization axes use groups of 128. - Fuse forward recomputation, backward computation, M-group quantization, and K-group quantization where the backend can handle it.
- Preserve the BF16 roundtrip before quantization; skipping it can change scales and INT8 values.
- Replace inner elementwise division with multiplication by precomputed inverse scale.
- Split backend routes because Triton codegen, cache behavior, warp size, and register pressure differ across the target chips.
Please read this as a competition artifact, not as a general-purpose training kernel library.
- It is tuned for the KernelGen Challenge 9 interface and official shape set.
- Platform detection is intentionally pragmatic and may not cover other devices.
- Some routes are chosen for the competition backends available at the time.
- Correctness depends on preserving the benchmark's BF16 and INT8 quantization semantics.
- Local tests require a CUDA-compatible or competition-like accelerator environment; CPU-only runs are not representative.
The repository keeps uv metadata for the lightweight Python test dependencies. Torch and Triton wheels are deliberately not pinned here because they should match the target accelerator environment.
In an environment that already has compatible torch and triton installed, a local sanity check is:
uv sync
uv run pytestThese tests compare dequantized INT8 outputs and scale tensors against the PyTorch reference under the tolerances used by the competition.