TankTechnology/kernelgen-challenge-9-task3-silu-quant

Stable source snapshot for KernelGen Challenge 9 Task 3

★ 1Forks 0PythonGitHub ↗Compare

README

KernelGen Challenge 9 Task 3: SwiGLU Backward + INT8 Quantization

This repository is a small source snapshot for our KernelGen Challenge 9 Task 3 implementation, whose required entry point is:

def silu_dot_fwd_bwd_quant_fuse(
    x,
    grad_y,
    grad_input_q,
    grad_input_s,
    y_q_t,
    y_s_t,
    group_size=128,
):
    ...

Official challenge page:

https://kernelgen.flagos.io/challenge/9?lang=zh&tab=readme

The implementation recomputes the SwiGLU forward value, computes the backward gradients, and writes two INT8-quantized products:

  • grad_input_q / grad_input_s: row-wise 128-channel quantization of [d_gate, d_up].
  • y_q_t / y_s_t: transposed 128-token-group quantization of y = silu(gate) * up.

Snapshot

The main solution.py file is exported from our stable submitted source:

Item Value
Submission sub_e8229df252bc
Source commit b13093a
Platform pass count 7/7
Average speedup about 10.46x
Date context June 2026 offline competition run

Per-platform scores recorded in notes/submission-log.md:

Platform Speedup
Ascend 1.14
Hygon / AMD 11.18
MetaX 10.33
MooreThreads 8.86
NVIDIA 13.57
T-Head 14.36
TianShu 13.76

What Is Included

  • solution.py: the competition implementation.
  • reference.py: a PyTorch reference used for local sanity checks.
  • tests/: local correctness tests over the official benchmark shapes.
  • notes/: submission log and optimization notes preserved for context.
  • pyproject.toml and uv.lock: local Python environment metadata.

The solution.py snapshot is fixed to the submitted commit above. The reference/test helper files are included so that the repository can be read and checked locally.

Main Ideas

The optimization is built around a few simple constraints:

  • Use 128x128 tiles because both quantization axes use groups of 128.
  • Fuse forward recomputation, backward computation, M-group quantization, and K-group quantization where the backend can handle it.
  • Preserve the BF16 roundtrip before quantization; skipping it can change scales and INT8 values.
  • Replace inner elementwise division with multiplication by precomputed inverse scale.
  • Split backend routes because Triton codegen, cache behavior, warp size, and register pressure differ across the target chips.

Caveats

Please read this as a competition artifact, not as a general-purpose training kernel library.

  • It is tuned for the KernelGen Challenge 9 interface and official shape set.
  • Platform detection is intentionally pragmatic and may not cover other devices.
  • Some routes are chosen for the competition backends available at the time.
  • Correctness depends on preserving the benchmark's BF16 and INT8 quantization semantics.
  • Local tests require a CUDA-compatible or competition-like accelerator environment; CPU-only runs are not representative.

Local Check

The repository keeps uv metadata for the lightweight Python test dependencies. Torch and Triton wheels are deliberately not pinned here because they should match the target accelerator environment.

In an environment that already has compatible torch and triton installed, a local sanity check is:

uv sync
uv run pytest

These tests compare dequantized INT8 outputs and scale tensors against the PyTorch reference under the tolerances used by the competition.

Contributors

TankTechnology

Issues