Rectangular mosaic texture in dense fur at 1344x768 is h3.c-specific (independent impl renders clean strands)

#70 · open · 1 comments

View on GitHub ↗

dominicletz

## Human Preamble I've experienced a very heavy mosaic structure on the generated videos using h3.c on my hardware. I'm surprised that nobody else reported this yet, as I couldn't really find a work-around configuration wise. I wonder if this is related to my Apple hardware or my copy of the model? Every video has those. I've generated below report on this and attached a sample side-by-side of the issue. ## Summary Dense fox-fur regions at 1344x768 render with a rectangular mosaic texture: squarish dark blotches with hard axis-aligned edges instead of coherent fur strands. It is present in the **raw pre-encode output** (`--frames-dir` PPM), so it is not x264. A second, independent Metal implementation (vpipe, own kernels, own VAE decode) renders the same prompt/geometry/seed with clean directional fur strands — so this looks h3.c-specific (DiT kernels or the video-VAE decode path), not model-inherent. ## Repro - Commit: `8974cc0` (predates PR #63) - Device: Apple M5 Pro, 48 GB, Metal 4 (`h3 --info`: applegpu_g17s, GPU family 10) - Model: MiniMax-H3 FL2VA, `-d ~/models/MiniMax-H3` - Must run with cwd at the shader dir (h3 loads `h3_shaders.metal` from cwd) ```sh cd ~/src/h3.c && ./h3 --profile -d ~/models/MiniMax-H3 \ -p "A red fox walks through fresh snow in a pine forest. Medium tracking shot, natural winter light, realistic fur, soft footsteps and wind." \ --width 1344 --height 768 --frames 22 --seed 42 \ --steps 20 --layers 50 --reuse 1 --use-slower-bf16-mlp \ --frames-dir /tmp/raw-frames -o /tmp/fox.mp4 ``` Look at the fox tail in any raw frame (e.g. frame 21): the mid-tail fur breaks into discrete dark squarish spots with visible block edges, roughly 15-30 px, instead of strands. Also visible (weaker) in `exact-default` and `balanced` configs. ## Ruled out - **Denoising steps**: identical mosaic at 4 / 20 / 50 steps (50-step sample kept). - **Precision**: int8 default and BF16-MLP (`--use-slower-bf16-mlp`) show identical structure. - **Encoder**: raw PPM frames show it; the old x264 defaults (fast/CRF 18) added *extra* blocking on top, fixed separately by encoding slow/CRF 14 — the underlying mosaic remains. - **VAE tile seams**: documented tiles are 256-320 px with 64 px overlap, an order of magnitude coarser than these blocks. - **Model**: vpipe cross-check below. ## Cross-check: vpipe renders clean fur vpipe (`tgo-app-dev/vpipe`, independent Metal kernels, own MiniMax-H3 loader/VAE decode) at the same prompt, 1344x768, 22 frames, seed 42, 16 steps, DiT + text encoder quantized to 8-bit group-64 **from the same local weights** (registered the diffusers dir directly, skipped the Comfy repack download). Its raw pre-encode PNGs show coherent directional fur strands with smooth tonal gradation in the same tail region — no blocks. Note vpipe runs *lower* effective precision (8-bit + lossy i8_gemm) than the h3.c BF16-MLP run that shows the mosaic, so precision cannot explain the difference either. (Poses differ between the two implementations, so this compares texture character, not pixels.) ## Suspects Remaining candidates on the h3.c side: DiT attention/MLP Metal kernels (note this checkout predates the PR #63 GQA-barrier fix — I also observe same-seed output variance consistent with issue #52) or the video-VAE decoder path at 768p. The related texture artifacts already documented in the README (woven texture in rejected schedules, 256 lattice fixed via RoPE halving) suggest this could be another member of that family.

Comments

dominicletz

Side-by-side, as promised — same tail region, raw pre-encode frames, full-res crop on top with 2x nearest-neighbor zoom below: ![h3.c (left, blocks) vs vpipe (right, strands), raw tail crops](https://files.catbox.moe/sudyr5.png) Left: h3.c `bf16-mlp` frame 21 — mid-tail fur breaks into discrete squarish blotches with axis-aligned edges. Right: vpipe 8-bit frame 21 — coherent directional strands, smooth gradation. (Poses differ between implementations, so compare texture character, not pixels. Full frames + clips available on request.)