DOCS · MILESTONES

Red Lite dev35 — long-context decode: grouped split-K attention

Status: done on branch dev/long-decode (from dev/bounded-prefill at 40e4fc0). Apple M4 Max 48 GiB only.

Finding#

Decode slows with position because of attention: at ~8400 positions the stage took 252.8 ms per 64 tokens (~4 ms/token, RL_ENGINE_PROFILE=1), against a floor of ~1 ms/token for reading the 12 layers' KV cache once (~413 MB at ~400 GB/s). The dev26 split-K kernel ran one threadgroup per (256-position block, query head): the 8 query heads sharing a KV head each read the same K and V rows.

Change#

attn_gqa_split_g: one threadgroup per (block, KV head) computes all its query heads (scores: one thread per position, float4 K loads, 8 accumulators; values: one thread per dimension, 8 accumulators); the merge kernel is shared and now takes the block length. RL_ENGINE_ATTN_GROUP=0 restores the dev26 kernel; RL_ENGINE_ATTN_BLK sets the block.

Attention stage at ~8400 positions (IQ2_XXS, full residency, 64 tokens, single profiled runs):

kernel ms
dev26 per query head, blocks of 256 252.8
grouped, blocks of 256 164.3
grouped, blocks of 128 (default) 112.8
grouped, blocks of 64 / 32 146.5 / 180.2
grouped, 4 threads per position (blocks of 256) 175.1 (reverted)

Decode, redlite-generate --raw on the fixture text repeated to P tokens, 128 tokens, full residency, cooled alternating A/B (3 runs each):

position dev26 kernel grouped runs
~4096 66.53 tok/s 67.79 tok/s 68.25/65.96/66.53 vs 67.32/67.79/69.21
~8192 57.80 tok/s 65.95 tok/s 59.06/57.68/57.80 vs 65.91/66.04/65.95

Parity#

Kernel self-test: the grouped kernel vs a double-precision GQA at 6 lengths (worst abs 7.6e-8). 1100-position context vs llama.cpp unchanged (IQ2_XXS 0.99 / KL 3.5e-3, IQ3_XXS 0.57 / KL 3.9e-4, full and 4 GiB); long_positions.sh: 4096 PASS (2.30 / KL 7.2e-3), 8192 PASS (4.48 / KL 2.2e-6).

Validation#

rm -rf .deps/redmetal, make redmetal, make native, make sanitize: 0 warnings; make test and ruff green; scripts/regress_m4.sh: IQ2_XXS 49/49, IQ3_XXS 36 passed, 0 failed, 13 skipped (the IQ2_XXS-layout dense stage tools).

Scope boundary#

  • M4 Max 48 GiB only; the same kernel serves the 4 GiB path (not re-benchmarked there).
  • Requires head_dim 256 and at most 8 query heads per KV head (Qwen3-Next: 16 / 2); otherwise the dev26 kernel runs.

SOURCE docs/REDLITE_DEV35_LONG_DECODE.md · updated 2026-10-01 · EDIT ON GITHUB