DOCS · MILESTONES
Red Lite dev67 — decode attention: the score pass and the merge
Status: done. Measured on the Apple M4 Max 48 GiB (90 W charger) and the Apple M4 Pro 24 GiB (AC), 2026-10-07.
dev66 brought the decode attention to about 415 GB/s on the M4 Max. Two parts of it still left time on the table.
1. The changes#
- The merge.
attn_gqa_mergecombines the per-block partial results of each query head. With 16 query heads it ran 16 threadgroups, each walking every block in turn and waiting on each load: 1,024 blocks at 256K positions. Without the merge the split pass alone took 25.4 ms of 29.6 at 256K (kernel-bench), so the merge cost 4.2 ms.- All threads of a head now compute the maximum over the blocks together (a maximum is exact in any order).
- The sums then run in the same order as before, with four blocks' loads in flight.
- The merge stays bit-identical and now costs about 1.3 ms at 256K.
- The score pass. Each thread read one key row of 1 KiB alone, so the 32 threads of a simdgroup touched 32
different cache lines at a time.
- Now eight threads share a row: each takes every eighth
float4, so together they read 128-byte lines whole. - Three shuffles per query head add the eight parts.
- The summation order changes. The self-test's worst absolute error against double precision is unchanged (7.6e-8).
- Now eight threads share a row: each takes every eighth
2. Measurements#
kernel-bench, M4 Max (12 layers, 12 distinct KV buffer pairs, 0.8.1 → dev67):
| Positions | 0.8.1 | dev67 | Change | Bandwidth |
|---|---|---|---|---|
| 65,536 | 7.55 ms | 6.50 ms | −14 % | 495 GB/s |
| 131,072 | 15.71 ms | 13.66 ms | −13 % | 472 GB/s |
| 262,144 | 31.05 ms | 26.67 ms | −14 % | 483 GB/s |
decode-bench, M4 Max (CF2, every expert resident, random state, median of 16 tokens, alternated):
| Position | 0.8.1 | dev67 | Change |
|---|---|---|---|
| 32,768 | 64.2–70.9 tok/s | 65.7–72.5 tok/s | within noise (two series disagree) |
| 65,536 | 53.3 / 53.5 tok/s | 55.5 / 56.1 tok/s | +4 % |
| 131,072 | 37.4 / 37.0 tok/s | 42.1 / 42.4 tok/s | +13 % |
| 262,000 | 24.6 / 24.0 tok/s | 26.5 / 27.4 tok/s | +11 % |
At 32K a first series measured dev67 4 % slower and a second, longer series 3 % faster; the spread of a single setting is about 10 % there. The decode is mostly experts and dense weights at that length, and attention is 3–4 ms of 15.
Through redlite-server (needles, CF2): 3 / 3 at 62K and 127K, the same answers. Decode 44.2 and 36.2 tok/s
(dev66: 47.6 and 33.3). Single runs: ingestion, whose code did not change, also moved by 5 % (490 against 513 tok/s
at 62K), so these are not a comparison.
M4 Pro 24 GiB (kernel-bench, one KV pair): 256K 55.8 → 50.6 ms, 64K 14.4 → 12.9, 8K 2.13 → 1.78 (−8 to −16 % across 4K–256K).
decode-bench with the 4 GiB expert cache is dominated by expert reads; the warm runs moved within noise (32K 24.4
→ 24.3–24.6 tok/s, 64K 20.6 → 21.2–21.4, 128K 16.0 → 16.6).
The 128/256 block threshold was rechecked: 256 wins at 16K and from 64K, 128 wins at 32K in kernel-bench
(3.38 against 3.76 ms), but in decode-bench at 32K the difference is inside the noise (68–71 tok/s against
66–72). The rule stays at 256 from 12,288 positions.
3. Validation#
- Kernel self-test: 6 lengths × 4 kernels against double-precision GQA, worst abs 7.56e-8.
scripts/regress_m4.shon the reference IQ2_XXS: 55 / 0 / 0.quick_parity.sh --long: 5 / 5. M4 Max.- Needles at 62K and 127K through the server, answers identical to dev65 and dev66.
scripts/dev/local_ci.sh: PASS.
4. Smaller findings#
- RoPE precision is not the cause of the drift from llama.cpp past 64K. A model-free probe computed the RoPE
angles of Qwen3-Next (64 rotary dims, base 1e7) on the GPU:
- Fast-math and
precise::sin/cos give identical errors. - The GPU's
powadds a little: against the exact angle, the error at 128K is 2.1e-3 where float can do 1.5e-3 (a frequency table computed in double on the CPU reaches 1.5e-3). At 256K both are 5.6e-3. - llama.cpp computes the angle with the same
powformula, so this cannot explain a KL that grows from 1.2e-3 at 64K to 0.025 at 128K. Not changed.
- Fast-math and
- Four key loads in flight in the old one-thread-per-row score pass: −2 to −3 %, replaced by the eight-lane read.
Scope boundary#
- Decode attention only (
attn_gqa_split_g,attn_gqa_merge). No other kernel, no format change. - The block threshold is the dev66 one, chosen on the M4 Max.
- No CPU-vs-Metal engine parity reaches 12K positions; the long path is covered by the kernel self-test and the needle runs.