DOCS · START HERE
What did not work
Everything that was tried on the native runtime and not kept, every trap that cost time, and every target that was not reached, in one place, with the measurement that decided it. All on the Apple M4 Max 48 GiB unless stated. Each entry links the milestone document with the details.
Method behind every "no gain": the change passed the parity gates, then an A/B against the previous build (cooled runs, alternated in the same session) did not beat run-to-run noise or was slower. Reverted changes leave no code behind; their environment switches were removed with them.
Performance attempts that were reverted#
| milestone | attempt | measurement | why it failed (best understanding) |
|---|---|---|---|
| dev20 | experts read in place from the mmap (RL_PREFILL_MAPPED_EXPERTS=1, still opt-in) |
numerically identical, 4–10× slower prefill | Metal re-established residency of each layer's ~350 MB expert window for every command buffer (~3.3 s per 48-layer chunk). Not re-tried since dev30's residency set (DEV20) |
| dev20 | batched expert kernel layouts: 8-pair register tile / adaptive 4-2-1 tiles / 4-row tile / grid over pairs | 981 / 370 / 295 / 256 ms vs 312 baseline (96-token chunk) | padding waste and load imbalance between experts; kept: slices of 4 pairs (DEV20) |
| dev22 | fewer Metal encoders alone (≈1 200 → 1 per token) | +1.9 % (57.6 → 58.7 tok/s) | encoder boundaries are cheap on this GPU; the time is inside kernels (DEV22) |
| dev23 | spin-waiting on MTLCommandBuffer.status instead of waitUntilCompleted |
4 GiB decode 44.6 → 27.5 tok/s | polling contends with the driver's completion path (DEV23) |
| dev23 | per-layer early-out as the only 4 GiB decode path | 8.3 tok/s | 45–87 % of tokens miss in every layer; no always-hit layers to chain (DEV23) |
| dev23 | periodic probing (32 synchronous tokens, then one GPU-routed try) | 39.5 tok/s vs 44+ | probes cost more than they gain; replaced by the load-free-token rule |
| dev26 | split-K decode attention at every length | short context 68.47 vs 70.20 tok/s | one block has the single-threadgroup kernel's parallelism plus a merge dispatch; split only above 256 positions (DEV26) |
| dev30 | router selection on the GPU (batched rl_route arithmetic) |
prefill 1100: median 571 vs 567 tok/s (runs 571/482/614 vs 577/567/513) | the CPU step fell 70–115 → 45 ms, but the gain is inside run-to-run spread (DEV30) |
| dev30 | matrix expert kernels: skip the empty 8-pair block of small slices | variable loop bound: experts 1 729 ms (accumulators spilled); constant bound + break: 901 ms vs 750–850 |
the extra branch and register pressure cost more than the saved MACs (DEV30) |
| dev30 | matrix expert kernels: split accumulators summed through threadgroup memory | experts 966 vs 777 ms | 26 KB of threadgroup memory per group cut occupancy; kept: summed in registers with an exact x·I + y (DEV30) |
| dev33 | IQ codebooks staged in threadgroup memory, 256-thread GEMV groups (what llama.cpp does) | IQ2_S 35.7 vs 33.2 µs, IQ3_XXS 38.2 vs 37.9 µs | divergent constant-memory reads were not the bound once the GPU was warm (DEV33) |
| dev33 | uchar4 codebook loads in the IQ2_XXS dense GEMV |
25.6 / 27.5 vs 27.3 / 25.1 µs | not the bound; the IQ2_XXS GEMV stays at ~169 GB/s (DEV33) |
| dev36 | IQ2_XXS dense GEMV with two rows per lane (one activation load serves both rows) | 25.4 vs 25.5 µs (8192 × 2048, warm kernel-bench) | third attempt on this kernel without gain; activation loads are not the bound. The IQ2_XXS dense GEMV stays at ~169 GB/s |
| dev37 | smarter expert replacement than LRU (simulated on routing traces, 4 GiB): decayed frequency (the Mira-style score), segmented LRU | best 60.9 vs 64.1 misses/token (−5 %); long half-lives and a 95 % protected segment are worse than LRU | decode routing is recency-dominated; the real waste was slot padding, fixed instead (dev37) |
| dev38 | two more concurrent-encoder groups ({ssm_out, conv-state copy}, {layer-output copy, next RMSNorm}) | 83.72 vs 83.77 tok/s | the removed barriers were not on the critical path |
| dev38 | (bound, not a change) concurrent decode encoder with no barriers at all | 74 → 118 tok/s, wrong output | most dispatches depend on the previous one; the kept version gains 2–6 % |
| dev39 | expert tail + next RMSNorm in one single-threadgroup kernel (−3 barriers per layer) | 79.1 vs 84.4 tok/s (5 pairs) | the 10-expert weighted sum on one GPU core costs more than the barriers; the parallel tail without the RMSNorm is kept (+1.6 %) |
| dev40 | prefill expert tiles of 32 (expert, token) pairs instead of 16 (each 32-row weight tile decoded once for twice the pairs), 128 threads with doubled per-thread accumulators | IQ3_XXS 1100 / 8192 tokens 400 / 404 vs 872 / 914 tok/s, bit-identical logits | register spills |
| dev40 | same with 256 threads (simdgroups 0-3 pairs 0-15, 4-7 pairs 16-31, per-thread work unchanged) | 1100: 886.5 / 866.2 / 878.0 vs 875.4 / 874.8 / 726.1; 8192: 921.6 / 906.8 / 909.2 vs 914.0 / 914.8 / 901.7 (medians 878.0 vs 874.8, 909.2 vs 914.0), bit-identical | re-decoding weights per 16-pair slice is not the prefill bound at these sizes |
| dev41 | decode expert kernels (gate/up, down) in threadgroups of 64 / 128 / 256 threads instead of 32 | 80.97 / 80.80-81.30 / 80.74-80.93 vs 81.30 / 81.51 tok/s (IQ2_XXS, full residency, two passes) | scheduling of the 5120-20480 one-row groups is not the bound |
| dev50 | 32-pair prefill expert tiles, 256 threads, simdgroups 4-7 skip empty pairs (bit-identical) | experts 4.21 vs 3.97 s (M4 Max, 8192 tokens) | the halved decode per pair did not pay for the larger threadgroup |
| dev50 | 32-pair tiles, 128 threads, two partial accumulators (register count of the 16-pair kernel); with outputs aliased on the weight tiles; with the branch unrolled; without the empty-pair skip | experts 6.87 / 8.01 / 8.00 / 9.57 vs 3.86 s | on this GPU 32-pair tiles lose with every accumulator layout tried (dev40 too); cause not understood |
| dev45 | relying on the compiler to share the weight decode between two dot calls (2-row verify) | 1.74x one vector instead of ~1.06x | not shared; explicit two-vector functions are used |
| dev45b | MTP draft through the trunk's Q5_K head instead of the MTP file's Q8_0 head | draft 1.56 → 1.51 ms, decode +0.3-0.8 % | inside noise; the head is not the draft's bound |
| dev35 | grouped decode attention, 4 threads per position for the scores | attention 175 vs 165 ms per 64 tokens at ~8400 positions | more threads per row did not add memory parallelism that mattered |
| dev35 | grouped decode attention with blocks of 256 / 64 / 32 positions | 164 / 146 / 180 ms vs 113 ms with 128 | 256: too few threadgroups (66); 64 and 32: more merge work and shorter loops (DEV35) |
Correctness traps (found, fixed or guarded)#
- Single accumulator in the matrix expert kernel (dev30). Rounding along the 2048-column reduction flipped an argmax at position 1135 of the long-context check (two logits 0.012 apart, after the documented router near-tie at 1035). Fixed with four partial accumulators; the gate caught it before commit.
- A stale dump passed a parity check (dev24).
quick_parity.shcompared an old file after a crash; the script now deletes its outputs first. In dev26compare_dumps.pywas found to accept an empty native dump; it now needs equal, non-zero token counts. - Silent truncation (dev26).
redlite-engine --tokenscut lists at 4 096 ids; now 65 536 and longer lists are rejected. - The 8192-position logit bound (dev30). The 4.0 max-logit bound set in dev26 from a 4096-position measurement is below the native self-consistency floor at 8192: the 0.3.0 binary's own token-by-token path differs by 4.29 from llama.cpp and 4.45 from its batched prefill there. The bound is 5.0 at ≥ 8192 positions; argmax and KL gates did not change.
- The
logitsCLI default chunk (dev30). Without--batchit ran token by token, so a "default chunk" parity run silently tested another path. It now uses the engine's chunk. - Benchmark artifacts. GPU clock ramp: the first kernel-bench numbers were 2–3× off until a warm-up pass was added (dev33). Thermal state: the same binary measured 64.7–71 tok/s at different hours of one night, and back-to-back prefill runs throttled to a third of the speed (dev24): only cooled, alternated, same-session A/B numbers are compared.
- Page cache state. 4 GiB-cache numbers depend on whether the GGUF is still cached: after
runs that allocate 22–29 GiB of expert slots, the first 4 GiB run read from SSD (IQ3_XXS 8192
tokens: 169 vs ~600 tok/s).
bench_m4.shwarms up before 4 GiB prefill runs. - Ad-hoc parity token lists (dev39).
redlite-engine parity --tokens 9707,11,1879,0,785,12884,13,151645,198 --cache-mib 2048fails at position 7 with a 2-id router mismatch and ~1e-5 logit errors, identically with the dev37 engine: a router near-tie on that sequence. Use the regression token lists. redlite-engine prefillon long IQ3_XXS sequences (dev39) reportsBATCHED PREFILL PARITY: NOat 1100 tokens (worst 0.027, 0 router mismatches, same argmax; dev37: 0.040). Its bound is sized for short sequences; the long-context check against llama.cpp is the gate.- Shared GPU (dev39). With other applications using the GPU, the same binary measured 30–45 tok/s decode and 335 tok/s prefill (instead of ~85 and ~800). Pairs run in such a window are excluded.
- A stale binary in a release benchmark (0.4.1). After reverting the dev40 kernels in the source,
.deps/redmetalwas not rebuilt, and the first 0.4.1 table was measured on the experimental build (IQ2_XXS prefill 854.3 / 858.3 tok/s instead of 0.4.0's 891.9 / 926.9). The table was re-measured afterrm -rf .deps/redmetaland a clean build. Rule: rebuild from a clean tree before any benchmark that goes into a record. - Battery power (dev47). On battery at 6 % the same binary decoded at 32 tok/s instead of ~81 and a
33.5K-token prompt took 17 minutes; every measurement of that window was discarded.
pmset -g battbefore benchmarks. - The token-by-token llama.cpp oracle at long positions (dev47).
redlite-ref-llama logitsdumps every position one decode at a time: 16K positions took about 2 hours, 32K would take about 5 and 60K 9. The 32K / 60K comparisons were not run. - The llama.cpp oracle on 24 GiB (dev47).
redlite-ref-llamaputs the whole 18 GiB model on Metal and fails withkIOGPUCommandBufferCallbackErrorOutOfMemoryon the M4 Pro. Correctness there is shown by bit-identity with the M4 Max's native outputs. - Command Line Tools with an SDK newer than their linker (dev47). On the M4 Pro (CLT 26.6, ld-1267) the
MacOSX27 SDK's
.tbdfiles read as "unknown architecture arm64e.x1";REDLITE_SDKpoints the builds at MacOSX26.5. - Chunk-parallel DeltaNet prefill (dev44): not built. Measured ceiling: the recurrence is at most 6.7 % of an 8387-token prefill, so the published 1.3–1.45× would give about 2 %.
- Float16 KV cache (dev53): built, measured, not kept. Every Metal attention kernel was switched to float16 storage
(
attn_fa_bwithsimdgroup_half8x8K/V tiles, mixed-precision multiply-accumulate), and the CPU oracle rounded K/V to float16 (identical to the hardware conversion on 4.3 M bit patterns). Model-free self-test,quick_parity.shIQ2_XXS and the 4K/8K/16K comparisons with llama.cpp passed. The median KL vs llama.cpp was slightly better (16K: 6.9e-9 vs 9.0e-9); single-logit outliers moved (16K: 9.66 vs 6.21, at positions with KL ~1e-7).- Gain: half the KV memory (64K context: footprint 21,567 → 20,030 MiB on the M4 Max). On the M4 Pro 24 GiB with every expert resident, MTP + a 32K context fit only with float16 (20.4 GiB; float32 ran out of GPU memory). But MTP there decoded 32.6 tok/s against 37.2 without it, and float32 without MTP already fit (19.9 GiB).
- No speed gain: M4 Max 33.5K-token prompt, decode 56.1 / 56.8 vs 54.5 / 57.4 tok/s; prefill 465* / 620 vs 566 / 587 tok/s (*disturbed run).
- Gates failed:
engine.parity.steered: 2 router id mismatches CPU vs Metal; logits 1.5e-4 vs 5.7e-6. The unsteered parity passed but its logits diff rose from ~1.5e-5 to 1.0e-4.- IQ3_XXS:
prefill.parity.96(6 router mismatches), andlogits.long_context_vs_llama/long_context_full_residency(argmax disagreement; KL 2.9e-3 within its bound).
- Why: CPU (double) and Metal (float) compute nearly equal K/V values. Float16 storage turns a tiny difference that straddles a rounding boundary into one float16 step (~1e-3 relative). Next to the router near-ties of these fixtures, that flips expert choices. A CPU/Metal router divergence is a hard failure, and the gain is memory alone, so it was reverted.
- The Qwen API as a greedy reference (dev48).
temperature: 0alone gave a different text in one run of three;top_k: 1is needed. Its logprobs are misaligned: the chosen token was missing from its own top 5 at 23 of 30 positions. Only text is compared. - A first single run overstated a gain (dev49). The vectorized expert decode measured −10 % expert time in one run against an earlier baseline run (4.40 → 3.94 s); alternated pairs gave −3.4 % (4.105 → 3.968 s). The baseline had been measured in a hotter state. Only alternated pairs are reported.
- Removing the multiplications to bound a kernel (dev49). With the simdgroup MACs deleted, the expert time fell to 266 ms because the compiler also dropped the now-unused weight decode. Only the "decode replaced by constants" variant (2.45 s of 4.40 s) is a valid bound.
- Steering doses (dev52). A difference-of-means vector at strength 1 over 16 layers, or 0.6+ over 8 layers, breaks the text into repeated tokens: the vector (|v| 4.35) is added at every steered layer, against a residual norm of about 10. Usable doses were 0.2–0.5 over 8–12 layers.
- Two GPU jobs at once (dev54). An API comparison at full residency (17 GiB) started while llama-perplexity held
another 19 GB file on the GPU. Request 8 failed with
kIOGPUCommandBufferCallbackErrorOutOfMemory. GPU-heavy runs are chained, never overlapped. - The pinned llama.cpp's IQ2_XXS recipe is not Bartowski's (dev54). It makes IQ2_XXS experts (not runnable in the
engine) and Q4_K attention. Custom mixes copy every tensor type from Bartowski's file
(
scripts/dev/quant_mix.py --like). - Expert mixes that did not help (dev54). IQ2_XS down projections on every layer with IQ1_M gate/up: +3 % size and perplexity 16.423 vs 16.370 (t = +1.13), and the engine cannot run it (type word 0x111d). Keeping layers 0–2 precise (E3b): no measurable gain (t = −1.28).
- Kernel panics in the GPU driver (2026-10-04, M4 Max 48 GiB, macOS 27.0.1).
- What happened. At 01:50:34, during
regress_m4.sh(checkengine.parity.gpu_routed: every expert resident, ~16 GB on the GPU), macOS panicked: "Kernel data abort" at address 0x10 insidecom.apple.iokit.IOGPUFamily(offset 0x2a31c, reached from 0xabd0), in a system call ofredlite-engine. Its second thread was in the kernel outside the driver. A user process cannot dereference a kernel null pointer, so this is a driver bug that our workload reached. - Not reproducible after a reboot. The same check passed 5 times in a row, and the whole regression passed 52/0/0 in the same order.
- State at the time (not proven to matter): 3 days of uptime with dozens of 17–29 GiB GPU load/unload cycles, swap in use, and a GPU out-of-memory error four hours earlier (two overlapped GPU jobs, dev54).
- It happened again at 12:37, 5 hours after a fresh boot, in
redlite-serverduringserver_check.py(1.6 GB resident;llama-quantizeand the local CI were running at the same time). - Same signature: the same driver offsets (0xabd0 → 0x2a31c), the same null address, and kernel stacks identical frame by frame. In both, two threads of our process were in the kernel: ours at default priority, in the driver, and a priority-46 thread outside it.
- macOS 27.0.1 (build 26A434) was installed on that Mac on 2026-09-30. The same checks had run hundreds of times on macOS 26 without a panic. The M4 Pro (macOS 26.7) has never panicked. A driver regression in 27.0.1 is the likely cause. Not proven: the kernel has no symbols, and the bug does not reproduce on demand.
- Practice.
- Never overlap GPU-heavy jobs.
- On macOS 27.0.1, run long Metal gates one at a time, and prefer a Mac on macOS 26 when available.
- Reports for Apple Feedback Assistant:
/Library/Logs/DiagnosticReports/panic-full-2026-10-04-113143.0002.panicandpanic-full-2026-10-04-123737.0002.panic.
- What happened. At 01:50:34, during
- zsh word splitting broke two benchmark loops (
set -- $vardoes not split in zsh); the measurement scripts run underbash.
Not reached, blocked or not attempted#
- GitHub CI green (dev28): not reached, then removed (2026-10-03). GitHub refused to start hosted jobs on
this account ("recent account payments have failed or your spending limit needs to be increased"). The
project does not pay for hosted runners. The same steps run locally with
scripts/dev/local_ci.sh. - IQ3_XXS prefill vs llama.cpp: 837 vs 861 tok/s at 1100 tokens with full residency (dev33). Re-measured 2026-10-01 at 547539f on an idle machine: native 859.4 (3 runs), 871.1 / 876.7 / 872.5 / 875.4 / 874.8 / 872.4 (A/B baselines), llama.cpp 878.8 (3 runs) on the same ids: about 0.6 % behind, inside run-to-run spread.
- 4 GiB-cache prefill on a 24 GiB Mac: measured in dev47. 280.8 tok/s (1100 tokens) and 349.8 tok/s (8192 tokens) on the M4 Pro 24 GiB. The dev34 gain alone (3.3× fewer expert bytes loaded) was not isolated there.
- Dense stage tools on the IQ3_XXS GGUF. The dev11–dev17 per-stage parity tools only accept
the IQ2_XXS dense layout;
regress_m4.shreports 13 SKIPs on the IQ3_XXS file. - IQ3_M on this Mac (dev36): runs, but is not worth it. The file is supported since dev36
(its Q4_K expert down projections got a decoder). Measured on the 48 GiB M4 Max:
- Full residency (34944 MiB) does not hold: GPU-routed decode ran at 46.8 tok/s, then 6.0 tok/s
on the next run, then failed with
kIOGPUCommandBufferCallbackErrorOutOfMemory(footprint 35.2 GiB). The planner's 75 % rule admitted it; the rule is now 70 %. - A 28 GiB cache gives 28.7 tok/s (97 % hits), against ~70 tok/s for IQ3_XXS with full residency.
- Perplexity on the frozen corpus is 14.05 ± 0.26 against 14.29 for IQ3_XXS: inside the error bar.
- Correctness is not the problem: with bounded caches it matches llama.cpp (logits, 24 greedy tokens, 1200-token long context). It is left out of the automatic model choice because on this Mac it is slower than IQ3_XXS for no measurable quality gain.
- Full residency (34944 MiB) does not hold: GPU-routed decode ran at 46.8 tok/s, then 6.0 tok/s
on the next run, then failed with
- IQ3_XS, IQ4_XS: not run. Header reads (no download) show IQ3_XS needs no new type and would fit at full residency only under the old 75 % rule; IQ4_XS (39168 MiB of experts) does not fit this Mac.
- Speculative decoding (dev37): not started then, done in dev45 with the MTP block from a separate file
(DEV45). The 2026-09 reasoning, still valid for 24 GiB Macs: the IQ2_XXS / IQ3 GGUFs carry no Qwen3-Next
multi-token-prediction layers (48 blocks, no
nextnkeys), so the MTP-based methods do not apply. Training-free drafts (prompt lookup, a reduced-expert self-draft) need a cheap multi-token verification, and the batched prefill path is slower than decode at small batches: 8-token chunks 137 ms against 8 decode steps 98 ms (full residency,redlite-engine prefill --batch 8). Papers on MoE speculation (EcoSpec 2607.12696, MoE-Spec 2602.16052, DraftExpert 2607.24434) also report that verifying several tokens loads the union of their experts, which is the 24 GiB bottleneck. - LensVLM (dev37): not implemented.
apple/lensvlm-9breads text rendered as compressed images (5–15× fewer tokens) and expands relevant pages on demand. It is a different model (Qwen3.5-9B + Qwen3-VL vision encoder, BF16 safetensors); Qwen3-Next-80B is text-only and cannot consume vision tokens, so the idea cannot be applied to this model without a second model family in the runtime. - OLED-MoE (2609.33385): not applicable. Its inter-iteration expert retention targets diffusion LLMs.
- Overlapping CPU encoding with GPU execution (dev38): not attempted. At full residency the CPU gap between GPU-routed tokens is 0.44 ms of 12.1 ms.
- 4096-token prefill chunks: not attempted. They exceed the 32 768 pairs one expert plan accepts and add ~0.5 GB of scratch (dev34).
- llama-perplexity in the oracle build tree. Rebuilding it there failed (OpenSSL target)
after relinking one oracle library from the same pinned source; it is built in its own tree
(
.deps/llama.cpp/build-ppl) byscripts/dev/perplexity.sh(dev31).