DOCS · MILESTONES

Red Lite dev34 — bounded-cache prefill: fewer expert reloads

Status: done on branch dev/bounded-prefill (from dev/iq3-kernels at 65cabd9). Apple M4 Max 48 GiB only. The change targets the 4 GiB configuration meant for 24 GiB Macs, which could not be measured on one (no M4 Pro here).

Finding#

With a bounded expert cache, the batched prefill routes every token of a chunk, loads the union of the selected experts layer by layer, and the next chunk finds them evicted by the other 47 layers: every chunk reloads nearly every expert of every layer. The chunk count, not the prompt length, sets the bytes loaded. redlite-engine logits … --cache-mib 4096 --batch N (IQ2_XXS, single runs; the prefill split line now reports the loads):

prompt chunk 512 (0.4.0 default) chunk 1024 chunk 2048
1100 tokens: experts loaded 37,022 (26.3 GB) 25,153 (17.9 GB) 16,784 (11.9 GB)
8192 tokens: experts loaded 226,117 (161 GB) 126,763 (90 GB) 69,071 (49 GB)

On this Mac the GGUF stays in the page cache, so a load is a memory copy (it also competes with the GPU for bandwidth). On a Mac whose page cache cannot hold the GGUF, loads are SSD reads, so 3.3× fewer bytes should matter more there; that is an expectation, not a measurement.

Change#

rl_engine_prefill_batch() returns 2048 for every cache size (dev30 used 2048 only with full residency; --batch still overrides). Footprint cost at 4 GiB with an 8387-token prompt: 5,231 → 5,802 MiB (+571 MiB, redlite-generate --json phys_footprint_mib). Chunks of 4096 would need more than the 32,768 pairs one expert plan accepts and another ~0.5 GB of scratch; not attempted.

Measurements#

Cooled, alternating (--batch 512 vs --batch 2048), IQ2_XXS, 4 GiB cache, 3 runs each:

prompt chunk 512 chunk 2048
1100 tokens 540.3 tok/s (540.3 / 543.0 / 520.9) 795.1 tok/s (795.1 / 783.2 / 807.0)
8192 tokens 639.5 tok/s (644.7 / 639.5 / 637.2) 885.9 tok/s (888.3 / 885.9 / 869.7)

The pinned llama.cpp, fully resident, does 858.3 / 893.7 tok/s on the same ids (dev28).

Parity#

1100-position context vs llama.cpp with the default chunk at 4 and 2 GiB: IQ2_XXS worst 0.99 / KL 3.5e-3, IQ3_XXS 0.57 / KL 3.9e-4 (as with 512-token chunks). quick_parity.sh: PASS. regress_m4.sh's generate.json now expects "batch":2048.

Validation#

rm -rf .deps/redmetal, make redmetal, make native, make sanitize: 0 warnings; make test and ruff green; scripts/regress_m4.sh: IQ2_XXS 49/49, IQ3_XXS 36 passed, 0 failed, 13 skipped (the IQ2_XXS-layout dense stage tools).

Scope boundary#

  • Not measured on a 24 GiB Mac; the benefit there (fewer SSD reads) is inferred from the load counts above.
  • The +571 MiB of scratch comes out of the memory a 24 GiB Mac has for the page cache.

SOURCE docs/REDLITE_DEV34_BOUNDED_PREFILL.md · updated 2026-09-30 · EDIT ON GITHUB