DOCS · REFERENCE

Field validation — Apple M4 Pro / 24 GiB

Date: 2026-09-02

Red Lite now has both an interactive field observation and a controlled context-depth sweep on its primary target class.

Hardware / model#

  • Apple M4 Pro
  • 24 GiB unified memory
  • 8 performance CPU cores detected
  • Qwen3-Next-80B-A3B-Instruct
  • Bartowski IQ2_XXS GGUF
  • actual local file size: 17.97 GiB
  • llama.cpp Metal backend
  • batch / ubatch: 256 / 128
  • K/V cache types: q8_0 / q8_0

Metal reported:

  • MTLGPUFamilyApple9
  • simdgroup matrix multiplication enabled
  • residency sets enabled
  • shared buffers enabled
  • recommended maximum working set: 21474.84 MB
  • Tensor API disabled on M4 Pro (expected for pre-M5 hardware in this llama.cpp build)

Interactive observation#

Interactive inference completed successfully at 2K context. Observed timings during the session were approximately:

  • prompt processing: 36.7–87.5 tokens/s
  • generation: 30.7–40.7 tokens/s

These interactive figures are not a controlled benchmark because prompt lengths, cache state and sampling differed between turns.

Controlled sweep#

Command:

redlite sweep models/Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf \
  --contexts 2048,4096,8192 \
  --prompt-tokens 256 \
  --gen-tokens 128 \
  --max-swap-delta 0.5 \
  --output benchmarks/m4pro-24gb.json

Observed results:

Context Policy status Estimated headroom Prompt tok/s Generation tok/s Swap delta
2048 CRITICAL +0.61 GiB 258.8 38.0 +0.00 GiB
4096 CRITICAL +0.50 GiB 247.2 36.4 +0.00 GiB
8192 CRITICAL +0.28 GiB 236.7 36.9 +0.00 GiB

All three context depths completed and no swap growth was observed during the sweep.

The result is notable because generation throughput stayed essentially flat from 4K to 8K while prompt throughput declined gradually. The policy still labels all three depths CRITICAL because its estimated resident margin is below 1 GiB; observed success does not turn a narrow memory margin into a generally safe one.

  • 2K — conservative: largest estimated margin; useful when other applications must remain open.
  • 4K — default: field-validated and preserves more margin than 8K.
  • 8K — experimental: completed with zero observed swap delta, but estimated headroom is only 0.28 GiB. Longer sessions and repeated runs are still needed before promoting it to the default.

Red Lite 0.2.1 therefore changes the automatic default to 4096 only for the observed Apple M4 Pro / 24 GiB / ~18 GiB resident profile. Other unvalidated 24 GiB machines remain on the conservative 2K default.

Planner status model#

Headroom Status Meaning
>= 2 GiB SAFE comfortable policy margin
1–2 GiB TIGHT usable, monitor pressure
0–1 GiB CRITICAL resident fit with little estimated margin
< 0 GiB UNSAFE reject resident mode

The controlled sweep is also stored in benchmarks/m4pro-24gb-sweep-2026-09-02.json and summarized in configs/mac-m4pro-24gb-observed.json.

SOURCE docs/FIELD_VALIDATION_M4PRO_24GB.md · updated 2026-09-02 · EDIT ON GITHUB