DOCS · START HERE
Documentation index
Every milestone document states what was implemented, how it was validated (machine, commit, command) and a Scope boundary of what it does not claim. "Validated" always names the machine: the M4 Pro 24 GiB (field target) or the M4 Max 48 GiB (development machine). "Tested synthetically" means model-free tests only.
Start here#
| Document | What it covers |
|---|---|
| ../README.md | The guide: install, FAQ, what runs, measured performance per Mac, correctness, limits, coding agents |
| FINDINGS.md | Read first. What we learned, with charts: where a token's time goes, speed by release vs llama.cpp, exact MTP speculation, 24 GiB cache options, long context, correctness, a better 2-bit file, coding agents |
| WHAT_DID_NOT_WORK.md | Every reverted attempt, correctness trap and unreached target, with the measurement that decided it |
| ROADMAP.md | Where things stand after 0.6.1 and the work in progress (agent quality, Qwen3-Coder at 2 bits, 48 GiB Macs) |
| ../CHANGELOG.md | Every release from 0.6.1 back to 0.3.0, then every dev milestone |
The project site, redlite.alfonsodaniello.it, renders these pages with a
search (scripts/dev/build_site.py).
Milestones of 0.5 and 0.6#
| Milestone | Document |
|---|---|
| dev62 routing quality, a full-precision agent reference, harder repo tasks, IQ3_XXS with an agent, two agents at once | REDLITE_DEV62_AGENT_TESTS.md |
dev61 24 GiB Mac at the default GPU limit: coding agent from the SSD, IQ3_XXS streamed, redlite setup-pi |
REDLITE_DEV61_STREAMING_24GB.md |
| dev60 the pi coding agent on the native server; agent-loop state reuse fixes | REDLITE_DEV60_CODING_AGENT.md |
| dev59 OpenAI tool calling in the server (Qwen3-Next JSON and Qwen3-Coder XML, exact rendering) | REDLITE_DEV59_TOOL_CALLS.md |
| dev58 dense weights at IQ3_XXS / Q4_K at E3's size, perplexity −5 % (the F2 file) | REDLITE_DEV58_DENSE_PRECISION.md |
dev56 two server requests decoded in one pass (--parallel 2) |
REDLITE_DEV56_PARALLEL_SERVER.md |
| dev55 long contexts at full residency on 24 GiB, where MTP stops paying, steering in the server | REDLITE_DEV55_LONG_CONTEXT_24GB.md |
| dev54 an expert-layer mix that beats Bartowski's IQ2_XXS at the same size | REDLITE_DEV54_QUANT_MIX.md |
dev52 activation steering (/steer in the chat, exact with MTP) |
REDLITE_DEV52_STEERING.md |
| dev51 decode on the 24 GiB Mac: cache size sweep, every expert resident, cache-aware routing by default | REDLITE_DEV51_24GB_DECODE.md |
| dev49 why the M4 Pro reads prompts at ~350 tok/s (GPU compute, not the disk) | REDLITE_DEV49_PREFILL_24GB.md |
| dev48 greedy answers compared word for word with Qwen's own API | REDLITE_DEV48_API_REFERENCE.md |
dev47 M4 Pro 24 GiB field session, long context (16K parity, 33.5K speed), REDLITE_SDK |
REDLITE_DEV47_LONG_CONTEXT.md |
| dev46 cache-aware routing bias, F_NOCACHE, engine perplexity, 24 GiB A/B script | REDLITE_DEV46_SMALL_MAC_OPTIONS.md |
| dev45 speculative decoding with the Qwen3-Next MTP block (+8–30 % decode, identical output) | REDLITE_DEV45_MTP.md |
dev43 session state checkpoints on disk (--state-dir) |
REDLITE_DEV43_STATE_CACHE.md |
| dev42 Qwen3-Coder-Next (same engine, validated against llama.cpp) | REDLITE_DEV42_CODER_NEXT.md |
Milestones of 0.4.0 and 0.4.1#
Developed on the dev/0.4 branch; M4 Max 48 GiB only.
| Milestone | Document |
|---|---|
| dev39 fused expert tail in GPU-routed decode (+1.6 %) | REDLITE_DEV39_EXPERT_TAIL.md |
| dev38 expert down lanes and concurrent decode encoders (IQ3_XXS 70.4 → 79.1 tok/s) | REDLITE_DEV38_DECODE_DISPATCH.md |
| dev37 expert slots of each layer's own size (4 GiB cache −24 % misses; full residency −3.9 GiB) | REDLITE_DEV37_EXPERT_SLOTS.md |
| dev36 IQ3_M GGUF supported (Q4_K experts), measured, not chosen on 48 GiB | REDLITE_DEV36_IQ3M.md |
| dev35 long-context decode: grouped split-K attention | REDLITE_DEV35_LONG_DECODE.md |
| dev34 bounded-cache prefill: fewer expert reloads | REDLITE_DEV34_BOUNDED_PREFILL.md |
| dev33 faster IQ kernels | REDLITE_DEV33_IQ_KERNELS.md |
| dev32 release 0.4.0 | REDLITE_DEV32_RELEASE.md |
| dev31 IQ3_XXS GGUF: new quant types, full residency on 48 GiB | REDLITE_DEV31_IQ3.md |
| dev30 batched prefill: tiled attention, matrix experts, faster dense pass | REDLITE_DEV30_PREFILL.md |
| dev29 server: state reuse, stop sequences, FIFO queue | REDLITE_DEV29_SERVER.md |
| dev28 Linux CI, release tarball, same-prompt llama.cpp baseline | REDLITE_DEV28_CI_RELEASE.md |
Native runtime milestones (v0.3, released as 0.3.0)#
Developed on the v0.3-streaming branch, which was merged into main with release 0.3.0.
dev9–dev17 built and validated one graph stage each. Their parity CLIs are now regression
tools, and redlite-engine (dev18 onward) supersedes them for inference. dev22–dev26 are
the 0.3.0 performance and robustness work.
| Milestone | Document |
|---|---|
| dev7–dev8 Red Metal resident experts | REDMETAL_DEV7.md, REDMETAL_DEV8.md |
| dev9 native foundation | REDLITE_DEV9_NATIVE.md |
| dev10 standalone Metal path | REDLITE_DEV10_NATIVE_METAL.md |
| dev11 CPU/GPU parity harness | REDLITE_DEV11_NATIVE_PARITY.md |
| dev12 field validation, offline hardening | REDLITE_DEV12_FIELD_VALIDATION.md, REDLITE_DEV12_OFFLINE_HARDENING.md |
| dev13 router | REDLITE_DEV13_ROUTER_AUDIT.md, REDLITE_DEV13_FIELD_VALIDATION.md |
| dev14 shared expert and complete FFN | REDLITE_DEV14_SHARED_AUDIT.md, REDLITE_DEV14_FIELD_VALIDATION.md, REDLITE_DEV14_COMPLETE_FFN_FIELD_VALIDATION.md |
| dev15 Gated DeltaNet | graph audit, projection, projection field, pre-state, state, tail, layer |
| dev16 full attention | REDLITE_DEV16_FULL_ATTENTION.md |
| dev17 48-layer decoder stack | REDLITE_DEV17_DECODER_STACK.md |
| dev18 end-to-end engine | REDLITE_DEV18_ENGINE.md |
dev19 redlite chat |
REDLITE_DEV19_CLI_CHAT.md |
| dev20 batched prefill | REDLITE_DEV20_BATCHED_PREFILL.md |
| dev21 GPU-routed decode | REDLITE_DEV21_GPU_ROUTED_DECODE.md |
| dev22 decode kernels (one encoder per token, sub-block GEMV) | REDLITE_DEV22_DECODE_KERNELS.md |
| dev23 per-layer early-out, pre-gated decode prefetch | REDLITE_DEV23_EARLY_OUT_PREFETCH.md |
| dev24 prefill: expert prefetch overlap, parallel routing | REDLITE_DEV24_PREFILL_OVERLAP.md |
| dev25 (partial) product surface: chat defaults, native server | REDLITE_DEV25_PRODUCT.md |
| dev26 robustness: GGUF fuzz, sanitizers, sampler parity | REDLITE_DEV26_ROBUSTNESS.md |
| dev26 completion: 4096/8192 positions, sanitized chat turn, kernel self-test, split-K attention | REDLITE_DEV26_LONG_CONTEXT.md |
| dev27 release 0.3.0: milestone results, final validation | REDLITE_DEV27_RELEASE.md |
Background and plans#
| Document | What it covers |
|---|---|
| STREAMING_V03.md, V0.3_STREAMING_DEV.md | v0.3 streaming design direction and dev1–dev8 notes (legacy: the Python streaming oracle redlite-stream / redlite-ffn / redlite-topk is frozen as a numerical reference) |
| METAL_STREAMING_ROADMAP.md | Native Metal + SSD expert streaming roadmap |
| DS4_ADAPTATION.md | Mapping from DwarfStar / DS4 ideas to Red Lite |
| RESEARCH_2026_10.md | October 2026 survey: current DwarfStar, Qwen3-Coder-Next, Qwen3-Next MTP GGUFs, llama.cpp/MLX/Metal techniques, ranked options |
| REDLITE_DEV18_ENGINE.md | The native end-to-end engine: design, validation, limits |
| ARCHITECTURE.md | Design goal and architecture of the v0.2 launcher (sparse 80B on a 24 GiB Mac) |
| MEMORY.md | Memory policy and headroom formula for 24 GiB Macs (launcher) |
| BENCHMARK.md | v0.2 benchmark protocol and context-depth sweep |
| FIELD_VALIDATION_M4PRO_24GB.md | v0.2 launcher field validation on the M4 Pro 24 GiB |