DOCS · REFERENCE
Red Lite v0.3 streaming development
Branch: v0.3-streaming (merged into main with release 0.3.0 and deleted; this page is the
historical dev1–dev8 log)
Status: historical. When this was written, the stable launcher was v0.2.1 on main.
Implemented in dev1#
- dependency-free GGUF tensor-directory parser;
- discovery of Qwen3-Next merged routed tensors named
blk.N.ffn_{gate,up,down}_exps.weight; - validation that the expert dimension is the outer GGUF dimension before physical slicing is enabled;
- byte-range mapping for individual experts inside merged tensors;
- bounded file-backed LRU using
mmap; MADV_WILLNEEDprefetch hints andMADV_DONTNEEDeviction hints where macOS/Python expose them;- synthetic top-k router probe to measure mmap/cache behavior without running a fake inference path;
- separate
redlite-streamCLI so stableredlitebehavior is unchanged.
Status (dev18)#
The native runtime now performs complete inference through the bounded expert
cache: redlite-generate tokenizes, embeds, runs all 48 layers on Metal with
persistent DeltaNet state and KV caches, produces logits, samples and decodes
text without llama.cpp or Python. Greedy output is token-identical to the
pinned llama.cpp on the validated prompts. See docs/REDLITE_DEV18_ENGINE.md
and CHANGELOG.md for the dev9–dev18 progression; the dev1 Python probes
below remain as the original mmap/residency experiment.
First real-model validation on the M4 Pro#
git switch main
git pull
./scripts/install.sh
redlite-stream plan models/Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf --cache-gib 4
The important fields are layers, routed tensors, slice safe, max expert triplet, and cache capacity.
If slice safe is YES, run a small file-residency probe:
redlite-stream probe models/Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf \
--cache-gib 4 \
--steps 2 \
--top-k 10
This will intentionally touch only a page-sized prefix of each selected expert window. It validates offsets, mmap behavior, LRU eviction and SSD-backed page residency without reading the entire model into Python memory.
Next milestones#
- Batched prompt prefill (the decoder currently ingests prompts token by token).
- Expert prefetch / speculative loading to overlap SSD reads with compute.
- Larger routed quantizations (IQ2_XS / IQ2_S / Q4_K_M experts) through the same bounded cache, benchmarked against the v0.2.1 IQ2_XXS resident baseline.
- Long-context throughput measurements and KV-cache layout work.