DOCS · REFERENCE

Red Lite v0.3 streaming development

Branch: v0.3-streaming (merged into main with release 0.3.0 and deleted; this page is the historical dev1–dev8 log)

Status: historical. When this was written, the stable launcher was v0.2.1 on main.

Implemented in dev1#

  • dependency-free GGUF tensor-directory parser;
  • discovery of Qwen3-Next merged routed tensors named blk.N.ffn_{gate,up,down}_exps.weight;
  • validation that the expert dimension is the outer GGUF dimension before physical slicing is enabled;
  • byte-range mapping for individual experts inside merged tensors;
  • bounded file-backed LRU using mmap;
  • MADV_WILLNEED prefetch hints and MADV_DONTNEED eviction hints where macOS/Python expose them;
  • synthetic top-k router probe to measure mmap/cache behavior without running a fake inference path;
  • separate redlite-stream CLI so stable redlite behavior is unchanged.

Status (dev18)#

The native runtime now performs complete inference through the bounded expert cache: redlite-generate tokenizes, embeds, runs all 48 layers on Metal with persistent DeltaNet state and KV caches, produces logits, samples and decodes text without llama.cpp or Python. Greedy output is token-identical to the pinned llama.cpp on the validated prompts. See docs/REDLITE_DEV18_ENGINE.md and CHANGELOG.md for the dev9–dev18 progression; the dev1 Python probes below remain as the original mmap/residency experiment.

First real-model validation on the M4 Pro#

git switch main
git pull
./scripts/install.sh

redlite-stream plan models/Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf --cache-gib 4

The important fields are layers, routed tensors, slice safe, max expert triplet, and cache capacity.

If slice safe is YES, run a small file-residency probe:

redlite-stream probe models/Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf \
  --cache-gib 4 \
  --steps 2 \
  --top-k 10

This will intentionally touch only a page-sized prefix of each selected expert window. It validates offsets, mmap behavior, LRU eviction and SSD-backed page residency without reading the entire model into Python memory.

Next milestones#

  1. Batched prompt prefill (the decoder currently ingests prompts token by token).
  2. Expert prefetch / speculative loading to overlap SSD reads with compute.
  3. Larger routed quantizations (IQ2_XS / IQ2_S / Q4_K_M experts) through the same bounded cache, benchmarked against the v0.2.1 IQ2_XXS resident baseline.
  4. Long-context throughput measurements and KV-cache layout work.

SOURCE docs/V0.3_STREAMING_DEV.md · updated 2026-09-30 · EDIT ON GITHUB