DOCS · REFERENCE

Mapping from DwarfStar / DS4 to Red Lite

Red Lite is not a source fork of DS4 v4. It is an adaptation of its engineering principles to a much smaller sparse model and a much smaller Mac.

DwarfStar idea Red Lite adaptation
Deliberately narrow model support Qwen3-Next only
Metal is the primary Mac backend Metal resident path for models that fit
SSD streaming for RAM-limited Macs Bounded mmap expert residency for oversized Q4
Asymmetric expert quantization matters Curated low-bit resident and Q4 quality profiles
Integrated CLI/server redlite run and redlite serve
Model-specific memory behavior Qwen3-Next recurrent/hybrid-aware conservative planner
Reproducible performance work pinned engine commits + benchmark protocol

Why Qwen3-Next instead of shrinking DeepSeek V4#

Qwen3-Next-80B-A3B already has the desired sparse shape: 80B total, ~3B active, 512 experts and top-10 routing. Reducing DeepSeek V4 itself would require pruning or retraining weights and validating the altered router/model quality. Using an existing 80B sparse checkpoint preserves an upstream-trained model while solving the hardware problem at the runtime layer.

Why v0.2 still uses two proven engines#

Writing a new Qwen3-Next Gated DeltaNet + full-attention + MoE Metal runtime from scratch is possible, but calling it complete before numerical comparisons and real Mac testing would be irresponsible. v0.2 therefore owns the Apple-specific policy, UX and reproducibility layer while using pinned, already-Qwen3-Next-capable kernels.

The architectural endpoint remains a native Red Metal expert-streaming backend, but that should only replace the current engines after reference-logit and long-run tests pass on actual Apple hardware.

SOURCE docs/DS4_ADAPTATION.md · updated 2026-09-02 · EDIT ON GITHUB