DOCS · MILESTONES

Red Lite dev13 — real router field validation

Target: Apple M4 Pro, 24 GiB unified memory

Model: Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf

This checkpoint validates the real Qwen3-Next F32 router and the router-selected routed-expert branch entirely through the standalone native C/Objective-C/Metal path. Python and ctypes are absent from the tested execution path.

Router audit#

All 48 routed layers contain exactly one blk.<layer>.ffn_gate_inp.weight tensor.

  • tensor count: 48
  • shape: (2048, 512) for every layer
  • GGML type: F32 (0) for all 48 tensors
  • physical span: 4.000 MiB per layer
  • total router payload: 192.000 MiB

Layer 0 router parity#

  • CPU router read: 1.311 ms on the first isolated run
  • CPU router matvec: 0.841 ms
  • GPU router read: 0.463 ms
  • GPU router matvec: 3.004 ms
  • max logit absolute error: 4.76837e-07
  • max logit relative error: 2.63783e-04
  • top-k IDs match: YES
  • max selected-weight absolute error: 8.9407e-08
  • router parity: YES

Selected experts:

438,41,204,493,139,451,412,200,367,154

Normalized weights:

0.2299825,0.1769590,0.1074829,0.0931625,0.0795658,0.0691245,0.0675194,0.0613993,0.0578358,0.0569682

Layer 6 router parity#

  • CPU router read: 1.467 ms on the first isolated run
  • CPU router matvec: 0.846 ms
  • GPU router read: 0.434 ms
  • GPU router matvec: 3.190 ms
  • max logit absolute error: 4.76837e-07
  • max logit relative error: 1.56213e-05
  • top-k IDs match: YES
  • max selected-weight absolute error: 2.98023e-08
  • router parity: YES

Selected experts:

480,288,135,505,143,344,485,182,111,83

Normalized weights:

0.1826100,0.1416743,0.1046222,0.1015359,0.0955648,0.0899121,0.0741731,0.0731347,0.0686038,0.0681691

Real-router routed-expert parity#

Layer 0 / IQ2_XS#

  • real router parity: YES
  • expert loads: 10
  • cache hits/misses: 0/10
  • positional expert reads: 30
  • SSD I/O during top-k compute: 0 bytes / 0 calls
  • GPU routed top-k: 9.364 ms
  • CPU routed reference: 22.736 ms
  • routed max absolute error: 2.20497e-08
  • routed max relative error: 4.09202e-06
  • routed parity: YES

Layer 6 / IQ1_M#

  • real router parity: YES
  • expert loads: 10
  • cache hits/misses: 0/10
  • positional expert reads: 30
  • SSD I/O during top-k compute: 0 bytes / 0 calls
  • GPU routed top-k: 8.721 ms
  • CPU routed reference: 23.397 ms
  • routed max absolute error: 1.19063e-08
  • routed max relative error: 2.62410e-06
  • routed parity: YES

What this proves#

The following path is field validated on the real model:

F32 router GGUF -> native CPU/Metal router matvec -> softmax(512) -> exact top-10 -> selected-weight renormalization -> transactional native LRU -> positional expert reads -> resident IQ2_XS/IQ1_M Metal FFNs -> GPU weighted accumulation

The native CPU oracle independently agrees with both the router selection/weights and routed output.

What it does not prove yet#

Qwen3-Next also computes a shared expert and multiplies it by its own sigmoided scalar gate before adding it to the routed MoE output. That branch is intentionally excluded from dev13 routed parity. Dev14 begins by auditing the shared-expert tensors and then will validate the complete FFN output.

SOURCE docs/REDLITE_DEV13_FIELD_VALIDATION.md · updated 2026-09-03 · EDIT ON GITHUB