DOCS · MILESTONES

Red Lite dev14 — complete Qwen3-Next FFN field validation

Target: Apple M4 Pro, 24 GiB unified memory, Bartowski Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf.

Dev14 closes the complete feed-forward block on the real model. The validated graph is:

F32 router -> softmax -> top-10 -> renormalized routed MoE + sigmoid-gated shared expert -> FFN output.

Both CPU and Metal paths are standalone native implementations with no Python or ctypes in the tested path.

Layer 0#

This layer exercises routed IQ2_XS experts plus a Q6_K shared expert.

  • router IDs match: YES;
  • router max logit error: 4.76837e-07;
  • router max selected-weight error: 8.9407e-08;
  • selected experts: 438,41,204,493,139,451,412,200,367,154;
  • routed cold loads: 10;
  • routed positional reads: 30;
  • routed SSD reads during Metal compute: 0 bytes / 0 calls;
  • routed GPU / CPU: 9.285 / 21.592 ms;
  • routed max abs / rel error: 2.20497e-08 / 4.09202e-06;
  • routed parity: YES;
  • shared quant: Q6_K gate/up/down;
  • shared scalar gate CPU / GPU: 0.544462628 / 0.544462562;
  • shared read calls: CPU=4, GPU=4;
  • shared SSD reads during Metal compute: 0 bytes / 0 calls;
  • shared GPU / CPU: 3.017 / 1.787 ms;
  • shared max abs / rel error: 1.23826e-07 / 1.26553e-05;
  • shared parity: YES;
  • complete FFN max abs / rel error: 1.32396e-07 / 6.31249e-05;
  • complete FFN parity: YES.

Layer 6#

This layer exercises routed IQ1_M experts plus an IQ2_XXS shared expert.

  • router IDs match: YES;
  • router max logit error: 4.76837e-07;
  • router max selected-weight error: 2.98023e-08;
  • selected experts: 480,288,135,505,143,344,485,182,111,83;
  • routed cold loads: 10;
  • routed positional reads: 30;
  • routed SSD reads during Metal compute: 0 bytes / 0 calls;
  • routed GPU / CPU: 8.301 / 22.242 ms;
  • routed max abs / rel error: 1.19063e-08 / 2.6241e-06;
  • routed parity: YES;
  • shared quant: IQ2_XXS gate/up/down;
  • shared scalar gate CPU / GPU: 0.407262824 / 0.407262862;
  • shared read calls: CPU=4, GPU=4;
  • shared SSD reads during Metal compute: 0 bytes / 0 calls;
  • shared GPU / CPU: 2.926 / 1.736 ms;
  • shared max abs / rel error: 8.18691e-08 / 4.46434e-06;
  • shared parity: YES;
  • complete FFN max abs / rel error: 7.44724e-08 / 1.63362e-05;
  • complete FFN parity: YES.

Q6_K oracle correction#

During dev14 bring-up, Q6_K Metal initially appeared to disagree with the native CPU reference. An isolated real-matrix probe showed Metal outputs at exactly twice the CPU result for affected rows. The actual defect was in the CPU FP16-to-FP32 conversion for half subnormal values: the exponent used 127 - 15 - shift instead of 127 - 14 - shift.

After fixing the oracle and adding a dedicated FP16-subnormal regression fixture:

  • Q6_K gate matrix parity: YES;
  • Q6_K up matrix parity: YES;
  • Q6_K down matrix parity: YES;
  • complete Q6_K shared branch parity: YES.

The routed dev12/dev13 reference converter used a different, correct half-subnormal implementation and is not invalidated by this fix.

Validation boundary#

Dev14 validates the complete Qwen3-Next FFN for representative real layers covering both routed quant families and both shared-expert quant families. It does not yet validate the complete decoder block. The next milestone must cover the surrounding graph:

RMSNorm -> DeltaNet/full attention -> residual -> RMSNorm -> complete FFN -> residual.

SOURCE docs/REDLITE_DEV14_COMPLETE_FFN_FIELD_VALIDATION.md · updated 2026-09-03 · EDIT ON GITHUB