DOCS · MILESTONES
Red Lite dev32 — release 0.4.0
Status: prepared as 0.4.0 on branch dev/0.4; pull request dev/0.4 → main open and
not merged (merging is the maintainer's decision). Machine: Apple M4 Max 48 GiB.
What the release changes#
- Version 0.4.0 in
VERSION,pyproject.toml,redlite/__init__.pyand the usage text of the product binaries (redlite-engine,redlite-generate,redlite-server). The stage parity CLIs keep the dev tag of the milestone that built them. - CHANGELOG: the "Unreleased" block became
0.4.0 — 2026-09-30, with a summary and the milestone targets of the cycle, then the dev28–dev31 entries. - README: the two supported GGUFs,
--cache-mib full, howredlite chatpicks the model, the server's reuse / stop / queue, a measured-performance table per model and per cache (native and pinned llama.cpp on the same ids), correctness for both files. - docs/README.md indexes dev28–dev32.
- Release tarball:
scripts/package_release.sh→dist/redlite-0.4.0-macos-arm64.tar.gzand its.sha256(not committed; ready to attach to a GitHub release).
No runtime code changed in dev32 apart from the version strings.
Milestone results of the 0.4.0 cycle (M4 Max 48 GiB)#
| Milestone | Target | Reached | Commits |
|---|---|---|---|
| dev28 CI and distribution | Linux CI on push/PR green on GitHub; mac workflow manual; arm64 tarball + SHA256 + install notes | tarball, manual mac workflow, CI workflow: yes. Green on GitHub: not reached — GitHub does not start hosted jobs on this account (billing); the same steps pass in a Linux container | 3912aec b84bdd8 f6d1d1f 16f0963 |
| dev29 server | reuse when the prompt extends the conversation (greedy identical to cold), TTFT before/after, stop sequences, FIFO queue, fake-engine tests, real check in regress | yes: identical answer (SHA-1 1f94f64826b2 in every configuration), TTFT 3833 → 173 ms (4 GiB), 3779 → 167 ms (22 GiB) | 4a44030 79b3572 788a316 b3d8df2 8dff5fd |
| dev30 prefill and experts | prefill 1100 ≥ llama.cpp same prompt; 8192 ≥ 280 tok/s; decode 22 GiB not below 0.3.0 | yes with full residency: 891.9 vs 858.34, 926.9, decode A/B 71.09 vs 70.47. At 4 GiB / 512 chunks 1100 tokens reach 545.8 (below llama.cpp) | d5d7fd8 f058a0f 65bce55 55e6cbe 7e96a3e aed1ec5 fca7d78 8cbdffe |
| dev31 quality | higher-quality GGUF resident on 48 GiB; missing quant types (CPU ref + Metal); greedy = llama.cpp, KL ≤ 1e-5; perplexity of both; chat picks the best model with ≥ 40 GiB | yes: IQ3_XXS (29,376 MiB resident), KL 1.4e-12, 24/24 greedy, PPL 14.29 vs 16.47 | 388853e … d75d080 (git log dev30-green..dev31-green) |
| dev32 release | 0.4.0 versions, CHANGELOG, README per model/cache, tarball, PR | yes; PR open, not merged | this series |
Final validation#
On the release tree (versions 0.4.0), in this order:
make test: OK (Ran 64 tests);ruff check redlite tests scripts --select E9,F63,F7,F82: clean.rm -rf .deps/redmetal, thenmake redmetal,make native,make sanitize: all OK, 0 compiler warnings.scripts/package_release.sh:dist/redlite-0.4.0-macos-arm64.tar.gz(528 KiB), SHA-2564c34a1c5b6825556a9ce2099455e651c48e81b6911e6848c56cee7ab0d67400c.scripts/regress_m4.shon the IQ2_XXS GGUF: 49/49; on the IQ3_XXS GGUF: 36 passed, 0 failed, 13 skipped (the dev11–dev17 dense stage tools, IQ2_XXS layout only).- GitHub CI: see the pull request; hosted jobs are refused for billing (dev28).
Scope boundary#
- Everything was measured and validated on the M4 Max 48 GiB only. The 4 GiB-cache configuration intended for the M4 Pro 24 GiB uses the same kernels but was not run there; neither was anything on an M4 Air 16 GiB.
- The release tarball was built and started on the M4 Max; the
-mcpu=apple-m1build has not run on older Apple Silicon. It is not signed or notarized. - GitHub CI is configured and would run on every push and PR, but GitHub refuses to start hosted jobs for this account until its Actions billing is fixed.