Publications

Official write-ups of the project's results to date — every number traceable to a committed result file with hardware manifest, every figure regenerated from raw data.

TR-01 Tue Aug 11 2026 20:00:00 GMT-0400 (heure avancée de l’Est) final

Margin-Gated Deferred Refinement: Streaming Higher-Precision LLM Quality Than Fits in Memory on Consumer Apple Silicon

Simon-Pierre Boucher

We ask whether an existing pretrained LLM whose preferred precision does not fit in a consumer Mac's unified memory can still contribute its quality locally. On an Apple M5 Max (48 GB), we first measure the substrate: the internal NVMe sustains ~13.1 GB/s of sequential reads but only 67 MB/s at 4 KiB random QD1, and saturated Metal GPU compute perturbs SSD streaming by less than 5%. We then show that the top-1 logit …

Read the full report →