Code / experiments/micro/expF_error_accumulation/hypothesis.md
experiments/micro/expF_error_accumulation/hypothesis.md
68 lines
---
project: localvm-research
document: expF_error_accumulation/hypothesis
author: Simon-Pierre Boucher
contact: contact@spboucher.ai
created: 2026-08-12
modified: 2026-08-12
status: reviewed
---
# Hypothesis — expF_error_accumulation
Follows expD: escalation currently means touching the FULL residual; if layer
sensitivity is concentrated, escalation can be restricted to a subset of
layers and candidate C1's bytes-per-escalation drops proportionally.
```text
Hypothesis
Layer sensitivity to quantization error is strongly non-uniform: degrading
a single depth-group to 4-bit (rest bf16) hurts agreement unevenly across
groups (≥3× spread between most and least sensitive), and symmetrically,
repairing only the most sensitive ~25% of layers (bf16 in-group, 4-bit
elsewhere) recovers a disproportionate share — ≥40% — of the agreement
lost by the all-4-bit model.
Falsification criterion
If per-group degradation effects are near-uniform (<2× spread), or if
repairing the best 25% of layers recovers <20% of the lost agreement
(i.e., error is diffuse and cooperative across depth), then
layer-restricted escalation cannot cut bytes-per-escalation materially
and C1 must rely entirely on temporal locality (expB) or block-level
selection (expE).
Method
Qwen3-1.7B bf16 reference, same 48-trajectory teacher-forced protocol
(benchmarks/datasets/eval_prompts.json, 128 tokens, greedy reference).
28 transformer layers → 7 contiguous depth groups of 4.
(i) DEGRADE-ONE: quantize (affine g64, 4-bit) all divisible Linear
layers of one group; rest bf16. 7 runs → sensitivity map.
(ii) REPAIR-ONE: all layers 4-bit except one group at bf16. 7 runs →
repair-value map, plus the all-4-bit floor (from expD stage0 B).
(iii) REPAIR-TOP-K: bf16 for the k most-repairing groups (k=1,2), 4-bit
elsewhere → cumulative repair curve vs bytes.
Metrics per config: agreement with reference, mean KL. Embedding/head
layers excluded (kept bf16 throughout, as in expD/expG).
Baseline
All-bf16 (agreement=1 by construction) and all-4-bit (87.5%, expD
stage0 B) bracket every configuration. No straw men.
Result
REFUTED. Sensitivity spread across depth groups: 1.5× (kill line: <2×).
Top-2 groups (29% of layers) repair only 24% of lost agreement —
sub-proportional; repair values sum to 7.4 of 12.5 points lost (repair
is cooperative, not concentrated). Full numbers:
results/expF_error_accumulation/20260812T043945Z/ and analysis.md.
Interpretation
Quantization damage to decisions is diffuse across depth. Layer-
restricted escalation gives no leverage; expD's late-layer hidden-state
error is decision-benign (norms ≠ decisions). C1's bytes-per-escalation
must come from temporal locality (expB), block-level selection
(expA/expE), or batch amortization (→ convergence with C2).
Next experiment
expA (block-granularity concentration) + expB (temporal stability).
```