- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
(two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
identity at ±30/±60 by construction and leave six-question metrics
unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
(all-pairs adds no dated probes); the bottleneck is monthly evaluation.
New finding recorded as BUG-1048 (investigating): _boundary_windows
year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
research docs; README board row -> 待验收. No production code, scoring,
thresholds, gates or Skill changed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8