- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays (two flips = 8 points = SEPARATION_LEAD). - R2: weights do apply (research scorer == production at V0); V1/V2 are identity at ±30/±60 by construction and leave six-question metrics unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not pass the gate. - R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted (all-pairs adds no dated probes); the bottleneck is monthly evaluation. New finding recorded as BUG-1048 (investigating): _boundary_windows year-straddle exemption and positional zip misalignment bypass the gate. - Dated errata appended (no deletions) to the 09-14/09-16 briefs and research docs; README board row -> 待验收. No production code, scoring, thresholds, gates or Skill changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
8.1 KiB
8.1 KiB
进度 · 生时校正三项离线研究(2026-09-26)
范围
- 任务书:
docs/tasks/TASK-rectification-offline-research-20260926.md - 分支:
codex/rectification-offline-research-20260926,worktree.worktrees/rectification-offline-research-20260926。开工基线origin/staging@abb05b67;测量在快进后的7ddce2c9上完成;提交最后变基到7203d94e(其间只新增scripts/research/fewer_probes_card_replay.py与文档,引擎与打分代码未变,测量结果不受影响)。本地提交,未推送。 - 执行方式:产品授权 Claude 子代理直接执行。
- 报告:
docs/research/rectification_offline_research_2026_09_26.md(三节 + 一句话结论 + 是否立实现单) - 性质:离线研究。未改线上代码、打分、阈值、出题闸门、Skill、workflow。 研究脚本在一次调用内临时替换模块属性(
event_probes._representative_pairs、event_probes._boundary_windows、event_probes._evaluate_contexts计数包装、研究模块precision_gate_lib.varga_factor),调用结束立即恢复,脚本末尾断言已恢复,并断言线上闸门仍是 45 / 30。 - 数据:v4 公开 AA 开放集(20 例)与虚构时刻。没有真实用户资料。开放集回放,不是盲测,不是准确率。
开工预检
- 读了
AGENTS.md(Part A、Part B B4)、CLAUDE.md、任务书、closure / minute_resolution / cluster_width / precision_gate / holdout_v4 / guided_collect_holdout 各文档、BUG-692、event_probes.py(_boundary_windows、_representative_pairs、_union_boundary_dates、_evaluate_contexts、guided_collect_windows)、decision_policy.py、v4 回放脚本、上一轮审计的shift.py/repo_shift.py、docs/research/pre_work_error_ledger.md(重点 ERR-110 / ERR-111)。 python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45:exit 0。remote verified,碎片扫描、外部引擎适配器诊断、5 个 focused 目标全部通过。本机没有复现台账里 09-24 / 09-25 的.workbuddy镜像断言失败,与 09-26 BUG-1047 预检记录一致。无新台账条目。- 本机没有
.venv,系统python3(带 swisseph、pytest)运行全部命令。
完成
| 项 | 脚本 | 结果 | 耗时 |
|---|---|---|---|
| R1 答错容错 | scripts/research/answer_flip_tolerance.py |
docs/research/answer_flip_tolerance_2026_09_26.json |
~1 分钟 |
| R2 V1/V2 重跑 | scripts/research/varga_sensitivity_rerun.py |
docs/research/varga_sensitivity_rerun_2026_09_26.json |
~16 分钟 |
| R3a 每分钟位移 | scripts/research/dasha_shift_per_minute.py |
docs/research/dasha_shift_per_minute_2026_09_26.json |
1 秒 |
| R3c 配对推断 | scripts/research/representative_pairs_probe.py |
docs/research/representative_pairs_probe_2026_09_26.json |
~6 分钟 |
| 共用纯函数 | scripts/research/offline_research_20260926_lib.py |
— | — |
| 回归测试 | tests/test_offline_research_20260926.py(5 条,只测研究纯函数) |
5 passed | — |
全部运行错误 0 例。
关键数字
- R1:不答错的基线与 09-14 G0 逐项一致(0.80 / 0.55 / 0.40;15 / 33 / 56;20/20)。答错 1 题:真值在区间 1.000 / 0.989 / 0.981,头名 0.58 / 0.33 / 0.24(同批不答错为 0.94 / 0.61 / 0.44)。答错 2 题:真值在区间 0.996 / 0.926 / 0.898,头名 0.40 / 0.08 / 0.06;±30 有 9/18 例、±60 有 12/18 例至少一种组合挤出真值。真值从未被淘汰(需 3 次强冲突)。随机抽样(种子 20260926,每格 30 次)与穷举一致。
- R2:研究计分器 V0 与线上 0 / 2060 候选有差。V1 / V2 在 ±30 / ±60 上 0 个候选分数变化(系数全截断到 1 / 不丢分盘),±10 上 168 / 159 个候选分数变化,六题后指标与 V0 完全相同 →
no_benefit(已实测)。补充 V1n 头名 0.85 / 0.60 / 0.50,宽度 15 / 33 / 73,不过门。 - R3a:每分钟边界位移中位 3.82 天(p10 1.58、p90 5.00、范围 1.34–5.85,n=2000),差分与解析式一致,仓库
_vim_start_dates中位 4.05(n=299 子样本),边界整体平移(对齐后差值 ≤1 天)。45 天闸门 ↔ 中位 11.8 分钟(7.7–33.5)。 - R3c:全部两两配对 vs 线上配对,带日期题(截断前)4.85 vs 4.80(±10 初始)、3.25 vs 3.30(±10 刷新),其余半径差异 ≤0.1;没有带日期题的例子数完全相同 → 推断被推翻。边界月份能分开代表的比例:初始 3.2–4.5%,刷新约 0.4–0.5%。严格闸门(研究补丁)下 ±10 初始没有带日期题的例子 4 → 8。
同机复跑(BUG-985:只做同机 A/B,不写死跨机哈希)
| 脚本 | 复跑方式 | 结果 |
|---|---|---|
| R1 | PYTHONHASHSEED=12345 再跑一次 |
JSON 逐字节一致 |
| R3a | PYTHONHASHSEED=999 再跑一次 |
JSON 逐字节一致 |
| R3c | PYTHONHASHSEED=4242 … --no-write |
summary 与已存 JSON 完全一致 |
| R2 | PYTHONHASHSEED=4242 … --no-write |
proof 摘要与 12 行指标与已存 JSON 完全一致 |
文档改动(只加勘误段,不删原文)
docs/tasks/TASK-rectification-precision-adaptive-boundary-research-20260914.md:1.1 天 / 122 天每度 / 换算表docs/tasks/TASK-rectification-open-collect-invite-20260914.md:D1b 依据句docs/tasks/TASK-rectification-precision-gate-guided-collect-20260916.md:实证第 5 条「20 分钟≈22 天,题池必然为空」docs/research/precision_gate_2026_09_14.md:M1b V1/V2 →no_benefit(已实测)docs/research/rectification_minute_resolution_closure_2026_09_14.md:新增 §7 勘误与补充docs/BUG_HISTORY.md:BUG-692 加一条补充;新增 BUG-1048(investigating,闸门跨年豁免 + zip 错位,未修)。开工时最大号是 BUG-1047;合入前请再核对一次编号,避免与并行分支冲突。docs/tasks/README.md:本单状态改为「待验收」。
线上引导文案未改;改写建议写在报告 §R3「改正与文案建议」。
测试
python3 -m pytest tests/test_offline_research_20260926.py tests/test_bug_history_workflow.py:7 passedpython3 -m pytest tests/test_repo_privacy_markers.py(新文件暂存后):全部 passedpython3 -m pytest tests/test_rectification_validation_integrity_gate.py tests/test_minute_resolution_research.py tests/test_cluster_width_research.py:全部 passed(没有改冻结评分文件,见 ERR-110)- 快速门
python3 scripts/run_quality_gate.py --profile quick:Python 部分 948 passed / 1 skipped(含隐私、校正完整性门禁);随后npm test步骤 exit 127,原因是本 worktree 没有frontend/node_modules(环境缺口,本单没有前端改动),整体报Quality gate failed。不记为通过。快速门没有改动任何跟踪文件。
偏离与注意
- R2 加了一个任务书没有的补充方案 V1n。 原因:V1 按定义在宽窗上截断成恒等,只报告「V1/V2 无变化」回答不了「按敏感度配权有没有用」。V1n 单列、标「补充」,不参与原判定。
- R3c 加了「严格闸门」对照,由此发现 BUG-1048。只做测量,不修。任何修改都要按 ERR-110 重新冻结 sealed holdout 记录。
scripts/research/precision_gate_sweep.py重跑会整篇覆盖docs/research/precision_gate_2026_09_14.md,冲掉 09-15 与 09-26 两段手写勘误(既有风险,本单未改那个脚本)。- 六题回放沿用 09-14 设计,题目不随作答重出。R1 因此测不到「答错后后续题被带偏」的影响,报告里已写明。
- v4 回放复现不出真机上「六题后问完了」的现象(±10 刷新阶段平均还有 3.25 道新的带日期题)。真机原因需要真机统计(
TASK-rectification-telemetry-20260926.md),本单不下结论。 scripts/research/**与tests/**在门禁路径里,推送会触发 staging 门禁。本单未推送。
未做
- 未推送、未合入 staging、未部署(按指令)。
- 未改线上引导文案(任务书要求只给建议)。
- 未评估 BUG-1048 两个漏洞放出来的题目在真人作答下是否可靠。