- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays (two flips = 8 points = SEPARATION_LEAD). - R2: weights do apply (research scorer == production at V0); V1/V2 are identity at ±30/±60 by construction and leave six-question metrics unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not pass the gate. - R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted (all-pairs adds no dated probes); the bottleneck is monthly evaluation. New finding recorded as BUG-1048 (investigating): _boundary_windows year-straddle exemption and positional zip misalignment bypass the gate. - Dated errata appended (no deletions) to the 09-14/09-16 briefs and research docs; README board row -> 待验收. No production code, scoring, thresholds, gates or Skill changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
80 lines
8.1 KiB
Markdown
80 lines
8.1 KiB
Markdown
# 进度 · 生时校正三项离线研究(2026-09-26)
|
||
|
||
## 范围
|
||
|
||
- 任务书:`docs/tasks/TASK-rectification-offline-research-20260926.md`
|
||
- 分支:`codex/rectification-offline-research-20260926`,worktree `.worktrees/rectification-offline-research-20260926`。开工基线 `origin/staging` @ `abb05b67`;测量在快进后的 `7ddce2c9` 上完成;提交最后变基到 `7203d94e`(其间只新增 `scripts/research/fewer_probes_card_replay.py` 与文档,引擎与打分代码未变,测量结果不受影响)。**本地提交,未推送。**
|
||
- 执行方式:产品授权 Claude 子代理直接执行。
|
||
- 报告:`docs/research/rectification_offline_research_2026_09_26.md`(三节 + 一句话结论 + 是否立实现单)
|
||
- 性质:离线研究。**未改线上代码、打分、阈值、出题闸门、Skill、workflow。** 研究脚本在一次调用内临时替换模块属性(`event_probes._representative_pairs`、`event_probes._boundary_windows`、`event_probes._evaluate_contexts` 计数包装、研究模块 `precision_gate_lib.varga_factor`),调用结束立即恢复,脚本末尾断言已恢复,并断言线上闸门仍是 45 / 30。
|
||
- 数据:v4 公开 AA 开放集(20 例)与虚构时刻。没有真实用户资料。**开放集回放,不是盲测,不是准确率。**
|
||
|
||
## 开工预检
|
||
|
||
- 读了 `AGENTS.md`(Part A、Part B B4)、`CLAUDE.md`、任务书、closure / minute_resolution / cluster_width / precision_gate / holdout_v4 / guided_collect_holdout 各文档、BUG-692、`event_probes.py`(`_boundary_windows`、`_representative_pairs`、`_union_boundary_dates`、`_evaluate_contexts`、`guided_collect_windows`)、`decision_policy.py`、v4 回放脚本、上一轮审计的 `shift.py` / `repo_shift.py`、`docs/research/pre_work_error_ledger.md`(重点 ERR-110 / ERR-111)。
|
||
- `python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45`:**exit 0**。remote verified,碎片扫描、外部引擎适配器诊断、5 个 focused 目标全部通过。本机没有复现台账里 09-24 / 09-25 的 `.workbuddy` 镜像断言失败,与 09-26 BUG-1047 预检记录一致。无新台账条目。
|
||
- 本机没有 `.venv`,系统 `python3`(带 swisseph、pytest)运行全部命令。
|
||
|
||
## 完成
|
||
|
||
| 项 | 脚本 | 结果 | 耗时 |
|
||
| --- | --- | --- | ---: |
|
||
| R1 答错容错 | `scripts/research/answer_flip_tolerance.py` | `docs/research/answer_flip_tolerance_2026_09_26.json` | ~1 分钟 |
|
||
| R2 V1/V2 重跑 | `scripts/research/varga_sensitivity_rerun.py` | `docs/research/varga_sensitivity_rerun_2026_09_26.json` | ~16 分钟 |
|
||
| R3a 每分钟位移 | `scripts/research/dasha_shift_per_minute.py` | `docs/research/dasha_shift_per_minute_2026_09_26.json` | 1 秒 |
|
||
| R3c 配对推断 | `scripts/research/representative_pairs_probe.py` | `docs/research/representative_pairs_probe_2026_09_26.json` | ~6 分钟 |
|
||
| 共用纯函数 | `scripts/research/offline_research_20260926_lib.py` | — | — |
|
||
| 回归测试 | `tests/test_offline_research_20260926.py`(5 条,只测研究纯函数) | 5 passed | — |
|
||
|
||
全部运行错误 0 例。
|
||
|
||
## 关键数字
|
||
|
||
- **R1**:不答错的基线与 09-14 G0 逐项一致(0.80 / 0.55 / 0.40;15 / 33 / 56;20/20)。答错 1 题:真值在区间 1.000 / 0.989 / 0.981,头名 0.58 / 0.33 / 0.24(同批不答错为 0.94 / 0.61 / 0.44)。答错 2 题:真值在区间 0.996 / 0.926 / 0.898,头名 0.40 / 0.08 / 0.06;±30 有 9/18 例、±60 有 12/18 例至少一种组合挤出真值。真值从未被淘汰(需 3 次强冲突)。随机抽样(种子 20260926,每格 30 次)与穷举一致。
|
||
- **R2**:研究计分器 V0 与线上 0 / 2060 候选有差。V1 / V2 在 ±30 / ±60 上 0 个候选分数变化(系数全截断到 1 / 不丢分盘),±10 上 168 / 159 个候选分数变化,六题后指标与 V0 完全相同 → `no_benefit`(已实测)。补充 V1n 头名 0.85 / 0.60 / 0.50,宽度 15 / 33 / 73,不过门。
|
||
- **R3a**:每分钟边界位移中位 3.82 天(p10 1.58、p90 5.00、范围 1.34–5.85,n=2000),差分与解析式一致,仓库 `_vim_start_dates` 中位 4.05(n=299 子样本),边界整体平移(对齐后差值 ≤1 天)。45 天闸门 ↔ 中位 11.8 分钟(7.7–33.5)。
|
||
- **R3c**:全部两两配对 vs 线上配对,带日期题(截断前)4.85 vs 4.80(±10 初始)、3.25 vs 3.30(±10 刷新),其余半径差异 ≤0.1;没有带日期题的例子数完全相同 → 推断被推翻。边界月份能分开代表的比例:初始 3.2–4.5%,刷新约 0.4–0.5%。严格闸门(研究补丁)下 ±10 初始没有带日期题的例子 4 → 8。
|
||
|
||
## 同机复跑(BUG-985:只做同机 A/B,不写死跨机哈希)
|
||
|
||
| 脚本 | 复跑方式 | 结果 |
|
||
| --- | --- | --- |
|
||
| R1 | `PYTHONHASHSEED=12345` 再跑一次 | JSON 逐字节一致 |
|
||
| R3a | `PYTHONHASHSEED=999` 再跑一次 | JSON 逐字节一致 |
|
||
| R3c | `PYTHONHASHSEED=4242 … --no-write` | summary 与已存 JSON 完全一致 |
|
||
| R2 | `PYTHONHASHSEED=4242 … --no-write` | proof 摘要与 12 行指标与已存 JSON 完全一致 |
|
||
|
||
## 文档改动(只加勘误段,不删原文)
|
||
|
||
- `docs/tasks/TASK-rectification-precision-adaptive-boundary-research-20260914.md`:1.1 天 / 122 天每度 / 换算表
|
||
- `docs/tasks/TASK-rectification-open-collect-invite-20260914.md`:D1b 依据句
|
||
- `docs/tasks/TASK-rectification-precision-gate-guided-collect-20260916.md`:实证第 5 条「20 分钟≈22 天,题池必然为空」
|
||
- `docs/research/precision_gate_2026_09_14.md`:M1b V1/V2 → `no_benefit`(已实测)
|
||
- `docs/research/rectification_minute_resolution_closure_2026_09_14.md`:新增 §7 勘误与补充
|
||
- `docs/BUG_HISTORY.md`:BUG-692 加一条补充;新增 **BUG-1048**(`investigating`,闸门跨年豁免 + zip 错位,未修)。开工时最大号是 BUG-1047;合入前请再核对一次编号,避免与并行分支冲突。
|
||
- `docs/tasks/README.md`:本单状态改为「待验收」。
|
||
|
||
线上引导文案**未改**;改写建议写在报告 §R3「改正与文案建议」。
|
||
|
||
## 测试
|
||
|
||
- `python3 -m pytest tests/test_offline_research_20260926.py tests/test_bug_history_workflow.py`:7 passed
|
||
- `python3 -m pytest tests/test_repo_privacy_markers.py`(新文件暂存后):全部 passed
|
||
- `python3 -m pytest tests/test_rectification_validation_integrity_gate.py tests/test_minute_resolution_research.py tests/test_cluster_width_research.py`:全部 passed(没有改冻结评分文件,见 ERR-110)
|
||
- 快速门 `python3 scripts/run_quality_gate.py --profile quick`:Python 部分 **948 passed / 1 skipped**(含隐私、校正完整性门禁);随后 `npm test` 步骤 exit 127,原因是本 worktree 没有 `frontend/node_modules`(环境缺口,本单没有前端改动),整体报 `Quality gate failed`。不记为通过。快速门没有改动任何跟踪文件。
|
||
|
||
## 偏离与注意
|
||
|
||
1. **R2 加了一个任务书没有的补充方案 V1n。** 原因:V1 按定义在宽窗上截断成恒等,只报告「V1/V2 无变化」回答不了「按敏感度配权有没有用」。V1n 单列、标「补充」,不参与原判定。
|
||
2. **R3c 加了「严格闸门」对照**,由此发现 BUG-1048。只做测量,不修。任何修改都要按 ERR-110 重新冻结 sealed holdout 记录。
|
||
3. `scripts/research/precision_gate_sweep.py` 重跑会整篇覆盖 `docs/research/precision_gate_2026_09_14.md`,冲掉 09-15 与 09-26 两段手写勘误(既有风险,本单未改那个脚本)。
|
||
4. 六题回放沿用 09-14 设计,题目不随作答重出。R1 因此测不到「答错后后续题被带偏」的影响,报告里已写明。
|
||
5. v4 回放复现不出真机上「六题后问完了」的现象(±10 刷新阶段平均还有 3.25 道新的带日期题)。真机原因需要真机统计(`TASK-rectification-telemetry-20260926.md`),本单不下结论。
|
||
6. `scripts/research/**` 与 `tests/**` 在门禁路径里,推送会触发 staging 门禁。本单未推送。
|
||
|
||
## 未做
|
||
|
||
- 未推送、未合入 staging、未部署(按指令)。
|
||
- 未改线上引导文案(任务书要求只给建议)。
|
||
- 未评估 BUG-1048 两个漏洞放出来的题目在真人作答下是否可靠。
|