Files
Jyotisha/docs/tasks/PROGRESS-rectification-offline-research-20260926.md
T
Jesse_ChenandClaude Opus 5.5 0f5442cea2 research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
  a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
  (two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
  identity at ±30/±60 by construction and leave six-question metrics
  unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
  pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
  gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
  (all-pairs adds no dated probes); the bottleneck is monthly evaluation.
  New finding recorded as BUG-1048 (investigating): _boundary_windows
  year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
  research docs; README board row -> 待验收. No production code, scoring,
  thresholds, gates or Skill changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:47:48 +08:00

80 lines
8.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 进度 · 生时校正三项离线研究(2026-09-26)
## 范围
- 任务书:`docs/tasks/TASK-rectification-offline-research-20260926.md`
- 分支:`codex/rectification-offline-research-20260926`,worktree `.worktrees/rectification-offline-research-20260926`。开工基线 `origin/staging` @ `abb05b67`;测量在快进后的 `7ddce2c9` 上完成;提交最后变基到 `7203d94e`(其间只新增 `scripts/research/fewer_probes_card_replay.py` 与文档,引擎与打分代码未变,测量结果不受影响)。**本地提交,未推送。**
- 执行方式:产品授权 Claude 子代理直接执行。
- 报告:`docs/research/rectification_offline_research_2026_09_26.md`(三节 + 一句话结论 + 是否立实现单)
- 性质:离线研究。**未改线上代码、打分、阈值、出题闸门、Skill、workflow。** 研究脚本在一次调用内临时替换模块属性(`event_probes._representative_pairs`、`event_probes._boundary_windows`、`event_probes._evaluate_contexts` 计数包装、研究模块 `precision_gate_lib.varga_factor`),调用结束立即恢复,脚本末尾断言已恢复,并断言线上闸门仍是 45 / 30。
- 数据:v4 公开 AA 开放集(20 例)与虚构时刻。没有真实用户资料。**开放集回放,不是盲测,不是准确率。**
## 开工预检
- 读了 `AGENTS.md`(Part A、Part B B4)、`CLAUDE.md`、任务书、closure / minute_resolution / cluster_width / precision_gate / holdout_v4 / guided_collect_holdout 各文档、BUG-692、`event_probes.py`(`_boundary_windows`、`_representative_pairs`、`_union_boundary_dates`、`_evaluate_contexts`、`guided_collect_windows`)、`decision_policy.py`、v4 回放脚本、上一轮审计的 `shift.py` / `repo_shift.py`、`docs/research/pre_work_error_ledger.md`(重点 ERR-110 / ERR-111)。
- `python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45`:**exit 0**。remote verified,碎片扫描、外部引擎适配器诊断、5 个 focused 目标全部通过。本机没有复现台账里 09-24 / 09-25 的 `.workbuddy` 镜像断言失败,与 09-26 BUG-1047 预检记录一致。无新台账条目。
- 本机没有 `.venv`,系统 `python3`(带 swisseph、pytest)运行全部命令。
## 完成
| 项 | 脚本 | 结果 | 耗时 |
| --- | --- | --- | ---: |
| R1 答错容错 | `scripts/research/answer_flip_tolerance.py` | `docs/research/answer_flip_tolerance_2026_09_26.json` | ~1 分钟 |
| R2 V1/V2 重跑 | `scripts/research/varga_sensitivity_rerun.py` | `docs/research/varga_sensitivity_rerun_2026_09_26.json` | ~16 分钟 |
| R3a 每分钟位移 | `scripts/research/dasha_shift_per_minute.py` | `docs/research/dasha_shift_per_minute_2026_09_26.json` | 1 秒 |
| R3c 配对推断 | `scripts/research/representative_pairs_probe.py` | `docs/research/representative_pairs_probe_2026_09_26.json` | ~6 分钟 |
| 共用纯函数 | `scripts/research/offline_research_20260926_lib.py` | — | — |
| 回归测试 | `tests/test_offline_research_20260926.py`(5 条,只测研究纯函数) | 5 passed | — |
全部运行错误 0 例。
## 关键数字
- **R1**:不答错的基线与 09-14 G0 逐项一致(0.80 / 0.55 / 0.40;15 / 33 / 56;20/20)。答错 1 题:真值在区间 1.000 / 0.989 / 0.981,头名 0.58 / 0.33 / 0.24(同批不答错为 0.94 / 0.61 / 0.44)。答错 2 题:真值在区间 0.996 / 0.926 / 0.898,头名 0.40 / 0.08 / 0.06;±30 有 9/18 例、±60 有 12/18 例至少一种组合挤出真值。真值从未被淘汰(需 3 次强冲突)。随机抽样(种子 20260926,每格 30 次)与穷举一致。
- **R2**:研究计分器 V0 与线上 0 / 2060 候选有差。V1 / V2 在 ±30 / ±60 上 0 个候选分数变化(系数全截断到 1 / 不丢分盘),±10 上 168 / 159 个候选分数变化,六题后指标与 V0 完全相同 → `no_benefit`(已实测)。补充 V1n 头名 0.85 / 0.60 / 0.50,宽度 15 / 33 / 73,不过门。
- **R3a**:每分钟边界位移中位 3.82 天(p10 1.58、p90 5.00、范围 1.34–5.85,n=2000),差分与解析式一致,仓库 `_vim_start_dates` 中位 4.05(n=299 子样本),边界整体平移(对齐后差值 ≤1 天)。45 天闸门 ↔ 中位 11.8 分钟(7.7–33.5)。
- **R3c**:全部两两配对 vs 线上配对,带日期题(截断前)4.85 vs 4.80(±10 初始)、3.25 vs 3.30(±10 刷新),其余半径差异 ≤0.1;没有带日期题的例子数完全相同 → 推断被推翻。边界月份能分开代表的比例:初始 3.2–4.5%,刷新约 0.4–0.5%。严格闸门(研究补丁)下 ±10 初始没有带日期题的例子 4 → 8。
## 同机复跑(BUG-985:只做同机 A/B,不写死跨机哈希)
| 脚本 | 复跑方式 | 结果 |
| --- | --- | --- |
| R1 | `PYTHONHASHSEED=12345` 再跑一次 | JSON 逐字节一致 |
| R3a | `PYTHONHASHSEED=999` 再跑一次 | JSON 逐字节一致 |
| R3c | `PYTHONHASHSEED=4242 … --no-write` | summary 与已存 JSON 完全一致 |
| R2 | `PYTHONHASHSEED=4242 … --no-write` | proof 摘要与 12 行指标与已存 JSON 完全一致 |
## 文档改动(只加勘误段,不删原文)
- `docs/tasks/TASK-rectification-precision-adaptive-boundary-research-20260914.md`:1.1 天 / 122 天每度 / 换算表
- `docs/tasks/TASK-rectification-open-collect-invite-20260914.md`:D1b 依据句
- `docs/tasks/TASK-rectification-precision-gate-guided-collect-20260916.md`:实证第 5 条「20 分钟≈22 天,题池必然为空」
- `docs/research/precision_gate_2026_09_14.md`:M1b V1/V2 → `no_benefit`(已实测)
- `docs/research/rectification_minute_resolution_closure_2026_09_14.md`:新增 §7 勘误与补充
- `docs/BUG_HISTORY.md`:BUG-692 加一条补充;新增 **BUG-1048**(`investigating`,闸门跨年豁免 + zip 错位,未修)。开工时最大号是 BUG-1047;合入前请再核对一次编号,避免与并行分支冲突。
- `docs/tasks/README.md`:本单状态改为「待验收」。
线上引导文案**未改**;改写建议写在报告 §R3「改正与文案建议」。
## 测试
- `python3 -m pytest tests/test_offline_research_20260926.py tests/test_bug_history_workflow.py`:7 passed
- `python3 -m pytest tests/test_repo_privacy_markers.py`(新文件暂存后):全部 passed
- `python3 -m pytest tests/test_rectification_validation_integrity_gate.py tests/test_minute_resolution_research.py tests/test_cluster_width_research.py`:全部 passed(没有改冻结评分文件,见 ERR-110)
- 快速门 `python3 scripts/run_quality_gate.py --profile quick`:Python 部分 **948 passed / 1 skipped**(含隐私、校正完整性门禁);随后 `npm test` 步骤 exit 127,原因是本 worktree 没有 `frontend/node_modules`(环境缺口,本单没有前端改动),整体报 `Quality gate failed`。不记为通过。快速门没有改动任何跟踪文件。
## 偏离与注意
1. **R2 加了一个任务书没有的补充方案 V1n。** 原因:V1 按定义在宽窗上截断成恒等,只报告「V1/V2 无变化」回答不了「按敏感度配权有没有用」。V1n 单列、标「补充」,不参与原判定。
2. **R3c 加了「严格闸门」对照**,由此发现 BUG-1048。只做测量,不修。任何修改都要按 ERR-110 重新冻结 sealed holdout 记录。
3. `scripts/research/precision_gate_sweep.py` 重跑会整篇覆盖 `docs/research/precision_gate_2026_09_14.md`,冲掉 09-15 与 09-26 两段手写勘误(既有风险,本单未改那个脚本)。
4. 六题回放沿用 09-14 设计,题目不随作答重出。R1 因此测不到「答错后后续题被带偏」的影响,报告里已写明。
5. v4 回放复现不出真机上「六题后问完了」的现象(±10 刷新阶段平均还有 3.25 道新的带日期题)。真机原因需要真机统计(`TASK-rectification-telemetry-20260926.md`),本单不下结论。
6. `scripts/research/**` 与 `tests/**` 在门禁路径里,推送会触发 staging 门禁。本单未推送。
## 未做
- 未推送、未合入 staging、未部署(按指令)。
- 未改线上引导文案(任务书要求只给建议)。
- 未评估 BUG-1048 两个漏洞放出来的题目在真人作答下是否可靠。