Files
Jyotisha/docs/tasks/PROGRESS-rectification-offline-research-20260926.md
T
Jesse_ChenandClaude Opus 5.5 0f5442cea2 research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
  a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
  (two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
  identity at ±30/±60 by construction and leave six-question metrics
  unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
  pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
  gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
  (all-pairs adds no dated probes); the bottleneck is monthly evaluation.
  New finding recorded as BUG-1048 (investigating): _boundary_windows
  year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
  research docs; README board row -> 待验收. No production code, scoring,
  thresholds, gates or Skill changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:47:48 +08:00

8.1 KiB
Raw Blame History

进度 · 生时校正三项离线研究(2026-09-26)

范围

  • 任务书:docs/tasks/TASK-rectification-offline-research-20260926.md
  • 分支:codex/rectification-offline-research-20260926,worktree .worktrees/rectification-offline-research-20260926。开工基线 origin/staging @ abb05b67;测量在快进后的 7ddce2c9 上完成;提交最后变基到 7203d94e(其间只新增 scripts/research/fewer_probes_card_replay.py 与文档,引擎与打分代码未变,测量结果不受影响)。本地提交,未推送。
  • 执行方式:产品授权 Claude 子代理直接执行。
  • 报告:docs/research/rectification_offline_research_2026_09_26.md(三节 + 一句话结论 + 是否立实现单)
  • 性质:离线研究。未改线上代码、打分、阈值、出题闸门、Skill、workflow。 研究脚本在一次调用内临时替换模块属性(event_probes._representative_pairs、event_probes._boundary_windows、event_probes._evaluate_contexts 计数包装、研究模块 precision_gate_lib.varga_factor),调用结束立即恢复,脚本末尾断言已恢复,并断言线上闸门仍是 45 / 30。
  • 数据:v4 公开 AA 开放集(20 例)与虚构时刻。没有真实用户资料。开放集回放,不是盲测,不是准确率。

开工预检

  • 读了 AGENTS.md(Part A、Part B B4)、CLAUDE.md、任务书、closure / minute_resolution / cluster_width / precision_gate / holdout_v4 / guided_collect_holdout 各文档、BUG-692、event_probes.py(_boundary_windows、_representative_pairs、_union_boundary_dates、_evaluate_contexts、guided_collect_windows)、decision_policy.py、v4 回放脚本、上一轮审计的 shift.py / repo_shift.py、docs/research/pre_work_error_ledger.md(重点 ERR-110 / ERR-111)。
  • python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45:exit 0。remote verified,碎片扫描、外部引擎适配器诊断、5 个 focused 目标全部通过。本机没有复现台账里 09-24 / 09-25 的 .workbuddy 镜像断言失败,与 09-26 BUG-1047 预检记录一致。无新台账条目。
  • 本机没有 .venv,系统 python3(带 swisseph、pytest)运行全部命令。

完成

项 脚本 结果 耗时
R1 答错容错 scripts/research/answer_flip_tolerance.py docs/research/answer_flip_tolerance_2026_09_26.json ~1 分钟
R2 V1/V2 重跑 scripts/research/varga_sensitivity_rerun.py docs/research/varga_sensitivity_rerun_2026_09_26.json ~16 分钟
R3a 每分钟位移 scripts/research/dasha_shift_per_minute.py docs/research/dasha_shift_per_minute_2026_09_26.json 1 秒
R3c 配对推断 scripts/research/representative_pairs_probe.py docs/research/representative_pairs_probe_2026_09_26.json ~6 分钟
共用纯函数 scripts/research/offline_research_20260926_lib.py — —
回归测试 tests/test_offline_research_20260926.py(5 条,只测研究纯函数) 5 passed —

全部运行错误 0 例。

关键数字

  • R1:不答错的基线与 09-14 G0 逐项一致(0.80 / 0.55 / 0.40;15 / 33 / 56;20/20)。答错 1 题:真值在区间 1.000 / 0.989 / 0.981,头名 0.58 / 0.33 / 0.24(同批不答错为 0.94 / 0.61 / 0.44)。答错 2 题:真值在区间 0.996 / 0.926 / 0.898,头名 0.40 / 0.08 / 0.06;±30 有 9/18 例、±60 有 12/18 例至少一种组合挤出真值。真值从未被淘汰(需 3 次强冲突)。随机抽样(种子 20260926,每格 30 次)与穷举一致。
  • R2:研究计分器 V0 与线上 0 / 2060 候选有差。V1 / V2 在 ±30 / ±60 上 0 个候选分数变化(系数全截断到 1 / 不丢分盘),±10 上 168 / 159 个候选分数变化,六题后指标与 V0 完全相同 → no_benefit(已实测)。补充 V1n 头名 0.85 / 0.60 / 0.50,宽度 15 / 33 / 73,不过门。
  • R3a:每分钟边界位移中位 3.82 天(p10 1.58、p90 5.00、范围 1.34–5.85,n=2000),差分与解析式一致,仓库 _vim_start_dates 中位 4.05(n=299 子样本),边界整体平移(对齐后差值 ≤1 天)。45 天闸门 ↔ 中位 11.8 分钟(7.7–33.5)。
  • R3c:全部两两配对 vs 线上配对,带日期题(截断前)4.85 vs 4.80(±10 初始)、3.25 vs 3.30(±10 刷新),其余半径差异 ≤0.1;没有带日期题的例子数完全相同 → 推断被推翻。边界月份能分开代表的比例:初始 3.2–4.5%,刷新约 0.4–0.5%。严格闸门(研究补丁)下 ±10 初始没有带日期题的例子 4 → 8。

同机复跑(BUG-985:只做同机 A/B,不写死跨机哈希)

脚本 复跑方式 结果
R1 PYTHONHASHSEED=12345 再跑一次 JSON 逐字节一致
R3a PYTHONHASHSEED=999 再跑一次 JSON 逐字节一致
R3c PYTHONHASHSEED=4242 … --no-write summary 与已存 JSON 完全一致
R2 PYTHONHASHSEED=4242 … --no-write proof 摘要与 12 行指标与已存 JSON 完全一致

文档改动(只加勘误段,不删原文)

  • docs/tasks/TASK-rectification-precision-adaptive-boundary-research-20260914.md:1.1 天 / 122 天每度 / 换算表
  • docs/tasks/TASK-rectification-open-collect-invite-20260914.md:D1b 依据句
  • docs/tasks/TASK-rectification-precision-gate-guided-collect-20260916.md:实证第 5 条「20 分钟≈22 天,题池必然为空」
  • docs/research/precision_gate_2026_09_14.md:M1b V1/V2 → no_benefit(已实测)
  • docs/research/rectification_minute_resolution_closure_2026_09_14.md:新增 §7 勘误与补充
  • docs/BUG_HISTORY.md:BUG-692 加一条补充;新增 BUG-1048(investigating,闸门跨年豁免 + zip 错位,未修)。开工时最大号是 BUG-1047;合入前请再核对一次编号,避免与并行分支冲突。
  • docs/tasks/README.md:本单状态改为「待验收」。

线上引导文案未改;改写建议写在报告 §R3「改正与文案建议」。

测试

  • python3 -m pytest tests/test_offline_research_20260926.py tests/test_bug_history_workflow.py:7 passed
  • python3 -m pytest tests/test_repo_privacy_markers.py(新文件暂存后):全部 passed
  • python3 -m pytest tests/test_rectification_validation_integrity_gate.py tests/test_minute_resolution_research.py tests/test_cluster_width_research.py:全部 passed(没有改冻结评分文件,见 ERR-110)
  • 快速门 python3 scripts/run_quality_gate.py --profile quick:Python 部分 948 passed / 1 skipped(含隐私、校正完整性门禁);随后 npm test 步骤 exit 127,原因是本 worktree 没有 frontend/node_modules(环境缺口,本单没有前端改动),整体报 Quality gate failed。不记为通过。快速门没有改动任何跟踪文件。

偏离与注意

  1. R2 加了一个任务书没有的补充方案 V1n。 原因:V1 按定义在宽窗上截断成恒等,只报告「V1/V2 无变化」回答不了「按敏感度配权有没有用」。V1n 单列、标「补充」,不参与原判定。
  2. R3c 加了「严格闸门」对照,由此发现 BUG-1048。只做测量,不修。任何修改都要按 ERR-110 重新冻结 sealed holdout 记录。
  3. scripts/research/precision_gate_sweep.py 重跑会整篇覆盖 docs/research/precision_gate_2026_09_14.md,冲掉 09-15 与 09-26 两段手写勘误(既有风险,本单未改那个脚本)。
  4. 六题回放沿用 09-14 设计,题目不随作答重出。R1 因此测不到「答错后后续题被带偏」的影响,报告里已写明。
  5. v4 回放复现不出真机上「六题后问完了」的现象(±10 刷新阶段平均还有 3.25 道新的带日期题)。真机原因需要真机统计(TASK-rectification-telemetry-20260926.md),本单不下结论。
  6. scripts/research/** 与 tests/** 在门禁路径里,推送会触发 staging 门禁。本单未推送。

未做

  • 未推送、未合入 staging、未部署(按指令)。
  • 未改线上引导文案(任务书要求只给建议)。
  • 未评估 BUG-1048 两个漏洞放出来的题目在真人作答下是否可靠。