research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays (two flips = 8 points = SEPARATION_LEAD). - R2: weights do apply (research scorer == production at V0); V1/V2 are identity at ±30/±60 by construction and leave six-question metrics unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not pass the gate. - R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted (all-pairs adds no dated probes); the bottleneck is monthly evaluation. New finding recorded as BUG-1048 (investigating): _boundary_windows year-straddle exemption and positional zip misalignment bypass the gate. - Dated errata appended (no deletions) to the 09-14/09-16 briefs and research docs; README board row -> 待验收. No production code, scoring, thresholds, gates or Skill changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
7203d94e9c
commit
0f5442cea2
@@ -37,7 +37,7 @@
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `TASK-rectification-fewer-probes-card-20260926.md` | [PROGRESS](PROGRESS-rectification-fewer-probes-card-20260926.md) | **减少无效追问 + 卡片区间为主**:引导补件离线新增达标 0 → 上限 6→2、定向题问完门槛未达直接出卡;卡片区间为主标题、代表分钟副标题,第一二名差距 ≥5 个百分点才显示百分比,否则写「目前区分不开」。不放宽任何置信度。先做 | 待验收 | `codex/rectification-fewer-probes-card-20260926`(未推送;D1 改在前端每校正 2 道,引擎常量因冻结评分身份未动;回放真值不降、宽度中位 ±1 分钟、提问 11.4→7.8;Skill 10.0.30) |
|
||||
| `TASK-rectification-telemetry-20260926.md` | — | **匿名聚合统计**:每会话一行只存数字 / 枚举(题数分类、宽度、差距、停止原因、门槛达标、耗时、版本),管理后台只看汇总、保留 180 天;动表须 test:db。排在 fewer-probes 后 | 待领取 | — |
|
||||
| `TASK-rectification-offline-research-20260926.md` | — | **三项离线研究**:答错 1–2 题的容错、V1/V2 分盘配权正确重跑、改正「1 分钟≈1.1 天」(实测中位 3.8 天)并核实 `_representative_pairs` 推断。不改线上 | 待领取 | — |
|
||||
| `TASK-rectification-offline-research-20260926.md` | `PROGRESS-rectification-offline-research-20260926.md` | **三项离线研究**:答错 1–2 题的容错、V1/V2 分盘配权正确重跑、改正「1 分钟≈1.1 天」(实测中位 3.8 天)并核实 `_representative_pairs` 推断。不改线上。结论(`docs/research/rectification_offline_research_2026_09_26.md`):R1 答错 1 题真值在区间 98–100%、头名降三到四成,答错 2 题 ±30/±60 挤出 7–10%(两道反答=8 分=淘汰线);R2 权重生效,V1/V2 在 ±30/±60 按定义恒等、±10 六题后指标不变 → `no_benefit`(已实测);R3 3.8 天/分钟复现,45 天闸≈8–34 分钟,`_representative_pairs` 推断被推翻(全配对题数不变,卡在逐月评估),另记 BUG-1048 `investigating`(闸门跨年豁免 + zip 错位)。三项均不建议立实现单 | 待验收 | `codex/rectification-offline-research-20260926`(本地未推) |
|
||||
| `TASK-rectification-code-split-20260926.md` | — | **代码拆分(只搬不改)**:聊天组件 2043 行 / 35 useState、`POST` 926 行、`runV9AgentTurn` 1047 行,269 处切源码测试;拆分 + 增长合同 + 切片测试改调用函数。排在 fewer-probes、telemetry 之后 | 待领取 | — |
|
||||
| `TASK-rectification-dup-question-20260926.md` | `PROGRESS-rectification-dup-question-20260926.md` | **同一轮问题出现两次(BUG-1045,复发自 BUG-585,BUG-969 拼回题干、去重只在刷新路径)+ 同一道选择题连画两张(BUG-1046:提交失败不回滚本地已答 + 兜底问题块条件过宽;H2 漏收回合)**。先于 latency 单 | 已验收(Claude 09-26 直接执行:子代理复现 A 与 B-H1(选择题提交 409 / 网络错误不回滚本地已答);Claude 变基到含 BUG-1043/1044 的 staging 后独立复验 tsc/lint 0、全量 3981 条失败名单与基线逐条一致、四路由 ○、gzip 不变) | `e4c1c7a3`(已部署 `f1d16405`,health 一致) |
|
||||
| `TASK-rectification-latency-20260926.md` | `PROGRESS-rectification-latency-20260926.md` | **每轮等待过长(BUG-1047)**:一轮打字回答串行 5 次开思考的模型调用,分类在开流前且无超时。产品定:收尾两步不动、分类保持思考只加 10 秒超时、发出后立即出确定性进度句并按阶段更新;补埋点;服务端无口吻优化须 A/B 逐位一致。排在 dup-question 之后 | 已验收(Claude 09-26 直接执行:子代理实现 D1–D4 与 D5 两项;Claude 用 Node 22 独立复验 tsc/lint 0、全量 4002 条失败名单与 Node 22 基线逐条一致(24 条均为 Docker/DB)、四路由 ○、gzip 不变。D5 第 1 项(探针复用)输出逐字节一致但触发冻结打分身份,Claude 建议暂不做) | `86ff9a40`(已部署 `8a409434`,health 一致) |
|
||||
|
||||
Reference in New Issue
Block a user