Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
141 lines
19 KiB
Markdown
141 lines
19 KiB
Markdown
# PROGRESS · 生时校正:旁白事实有据 + 去掉重试按钮 + 给模型的资料瘦身 · 2026-09-27
|
||
|
||
- 执行方式:直接执行(产品负责人授权子代理执行;Claude 事后独立验收)
|
||
- 任务书:`docs/tasks/TASK-rectification-grounding-20260927.md`(`a98d3b20`)
|
||
- 基线:开工 `origin/staging` `a98d3b20`(诊断基于 `599fe7a9`);中途 staging 合入数据卡单(`cd4dde9d`),本分支已 rebase 到 `cd4dde9d`,最终对比基线全部在 `cd4dde9d` 上重跑;随后再 rebase 到 `19376089`(只改 `docs/BUG_HISTORY.md` 的纯文档提交,代码树与 `cd4dde9d` 相同)。
|
||
- 分支:`codex/rectification-grounding-20260927`,工作树 `.worktrees/rectification-grounding-20260927`(本地提交,未推送)
|
||
- BUG 编号:预留 BUG-1055~1058,开工与提交时核对(开工最大号 1054;rebase 后 staging 有 1059,本单 1055~1058 未被占用)
|
||
- Skill:不 bump。`jyotish-birth-time-rectification` 10.0.31 文本未改,只是证据轮按章节切片发送;registry 不动。
|
||
- 数据库:不动。`regenerate_agentic_rectification_turn` 保留,**待另开单退役**(AGENTS §7.6 两轮规则:本轮先下线代码)。
|
||
|
||
## 提交
|
||
|
||
| 任务 | 提交 | 内容 |
|
||
| --- | --- | --- |
|
||
| T1 | `0e0baaa7` | BUG-1055:事实句不参与裁句、不先流后撤;P3 数字白名单(`spoken-grounding.ts`);batch 返回 `range_after_rescore`;系统提示「范围由系统说」 |
|
||
| T4 | `5a808449` | BUG-1058:`deferFollowup` 路径把答题前范围交给 Agent 轮(`rangeBeforeTurn`);路由级回归 |
|
||
| T2 | `84eb2315` | BUG-1056:删除校正重新生成(路由、服务端模块、Agent、只读工具集、客户端动作与状态);`ChatMessageActions` 无 `onRegenerate` 不画按钮;DESIGN / VOICE |
|
||
| T3 | `13d93020`、`5b9fbe5b` | BUG-1057:read-case 候选四字段、去重复列表、超限先砍候选明细;(后者只修测试里的类型转换) |
|
||
| T5 | `1ed3e555` | R1:offer 走 compare 同一套剥离;报告 Markdown 只留一份;引擎宽度不给模型 |
|
||
| T6 | `c4bb394b` | R4 + Skill 切片:证据轮只附 §5/§7;关 Mastra `<available_skills>` 注入 |
|
||
| — | `717a378b` | T3/T5 跟进:同样行为,改写法避免新增 3 条 unused-var 警告 |
|
||
| T7 | 本提交 | BUG 历史、CHANGELOG、本记录、真机清单、任务索引 |
|
||
|
||
T4 先于 T2 提交(T1 的 attempt 改动已带 `rangeBeforeTurn` 选项,T4 只接路由)。
|
||
|
||
## 各项怎么修的
|
||
|
||
| 项 | 做法 |
|
||
| --- | --- |
|
||
| D1 / BUG-1055 | attempt 结算时不再把「范围从…变为…」「这次没有重新比较」「候选比较这次没跑成」接进正文、也不流式放出;只把结构化事实(`TurnServerFacts`)和当轮数字白名单(`SpokenFactWhitelist`:结算后可信范围与区间、代表分钟、决策范围、吻合率)交给 finish。finish 顺序:① 证据轮对模型正文跑白名单(任何一句的时刻 / 时刻区间 / 百分比不在白名单就整句丢;一句不剩用 batch 回执复述)→ ② BUG-606 / BUG-615 裁句(只作用于模型正文)→ ③ `composeSpokenWithServerFacts` 接事实句(沿用原 `withCompareFailedRetryNotice` / `withRescoreSkippedNotice` / `withRangeChangedAfterEvidence` 的格式)→ ④ 交付时的缺口句 → 与已流式文本不同则一次 `replace`。batch 返回 `range_after_rescore`(`credible_range`、`representative_time`、`event_fit_percent`、`delivers_range_this_turn` = 无下一问且 `deliveryNarrationAllowed`),回执指纹仍按旧形状算。系统提示:删「有收窄就说进度」,改为范围由系统说、除出牌轮外不写时刻和百分比;出牌三句只在 `delivers_range_this_turn=true` 时写、数字只抄它(R3)。 |
|
||
| D2 / BUG-1056 | 路由只服务校正 → 整条删除(Next 返回 404)。删 `regenerate-turn.ts`、`getRectificationV9RegenerationAgent`、`createRectificationV9ReadOnlyTools`、交付守卫的 `"regenerate"` 触发源、`insertV9SkillRunReceipt` 的 `"regeneration"` 分支类型;客户端删 `regenerateRectificationMessage`、`regeneratingMessageKey`、`regenerating` / `canRegenerate` / `latestRegeneratableKey`、问题空档输入的 `regenerating`。`ChatMessageActions` 的 `onRegenerate` / `canRegenerate` 改可选,不传不画按钮;普通对话照旧传。 |
|
||
| D3 / BUG-1057 | `modelVisibleInference`:候选只留 `time` / `score`(= 推断概率,保留三位小数)/ `status` / `cluster_range`;上一轮 `eliminated_ids` / `score_deltas` 改按时刻(`eliminated_times`)。删 `candidate_summary.candidates`。`enforceTurnDecisionBudget` 新顺序:去 cluster_range → 最好 6 个候选(未淘汰优先,`candidates_omitted`)→ 对话 4 条 / 证据 3 条 → 清空(`truncated`)。compare / offer 的 `inference_state` 用同一个瘦身视图。 |
|
||
| D4 / BUG-1058 | 二选一选「把答题前范围交给 Agent 轮」:只出一条服务端范围句、走 D1 的事实句通道(不裁、不撤回),答题与经历两次变化合成一句;若把 `applied.narration` 并进正文,同一条消息会有两句范围。实现:typed-message 在点选焦点 `deferFollowup` 时写 `turnState.rangeBeforeTurn = previous?.credible_range`,Agent 轮用它当开轮范围。收集焦点拒答的 `deferFollowup` 不改推断,不传。 |
|
||
| R1 | `offerModelProjection` = `agentVisibleLatestProjection(payload)`;后者再去掉 `engine_indistinguishable_width_minutes` 与 `range_delivery.verification_markdown`(与 `skill_verification_report.markdown` 逐字相同)。病例 API 的 `latestResultToolProjection` 与回执不变。 |
|
||
| R3 | 见 D1:写正文时是否交付由 batch 的 `delivers_range_this_turn` 告诉模型,吻合率也在里面;数字仍过白名单。 |
|
||
| R4 | 校正 Agent 加 `inputProcessors: [rectificationSkillBoundProcessor]`(`providesSkillDiscovery: "on-demand"`,与咨询 Agent 的做法相同):Mastra 不再加 `<available_skills>`(含临时目录路径)与「请调用 skill 工具」两条系统消息,`agent.getSkill` 仍按绑定版本加载。 |
|
||
| Skill 切片 | `skill-slice.ts`,规则常量 `RECTIFICATION_SKILL_SECTIONS_BY_ACTION`:证据轮按标题取「ConversationFocus 与意图承接」「批量证据与日期真实性」(10.0.x 的 §5、§7),前面加一句「本轮是证据轮…其余章节不适用」;开场、只读等其余动作发整份。按标题不按编号:9.0.0 的编号不同(它的 §5 是候选语言),找不齐标题就整份发,历史 Case 仍按绑定版本原文运行(BUG-621)。§6(全是 `full_diagnostics` 才有的字段:`method_followup_plan`、`guided_collect_windows`…)在证据轮不再发送。 |
|
||
|
||
## 红线测试清单
|
||
|
||
| 红线 | 测试 |
|
||
| --- | --- |
|
||
| 1 事实句不裁、流式终态 = 落库、不先出后撤 | `rectification-grounding-20260927.test.ts`:repro 1(两句)、repro 1b(三句)、repro 2、「这次没有重新比较」、「候选比较这次没跑成」——都断言 `states.at(-1) === persisted` 且事实句一旦出现不再消失(`assertNeverWithdrawn`);`rectification-grounding-route-20260927.test.ts` 两条(路由级 `shown === persisted`) |
|
||
| 2 时刻 / 区间 / 百分比 = 当轮事实 | 同文件:首句旧范围整句不落库、编造吻合率 90% 被丢 / 服务端 80% 保留、重算后范围与代表分钟可复述、白名单单元(整句门、不在句中删词) |
|
||
| 3 9 个候选仍保留对话与证据 | `rectification-grounding-read-case-20260927.test.ts`:真实引擎 9 候选 + 6 条对话;超限先砍候选明细;第一步只去 cluster_range |
|
||
| 5 历史会话(BUG-621) | `rectification-grounding-prompt-20260927.test.ts`:9.0.0 快照仍可 `resolveExactSkillPackage`,切片对它整份返回;Skill 未 bump,既有 `skill-registry` / `rectification-history-open-20260909` 测试照旧通过 |
|
||
| D2 | `rectification-grounding-regenerate-20260927.test.ts`(校正无重新生成、普通对话有、路由与代码已删、DB 函数仍在迁移里) |
|
||
| R1 | `rectification-grounding-projection-20260927.test.ts`(offer / compare 模型视图,真实引擎响应) |
|
||
| R4 / 切片 / 体量 | `rectification-grounding-prompt-20260927.test.ts`(真实 `runV9AgentTurn` + 真实 Agent + 记录提示词的假模型:无 `<available_skills>`、只含 §5/§7、每次调用固定开销 < 8K) |
|
||
|
||
未改动基线上:`rectification-grounding-20260927` 的 8 条运行级用例全部失败;路由级 2 条在去掉修复后失败。
|
||
|
||
fixture:`frontend/tests/fixtures/rectification-grounding-aa-30min.golden.json`,诊断时本机真实引擎对公开 AA 盘(30 分钟窗、3 件公开事件)的 `runV9CandidateScore` 请求与响应(388KB,AGENTS §7.4)。
|
||
|
||
## 体量前后对比
|
||
|
||
测量方法:同一个脚本分别在 `cd4dde9d`(改前)与本分支(改后)上跑:真实 `runV9AgentTurn`(证据轮)+ 真实 `getRectificationV9Agent`(真实绑定 Skill 10.0.31、真实工具)+ 记录每次调用提示词的假模型;数据是上面的公开 AA 盘 fixture(9 个候选)。token 为估算:每个 CJK 字符(含全角标点)记 1,其余字符每 4 个记 1(`estimateTokens`,不是供应商分词器)。「固定开销」= 每次调用都会重发、与工具结果无关的部分:全部 system 消息 + 工具 schema。
|
||
|
||
| 每次调用 | 改前 | 改后 |
|
||
| --- | --- | --- |
|
||
| system 消息字符数 | 13,083 | 3,925 |
|
||
| system 消息 token | 8,211 | 2,680 |
|
||
| 工具 schema(第 1 步只有 read-case) | 797 B / 229 | 797 B / 229 |
|
||
| 工具 schema(第 2 步起 14 个工具) | 13,796 B / 3,791 | 13,796 B / 3,791 |
|
||
| **固定开销,第 1 步** | **8,440** | **2,909** |
|
||
| **固定开销,第 2 步起** | **12,002** | **6,471**(目标 < 8K) |
|
||
| 第 1 次调用合计输入 | 8,556 | 3,025 |
|
||
| 第 2 次调用合计输入(含 read-case 结果) | 13,605 | 7,417 |
|
||
|
||
system 消息改前 = 系统提示 1.5K 字 + Skill 全文 10.7K 字 + Mastra 技能清单与「调用 skill 工具」约 0.95K 字;改后 = 系统提示 + Skill §5/§7。工具 schema 没动(按轮次收窄 schema 是可选项,本单未做,见让步)。
|
||
|
||
典型「read-case → compare → 写正文」三次调用合计(前两次实测,第三次 = 第二次 + compare 返回):改前约 8,556 + 13,605 + (13,605 + 10,488) ≈ 46.3K;改后约 3,025 + 7,417 + (7,417 + 8,129) ≈ 26.0K。
|
||
|
||
| 工具返回(给模型看的) | 改前 | 改后 |
|
||
| --- | --- | --- |
|
||
| read-case,9 候选 + 2 条对话 | 5,639 B,`truncated`,对话 0 / 证据 0,候选明细 3,778 B | 3,022 B,对话 2 / 证据 3,候选 754 B |
|
||
| read-case,9 候选 + 6 条对话 | 5,639 B,`truncated`,对话 0 / 证据 0 | 3,420 B,对话 6 / 证据 3 |
|
||
| compare | 38,026 B(约 10.5K token) | 29,698 B(约 8.1K token) |
|
||
| offer | 100,012 B | 28,799 B |
|
||
|
||
read-case 内容对比(9 候选):改前每个候选带 `id`、`probability`、`posterior_score`、`rank`、窗口定位五项(`candidate_date` / `window_index` / `window_offset_minutes` / `segment_index` / `cluster_intervals`),另有 `candidate_summary.candidates` 重复一份前 6 个;超限后对话与证据被清空。改后每个候选只有 `time` / `score` / `status` / `cluster_range`,无重复列表,对话与证据都在。
|
||
|
||
## 测试与构建(最终在 rebase 后的分支上跑)
|
||
|
||
| 项 | 基线 `cd4dde9d` | 本分支 | 结论 |
|
||
| --- | --- | --- | --- |
|
||
| `tsc --noEmit` | — | 0 错 | 通过 |
|
||
| `npm run lint` | 0 error / 127 warning | 0 error / 127 warning(逐条同名单) | 通过;中途多出的 3 条 unused-var 由 `717a378b` 改写法消掉 |
|
||
| 全量前端测试(Node 22.14,TAP 写文件避免 stdout 吞行) | 4136 tests,fail 24 | 4152 tests,fail 24 | 失败名单逐条一致(24 条均为无 Docker 的 DB / 部署套件);新增 27 个测试名 |
|
||
| 消失的测试名 | — | 11 个 | 8 个是删掉的 `rectification-v9-regenerate.test.ts`;3 个改名:`completed Agent replies restore feedback, copy and safe in-place regeneration actions`、`question gap: nothing is shown while busy, readonly, regenerating, or for non-resumable cases`、`reply regeneration is a separate Jyotisha agent with only read-case access`(三栏见下表) |
|
||
| 全量 Python 门禁集(`gate-pytest-args.txt`,含 `test_api_server_growth_contract.py`) | 948 passed / 1 skipped | 948 passed / 1 skipped | 通过;未改 Python、`scripts/jyotish_api_server.py` 未动,伪造点与方法数不变 |
|
||
| `npm run build -- --webpack` | `/`、`/chart`、`/ephemeris`、`/people` ○ Static | 同左 ○ Static | 通过 |
|
||
| rootMainFiles gzip(level 9,4 个文件) | 131,145 B | 130,933 B | −0.16%,在 ±2% 内 |
|
||
| 首页 index.html 引用的 36 个 js/css gzip | 661,107 B | 660,413 B | −0.10% |
|
||
| `tests/*.py` 读前端文本 | — | 已 grep 本单改动的文件名与文案(重新生成、`有收窄就说进度`、`candidate_summary`、消息组件等),无 Python 合同命中 | — |
|
||
|
||
失败名单对比方法:`tsx --test --test-reporter=tap --test-reporter-destination=<file> tests/*.test.ts tests/*.test.tsx`,取 `^(not )?ok N - ` 行排序后 `comm`;基线在 `cd4dde9d` 的一次性工作树上跑。构建后已删 `frontend/frontend/`。
|
||
|
||
## 改动的既有断言(原值 / 新值 / 原因)
|
||
|
||
每处在测试文件里都写了三栏注释,这里汇总:
|
||
|
||
| 文件 · 用例 | 原值 | 新值 | 原因 |
|
||
| --- | --- | --- | --- |
|
||
| `rectification-v9-agent` · set-focus after spoken text… | 第二句「范围收到 05:00 到 05:10。」 | 「这件事拿去和星盘对照了。」 | P3:夹具无候选快照,手写时刻整句不落库;本用例锁 set-focus 之后不接第二段 |
|
||
| `rectification-v9-agent` · agent-run strips question sentences… | 首句「范围已经收到,收在 05:00–05:10。」 | 「2016 年 9 月去北京工作,这件事拿去和星盘对照。」 | 同上;剪问句断言不变 |
|
||
| `rectification-v9-agent` · delivery persist keeps three sentences… | 首句「目前范围 04:49–04:53。」,回执无吻合率 | 「代表分钟是 05:02。」,回执加 `event_fit_rate` 80% | P3 只放行当轮事实;三句 / 四句裁三句的断言不变 |
|
||
| `rectification-v9-agent` · reply regeneration is a separate… | 断言重新生成 Agent 与只读工具 | 改名「rectification has no reply-regeneration agent…」,断言两者不存在 | BUG-1056 删除 |
|
||
| `rectification-surface-state` · question gap: nothing is shown while busy, readonly, regenerating… | 含 `regenerating: true` → idle | 去掉 regenerating 与该断言(改名) | 输入已无此项 |
|
||
| `rectification-agentic-entry` · completed Agent replies restore feedback, copy and safe in-place regeneration… | 容器有重新生成请求、路由不计费且按绑定 Skill | 改名「…feedback and copy, no regeneration」;断言校正不传 `onRegenerate`、无请求、路由文件不存在 | BUG-1056 |
|
||
| `api-service-unavailable-20260904` | 路由清单含 regenerate | 移出 | 文件已删 |
|
||
| `chat-composer-queue` | `/if \(busy \|\| regeneratingMessageKey\) \{/` | `/if \(busy\) \{/` | 排队行为不变 |
|
||
| `rectification-answer-choice` · attempt timeout stays under… | 另断言 regenerate 路由 maxDuration | 只断言 agent 路由 | 文件已删 |
|
||
| `rectification-delivery-ui-simplify` · already_delivered… | 断言 regenerate 路由 `trigger: "regenerate"` | 删该断言 | 触发源随之删除 |
|
||
| `rectification-request-dossier-cache` | 路由清单含 regenerate | 移出 | 文件已删 |
|
||
| `rectification-v9-stream` · whole-run budget is less than both route maxDuration values | 另断言 regenerate 路由 | 只断言 agent 路由(名称保留) | 文件已删 |
|
||
| `rectification-settled-render-split` | 传 `regenerating` / `canRegenerate` / `latestRegeneratableKey` / `regeneratingMessageKey` | 去掉 | 属性已删;渲染计数断言不变 |
|
||
| 11 个文件里的问题空档输入 `regenerating: false,`(delivery-vs-collect、open-collect-invite、post-adopt-verify、probe-pool-exhausted、replay-20260911、spoken-orphan、surface-state、targeted-card-live、targeted-spoken-focus-recovery、tiebreak-before-card、tied-first-fix)与 `rectification-chat-run-fixtures` 的 `regeneratingMessageKey: null` | 有这一项 | 删除这一行 | 输入类型已无此项(tsc 会拒绝多余属性);各断言不变 |
|
||
| `rectification-v9-regenerate.test.ts` | 8 条重新生成用例 | 整份删除 | 测的是已删除的接口;替代为 `rectification-grounding-regenerate-20260927.test.ts` |
|
||
|
||
BUG-588 原测试(`rectification-probe-replay-loss-20260908`)、BUG-593 原测试(`rectification-delivery-report-facts`)、BUG-615 源码合同(`rectification-delivery-ui-simplify` 的裁句顺序)未改、仍通过。
|
||
|
||
## 让步与未做
|
||
|
||
- T6 可选项「按轮次收窄每步工具 schema」未做(第 2 步起 14 个工具 3,791 token 仍是固定开销的大头;固定开销已低于 8K,留下一单)。
|
||
- compare 返回仍约 29.7KB(`range_delivery.columns`、`window_scan`、`candidates`、审计表等),本单只按任务书剥离了推断前数据、对照包、探针、重复 Markdown 与引擎宽度。
|
||
- `delivers_range_this_turn` 是 batch 时刻按服务端交付判定(无下一问 + `deliveryNarrationAllowed`)给出的;finish 里 `persistNextInterviewIfIdle` 若再触发陈旧快照重算,最终是否交付以它为准(模型正文仍过白名单与三句上限)。
|
||
- 证据轮切片不含 §9(候选语言);交付轮的边界句由系统提示第 5 条与白名单保证。若验收认为交付轮需要 §9,可在常量里加标题。
|
||
- 数据库函数 `regenerate_agentic_rectification_turn` 待另开单退役。
|
||
|
||
## 环境缺口
|
||
|
||
- 无 Docker:24 条 DB / 部署套件失败,与基线逐条一致(见上)。
|
||
- 无真实模型凭据:真实模型下的旁白质量、`delivers_range_this_turn` 时的三句写法,需部署后按 `docs/testing/rectification-grounding-20260927.md` 走真机。
|
||
|
||
## 验收(Claude,2026-09-27)
|
||
|
||
基于 `19376089` 独立复跑(Node 22.14):tsc 0;lint 0 error;`npm test` 4152 / 24 fail / 28 skip,失败名单与 `cd4dde9d` 基线逐条一致;消失的 11 个测试名 = 随重试功能删除的 `rectification-v9-regenerate.test.ts` 8 条 + 3 条改名(有三栏说明);Python 门禁集退出 0;`/`、`/chart`、`/ephemeris`、`/people` ○ Static;rootMainFiles gzip(level 9)130933 B。未碰 `.gitea/`、`vendor/`、`deploy/`、迁移、任何 Python;删除的 3 个文件均为校正重试功能本身,源码里已无校正重试引用,普通对话重试按钮由 `onRegenerate` 控制仍在。数据库函数 `regenerate_agentic_rectification_turn` 保留,待另开单退役。
|
||
|
||
观察(不阻塞):数字白名单会把模型复述用户经历里自带的钟点(如「10:30 手术」)那一整句丢掉、由服务端 recap 顶上,措辞会略生硬;真机若常见再调。
|