Files
Jyotisha/docs/tasks/PROGRESS-rectification-grounding-20260927.md
T
Jesse_ChenandClaude Opus 5.5 261d7b2e04
Independent Staging Quality Gate / validate (push) Successful in 11m43s
Independent Staging Quality Gate / publish (push) Successful in 3m27s
docs(tasks): rectification grounding acceptance
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 03:36:29 +08:00

141 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PROGRESS · 生时校正:旁白事实有据 + 去掉重试按钮 + 给模型的资料瘦身 · 2026-09-27
- 执行方式:直接执行(产品负责人授权子代理执行;Claude 事后独立验收)
- 任务书:`docs/tasks/TASK-rectification-grounding-20260927.md`(`a98d3b20`)
- 基线:开工 `origin/staging` `a98d3b20`(诊断基于 `599fe7a9`);中途 staging 合入数据卡单(`cd4dde9d`),本分支已 rebase 到 `cd4dde9d`,最终对比基线全部在 `cd4dde9d` 上重跑;随后再 rebase 到 `19376089`(只改 `docs/BUG_HISTORY.md` 的纯文档提交,代码树与 `cd4dde9d` 相同)。
- 分支:`codex/rectification-grounding-20260927`,工作树 `.worktrees/rectification-grounding-20260927`(本地提交,未推送)
- BUG 编号:预留 BUG-1055~1058,开工与提交时核对(开工最大号 1054;rebase 后 staging 有 1059,本单 1055~1058 未被占用)
- Skill:不 bump。`jyotish-birth-time-rectification` 10.0.31 文本未改,只是证据轮按章节切片发送;registry 不动。
- 数据库:不动。`regenerate_agentic_rectification_turn` 保留,**待另开单退役**(AGENTS §7.6 两轮规则:本轮先下线代码)。
## 提交
| 任务 | 提交 | 内容 |
| --- | --- | --- |
| T1 | `0e0baaa7` | BUG-1055:事实句不参与裁句、不先流后撤;P3 数字白名单(`spoken-grounding.ts`);batch 返回 `range_after_rescore`;系统提示「范围由系统说」 |
| T4 | `5a808449` | BUG-1058:`deferFollowup` 路径把答题前范围交给 Agent 轮(`rangeBeforeTurn`);路由级回归 |
| T2 | `84eb2315` | BUG-1056:删除校正重新生成(路由、服务端模块、Agent、只读工具集、客户端动作与状态);`ChatMessageActions` 无 `onRegenerate` 不画按钮;DESIGN / VOICE |
| T3 | `13d93020`、`5b9fbe5b` | BUG-1057:read-case 候选四字段、去重复列表、超限先砍候选明细;(后者只修测试里的类型转换) |
| T5 | `1ed3e555` | R1:offer 走 compare 同一套剥离;报告 Markdown 只留一份;引擎宽度不给模型 |
| T6 | `c4bb394b` | R4 + Skill 切片:证据轮只附 §5/§7;关 Mastra `<available_skills>` 注入 |
| — | `717a378b` | T3/T5 跟进:同样行为,改写法避免新增 3 条 unused-var 警告 |
| T7 | 本提交 | BUG 历史、CHANGELOG、本记录、真机清单、任务索引 |
T4 先于 T2 提交(T1 的 attempt 改动已带 `rangeBeforeTurn` 选项,T4 只接路由)。
## 各项怎么修的
| 项 | 做法 |
| --- | --- |
| D1 / BUG-1055 | attempt 结算时不再把「范围从…变为…」「这次没有重新比较」「候选比较这次没跑成」接进正文、也不流式放出;只把结构化事实(`TurnServerFacts`)和当轮数字白名单(`SpokenFactWhitelist`:结算后可信范围与区间、代表分钟、决策范围、吻合率)交给 finish。finish 顺序:① 证据轮对模型正文跑白名单(任何一句的时刻 / 时刻区间 / 百分比不在白名单就整句丢;一句不剩用 batch 回执复述)→ ② BUG-606 / BUG-615 裁句(只作用于模型正文)→ ③ `composeSpokenWithServerFacts` 接事实句(沿用原 `withCompareFailedRetryNotice` / `withRescoreSkippedNotice` / `withRangeChangedAfterEvidence` 的格式)→ ④ 交付时的缺口句 → 与已流式文本不同则一次 `replace`。batch 返回 `range_after_rescore`(`credible_range`、`representative_time`、`event_fit_percent`、`delivers_range_this_turn` = 无下一问且 `deliveryNarrationAllowed`),回执指纹仍按旧形状算。系统提示:删「有收窄就说进度」,改为范围由系统说、除出牌轮外不写时刻和百分比;出牌三句只在 `delivers_range_this_turn=true` 时写、数字只抄它(R3)。 |
| D2 / BUG-1056 | 路由只服务校正 → 整条删除(Next 返回 404)。删 `regenerate-turn.ts`、`getRectificationV9RegenerationAgent`、`createRectificationV9ReadOnlyTools`、交付守卫的 `"regenerate"` 触发源、`insertV9SkillRunReceipt` 的 `"regeneration"` 分支类型;客户端删 `regenerateRectificationMessage`、`regeneratingMessageKey`、`regenerating` / `canRegenerate` / `latestRegeneratableKey`、问题空档输入的 `regenerating`。`ChatMessageActions` 的 `onRegenerate` / `canRegenerate` 改可选,不传不画按钮;普通对话照旧传。 |
| D3 / BUG-1057 | `modelVisibleInference`:候选只留 `time` / `score`(= 推断概率,保留三位小数)/ `status` / `cluster_range`;上一轮 `eliminated_ids` / `score_deltas` 改按时刻(`eliminated_times`)。删 `candidate_summary.candidates`。`enforceTurnDecisionBudget` 新顺序:去 cluster_range → 最好 6 个候选(未淘汰优先,`candidates_omitted`)→ 对话 4 条 / 证据 3 条 → 清空(`truncated`)。compare / offer 的 `inference_state` 用同一个瘦身视图。 |
| D4 / BUG-1058 | 二选一选「把答题前范围交给 Agent 轮」:只出一条服务端范围句、走 D1 的事实句通道(不裁、不撤回),答题与经历两次变化合成一句;若把 `applied.narration` 并进正文,同一条消息会有两句范围。实现:typed-message 在点选焦点 `deferFollowup` 时写 `turnState.rangeBeforeTurn = previous?.credible_range`,Agent 轮用它当开轮范围。收集焦点拒答的 `deferFollowup` 不改推断,不传。 |
| R1 | `offerModelProjection` = `agentVisibleLatestProjection(payload)`;后者再去掉 `engine_indistinguishable_width_minutes` 与 `range_delivery.verification_markdown`(与 `skill_verification_report.markdown` 逐字相同)。病例 API 的 `latestResultToolProjection` 与回执不变。 |
| R3 | 见 D1:写正文时是否交付由 batch 的 `delivers_range_this_turn` 告诉模型,吻合率也在里面;数字仍过白名单。 |
| R4 | 校正 Agent 加 `inputProcessors: [rectificationSkillBoundProcessor]`(`providesSkillDiscovery: "on-demand"`,与咨询 Agent 的做法相同):Mastra 不再加 `<available_skills>`(含临时目录路径)与「请调用 skill 工具」两条系统消息,`agent.getSkill` 仍按绑定版本加载。 |
| Skill 切片 | `skill-slice.ts`,规则常量 `RECTIFICATION_SKILL_SECTIONS_BY_ACTION`:证据轮按标题取「ConversationFocus 与意图承接」「批量证据与日期真实性」(10.0.x 的 §5、§7),前面加一句「本轮是证据轮…其余章节不适用」;开场、只读等其余动作发整份。按标题不按编号:9.0.0 的编号不同(它的 §5 是候选语言),找不齐标题就整份发,历史 Case 仍按绑定版本原文运行(BUG-621)。§6(全是 `full_diagnostics` 才有的字段:`method_followup_plan`、`guided_collect_windows`…)在证据轮不再发送。 |
## 红线测试清单
| 红线 | 测试 |
| --- | --- |
| 1 事实句不裁、流式终态 = 落库、不先出后撤 | `rectification-grounding-20260927.test.ts`:repro 1(两句)、repro 1b(三句)、repro 2、「这次没有重新比较」、「候选比较这次没跑成」——都断言 `states.at(-1) === persisted` 且事实句一旦出现不再消失(`assertNeverWithdrawn`);`rectification-grounding-route-20260927.test.ts` 两条(路由级 `shown === persisted`) |
| 2 时刻 / 区间 / 百分比 = 当轮事实 | 同文件:首句旧范围整句不落库、编造吻合率 90% 被丢 / 服务端 80% 保留、重算后范围与代表分钟可复述、白名单单元(整句门、不在句中删词) |
| 3 9 个候选仍保留对话与证据 | `rectification-grounding-read-case-20260927.test.ts`:真实引擎 9 候选 + 6 条对话;超限先砍候选明细;第一步只去 cluster_range |
| 5 历史会话(BUG-621) | `rectification-grounding-prompt-20260927.test.ts`:9.0.0 快照仍可 `resolveExactSkillPackage`,切片对它整份返回;Skill 未 bump,既有 `skill-registry` / `rectification-history-open-20260909` 测试照旧通过 |
| D2 | `rectification-grounding-regenerate-20260927.test.ts`(校正无重新生成、普通对话有、路由与代码已删、DB 函数仍在迁移里) |
| R1 | `rectification-grounding-projection-20260927.test.ts`(offer / compare 模型视图,真实引擎响应) |
| R4 / 切片 / 体量 | `rectification-grounding-prompt-20260927.test.ts`(真实 `runV9AgentTurn` + 真实 Agent + 记录提示词的假模型:无 `<available_skills>`、只含 §5/§7、每次调用固定开销 < 8K) |
未改动基线上:`rectification-grounding-20260927` 的 8 条运行级用例全部失败;路由级 2 条在去掉修复后失败。
fixture:`frontend/tests/fixtures/rectification-grounding-aa-30min.golden.json`,诊断时本机真实引擎对公开 AA 盘(30 分钟窗、3 件公开事件)的 `runV9CandidateScore` 请求与响应(388KB,AGENTS §7.4)。
## 体量前后对比
测量方法:同一个脚本分别在 `cd4dde9d`(改前)与本分支(改后)上跑:真实 `runV9AgentTurn`(证据轮)+ 真实 `getRectificationV9Agent`(真实绑定 Skill 10.0.31、真实工具)+ 记录每次调用提示词的假模型;数据是上面的公开 AA 盘 fixture(9 个候选)。token 为估算:每个 CJK 字符(含全角标点)记 1,其余字符每 4 个记 1(`estimateTokens`,不是供应商分词器)。「固定开销」= 每次调用都会重发、与工具结果无关的部分:全部 system 消息 + 工具 schema。
| 每次调用 | 改前 | 改后 |
| --- | --- | --- |
| system 消息字符数 | 13,083 | 3,925 |
| system 消息 token | 8,211 | 2,680 |
| 工具 schema(第 1 步只有 read-case) | 797 B / 229 | 797 B / 229 |
| 工具 schema(第 2 步起 14 个工具) | 13,796 B / 3,791 | 13,796 B / 3,791 |
| **固定开销,第 1 步** | **8,440** | **2,909** |
| **固定开销,第 2 步起** | **12,002** | **6,471**(目标 < 8K) |
| 第 1 次调用合计输入 | 8,556 | 3,025 |
| 第 2 次调用合计输入(含 read-case 结果) | 13,605 | 7,417 |
system 消息改前 = 系统提示 1.5K 字 + Skill 全文 10.7K 字 + Mastra 技能清单与「调用 skill 工具」约 0.95K 字;改后 = 系统提示 + Skill §5/§7。工具 schema 没动(按轮次收窄 schema 是可选项,本单未做,见让步)。
典型「read-case → compare → 写正文」三次调用合计(前两次实测,第三次 = 第二次 + compare 返回):改前约 8,556 + 13,605 + (13,605 + 10,488) ≈ 46.3K;改后约 3,025 + 7,417 + (7,417 + 8,129) ≈ 26.0K。
| 工具返回(给模型看的) | 改前 | 改后 |
| --- | --- | --- |
| read-case,9 候选 + 2 条对话 | 5,639 B,`truncated`,对话 0 / 证据 0,候选明细 3,778 B | 3,022 B,对话 2 / 证据 3,候选 754 B |
| read-case,9 候选 + 6 条对话 | 5,639 B,`truncated`,对话 0 / 证据 0 | 3,420 B,对话 6 / 证据 3 |
| compare | 38,026 B(约 10.5K token) | 29,698 B(约 8.1K token) |
| offer | 100,012 B | 28,799 B |
read-case 内容对比(9 候选):改前每个候选带 `id`、`probability`、`posterior_score`、`rank`、窗口定位五项(`candidate_date` / `window_index` / `window_offset_minutes` / `segment_index` / `cluster_intervals`),另有 `candidate_summary.candidates` 重复一份前 6 个;超限后对话与证据被清空。改后每个候选只有 `time` / `score` / `status` / `cluster_range`,无重复列表,对话与证据都在。
## 测试与构建(最终在 rebase 后的分支上跑)
| 项 | 基线 `cd4dde9d` | 本分支 | 结论 |
| --- | --- | --- | --- |
| `tsc --noEmit` | — | 0 错 | 通过 |
| `npm run lint` | 0 error / 127 warning | 0 error / 127 warning(逐条同名单) | 通过;中途多出的 3 条 unused-var 由 `717a378b` 改写法消掉 |
| 全量前端测试(Node 22.14,TAP 写文件避免 stdout 吞行) | 4136 tests,fail 24 | 4152 tests,fail 24 | 失败名单逐条一致(24 条均为无 Docker 的 DB / 部署套件);新增 27 个测试名 |
| 消失的测试名 | — | 11 个 | 8 个是删掉的 `rectification-v9-regenerate.test.ts`;3 个改名:`completed Agent replies restore feedback, copy and safe in-place regeneration actions`、`question gap: nothing is shown while busy, readonly, regenerating, or for non-resumable cases`、`reply regeneration is a separate Jyotisha agent with only read-case access`(三栏见下表) |
| 全量 Python 门禁集(`gate-pytest-args.txt`,含 `test_api_server_growth_contract.py`) | 948 passed / 1 skipped | 948 passed / 1 skipped | 通过;未改 Python、`scripts/jyotish_api_server.py` 未动,伪造点与方法数不变 |
| `npm run build -- --webpack` | `/`、`/chart`、`/ephemeris`、`/people` ○ Static | 同左 ○ Static | 通过 |
| rootMainFiles gzip(level 9,4 个文件) | 131,145 B | 130,933 B | −0.16%,在 ±2% 内 |
| 首页 index.html 引用的 36 个 js/css gzip | 661,107 B | 660,413 B | −0.10% |
| `tests/*.py` 读前端文本 | — | 已 grep 本单改动的文件名与文案(重新生成、`有收窄就说进度`、`candidate_summary`、消息组件等),无 Python 合同命中 | — |
失败名单对比方法:`tsx --test --test-reporter=tap --test-reporter-destination=<file> tests/*.test.ts tests/*.test.tsx`,取 `^(not )?ok N - ` 行排序后 `comm`;基线在 `cd4dde9d` 的一次性工作树上跑。构建后已删 `frontend/frontend/`。
## 改动的既有断言(原值 / 新值 / 原因)
每处在测试文件里都写了三栏注释,这里汇总:
| 文件 · 用例 | 原值 | 新值 | 原因 |
| --- | --- | --- | --- |
| `rectification-v9-agent` · set-focus after spoken text… | 第二句「范围收到 05:00 到 05:10。」 | 「这件事拿去和星盘对照了。」 | P3:夹具无候选快照,手写时刻整句不落库;本用例锁 set-focus 之后不接第二段 |
| `rectification-v9-agent` · agent-run strips question sentences… | 首句「范围已经收到,收在 05:00–05:10。」 | 「2016 年 9 月去北京工作,这件事拿去和星盘对照。」 | 同上;剪问句断言不变 |
| `rectification-v9-agent` · delivery persist keeps three sentences… | 首句「目前范围 04:49–04:53。」,回执无吻合率 | 「代表分钟是 05:02。」,回执加 `event_fit_rate` 80% | P3 只放行当轮事实;三句 / 四句裁三句的断言不变 |
| `rectification-v9-agent` · reply regeneration is a separate… | 断言重新生成 Agent 与只读工具 | 改名「rectification has no reply-regeneration agent…」,断言两者不存在 | BUG-1056 删除 |
| `rectification-surface-state` · question gap: nothing is shown while busy, readonly, regenerating… | 含 `regenerating: true` → idle | 去掉 regenerating 与该断言(改名) | 输入已无此项 |
| `rectification-agentic-entry` · completed Agent replies restore feedback, copy and safe in-place regeneration… | 容器有重新生成请求、路由不计费且按绑定 Skill | 改名「…feedback and copy, no regeneration」;断言校正不传 `onRegenerate`、无请求、路由文件不存在 | BUG-1056 |
| `api-service-unavailable-20260904` | 路由清单含 regenerate | 移出 | 文件已删 |
| `chat-composer-queue` | `/if \(busy \|\| regeneratingMessageKey\) \{/` | `/if \(busy\) \{/` | 排队行为不变 |
| `rectification-answer-choice` · attempt timeout stays under… | 另断言 regenerate 路由 maxDuration | 只断言 agent 路由 | 文件已删 |
| `rectification-delivery-ui-simplify` · already_delivered… | 断言 regenerate 路由 `trigger: "regenerate"` | 删该断言 | 触发源随之删除 |
| `rectification-request-dossier-cache` | 路由清单含 regenerate | 移出 | 文件已删 |
| `rectification-v9-stream` · whole-run budget is less than both route maxDuration values | 另断言 regenerate 路由 | 只断言 agent 路由(名称保留) | 文件已删 |
| `rectification-settled-render-split` | 传 `regenerating` / `canRegenerate` / `latestRegeneratableKey` / `regeneratingMessageKey` | 去掉 | 属性已删;渲染计数断言不变 |
| 11 个文件里的问题空档输入 `regenerating: false,`(delivery-vs-collect、open-collect-invite、post-adopt-verify、probe-pool-exhausted、replay-20260911、spoken-orphan、surface-state、targeted-card-live、targeted-spoken-focus-recovery、tiebreak-before-card、tied-first-fix)与 `rectification-chat-run-fixtures` 的 `regeneratingMessageKey: null` | 有这一项 | 删除这一行 | 输入类型已无此项(tsc 会拒绝多余属性);各断言不变 |
| `rectification-v9-regenerate.test.ts` | 8 条重新生成用例 | 整份删除 | 测的是已删除的接口;替代为 `rectification-grounding-regenerate-20260927.test.ts` |
BUG-588 原测试(`rectification-probe-replay-loss-20260908`)、BUG-593 原测试(`rectification-delivery-report-facts`)、BUG-615 源码合同(`rectification-delivery-ui-simplify` 的裁句顺序)未改、仍通过。
## 让步与未做
- T6 可选项「按轮次收窄每步工具 schema」未做(第 2 步起 14 个工具 3,791 token 仍是固定开销的大头;固定开销已低于 8K,留下一单)。
- compare 返回仍约 29.7KB(`range_delivery.columns`、`window_scan`、`candidates`、审计表等),本单只按任务书剥离了推断前数据、对照包、探针、重复 Markdown 与引擎宽度。
- `delivers_range_this_turn` 是 batch 时刻按服务端交付判定(无下一问 + `deliveryNarrationAllowed`)给出的;finish 里 `persistNextInterviewIfIdle` 若再触发陈旧快照重算,最终是否交付以它为准(模型正文仍过白名单与三句上限)。
- 证据轮切片不含 §9(候选语言);交付轮的边界句由系统提示第 5 条与白名单保证。若验收认为交付轮需要 §9,可在常量里加标题。
- 数据库函数 `regenerate_agentic_rectification_turn` 待另开单退役。
## 环境缺口
- 无 Docker:24 条 DB / 部署套件失败,与基线逐条一致(见上)。
- 无真实模型凭据:真实模型下的旁白质量、`delivers_range_this_turn` 时的三句写法,需部署后按 `docs/testing/rectification-grounding-20260927.md` 走真机。
## 验收(Claude,2026-09-27)
基于 `19376089` 独立复跑(Node 22.14):tsc 0;lint 0 error;`npm test` 4152 / 24 fail / 28 skip,失败名单与 `cd4dde9d` 基线逐条一致;消失的 11 个测试名 = 随重试功能删除的 `rectification-v9-regenerate.test.ts` 8 条 + 3 条改名(有三栏说明);Python 门禁集退出 0;`/`、`/chart`、`/ephemeris`、`/people` ○ Static;rootMainFiles gzip(level 9)130933 B。未碰 `.gitea/`、`vendor/`、`deploy/`、迁移、任何 Python;删除的 3 个文件均为校正重试功能本身,源码里已无校正重试引用,普通对话重试按钮由 `onRegenerate` 控制仍在。数据库函数 `regenerate_agentic_rectification_turn` 保留,待另开单退役。
观察(不阻塞):数字白名单会把模型复述用户经历里自带的钟点(如「10:30 手术」)那一整句丢掉、由服务端 recap 顶上,措辞会略生硬;真机若常见再调。