Files
Jyotisha/docs/tasks/PROGRESS-rectification-grounding-20260927.md
T
Jesse_ChenandClaude Opus 5.5 261d7b2e04
Independent Staging Quality Gate / validate (push) Successful in 11m43s
Independent Staging Quality Gate / publish (push) Successful in 3m27s
docs(tasks): rectification grounding acceptance
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 03:36:29 +08:00

19 KiB
Raw Blame History

PROGRESS · 生时校正:旁白事实有据 + 去掉重试按钮 + 给模型的资料瘦身 · 2026-09-27

  • 执行方式:直接执行(产品负责人授权子代理执行;Claude 事后独立验收)
  • 任务书:docs/tasks/TASK-rectification-grounding-20260927.md(a98d3b20)
  • 基线:开工 origin/staging a98d3b20(诊断基于 599fe7a9);中途 staging 合入数据卡单(cd4dde9d),本分支已 rebase 到 cd4dde9d,最终对比基线全部在 cd4dde9d 上重跑;随后再 rebase 到 19376089(只改 docs/BUG_HISTORY.md 的纯文档提交,代码树与 cd4dde9d 相同)。
  • 分支:codex/rectification-grounding-20260927,工作树 .worktrees/rectification-grounding-20260927(本地提交,未推送)
  • BUG 编号:预留 BUG-1055~1058,开工与提交时核对(开工最大号 1054;rebase 后 staging 有 1059,本单 1055~1058 未被占用)
  • Skill:不 bump。jyotish-birth-time-rectification 10.0.31 文本未改,只是证据轮按章节切片发送;registry 不动。
  • 数据库:不动。regenerate_agentic_rectification_turn 保留,待另开单退役(AGENTS §7.6 两轮规则:本轮先下线代码)。

提交

任务 提交 内容
T1 0e0baaa7 BUG-1055:事实句不参与裁句、不先流后撤;P3 数字白名单(spoken-grounding.ts);batch 返回 range_after_rescore;系统提示「范围由系统说」
T4 5a808449 BUG-1058:deferFollowup 路径把答题前范围交给 Agent 轮(rangeBeforeTurn);路由级回归
T2 84eb2315 BUG-1056:删除校正重新生成(路由、服务端模块、Agent、只读工具集、客户端动作与状态);ChatMessageActions 无 onRegenerate 不画按钮;DESIGN / VOICE
T3 13d93020、5b9fbe5b BUG-1057:read-case 候选四字段、去重复列表、超限先砍候选明细;(后者只修测试里的类型转换)
T5 1ed3e555 R1:offer 走 compare 同一套剥离;报告 Markdown 只留一份;引擎宽度不给模型
T6 c4bb394b R4 + Skill 切片:证据轮只附 §5/§7;关 Mastra <available_skills> 注入
— 717a378b T3/T5 跟进:同样行为,改写法避免新增 3 条 unused-var 警告
T7 本提交 BUG 历史、CHANGELOG、本记录、真机清单、任务索引

T4 先于 T2 提交(T1 的 attempt 改动已带 rangeBeforeTurn 选项,T4 只接路由)。

各项怎么修的

项 做法
D1 / BUG-1055 attempt 结算时不再把「范围从…变为…」「这次没有重新比较」「候选比较这次没跑成」接进正文、也不流式放出;只把结构化事实(TurnServerFacts)和当轮数字白名单(SpokenFactWhitelist:结算后可信范围与区间、代表分钟、决策范围、吻合率)交给 finish。finish 顺序:① 证据轮对模型正文跑白名单(任何一句的时刻 / 时刻区间 / 百分比不在白名单就整句丢;一句不剩用 batch 回执复述)→ ② BUG-606 / BUG-615 裁句(只作用于模型正文)→ ③ composeSpokenWithServerFacts 接事实句(沿用原 withCompareFailedRetryNotice / withRescoreSkippedNotice / withRangeChangedAfterEvidence 的格式)→ ④ 交付时的缺口句 → 与已流式文本不同则一次 replace。batch 返回 range_after_rescore(credible_range、representative_time、event_fit_percent、delivers_range_this_turn = 无下一问且 deliveryNarrationAllowed),回执指纹仍按旧形状算。系统提示:删「有收窄就说进度」,改为范围由系统说、除出牌轮外不写时刻和百分比;出牌三句只在 delivers_range_this_turn=true 时写、数字只抄它(R3)。
D2 / BUG-1056 路由只服务校正 → 整条删除(Next 返回 404)。删 regenerate-turn.ts、getRectificationV9RegenerationAgent、createRectificationV9ReadOnlyTools、交付守卫的 "regenerate" 触发源、insertV9SkillRunReceipt 的 "regeneration" 分支类型;客户端删 regenerateRectificationMessage、regeneratingMessageKey、regenerating / canRegenerate / latestRegeneratableKey、问题空档输入的 regenerating。ChatMessageActions 的 onRegenerate / canRegenerate 改可选,不传不画按钮;普通对话照旧传。
D3 / BUG-1057 modelVisibleInference:候选只留 time / score(= 推断概率,保留三位小数)/ status / cluster_range;上一轮 eliminated_ids / score_deltas 改按时刻(eliminated_times)。删 candidate_summary.candidates。enforceTurnDecisionBudget 新顺序:去 cluster_range → 最好 6 个候选(未淘汰优先,candidates_omitted)→ 对话 4 条 / 证据 3 条 → 清空(truncated)。compare / offer 的 inference_state 用同一个瘦身视图。
D4 / BUG-1058 二选一选「把答题前范围交给 Agent 轮」:只出一条服务端范围句、走 D1 的事实句通道(不裁、不撤回),答题与经历两次变化合成一句;若把 applied.narration 并进正文,同一条消息会有两句范围。实现:typed-message 在点选焦点 deferFollowup 时写 turnState.rangeBeforeTurn = previous?.credible_range,Agent 轮用它当开轮范围。收集焦点拒答的 deferFollowup 不改推断,不传。
R1 offerModelProjection = agentVisibleLatestProjection(payload);后者再去掉 engine_indistinguishable_width_minutes 与 range_delivery.verification_markdown(与 skill_verification_report.markdown 逐字相同)。病例 API 的 latestResultToolProjection 与回执不变。
R3 见 D1:写正文时是否交付由 batch 的 delivers_range_this_turn 告诉模型,吻合率也在里面;数字仍过白名单。
R4 校正 Agent 加 inputProcessors: [rectificationSkillBoundProcessor](providesSkillDiscovery: "on-demand",与咨询 Agent 的做法相同):Mastra 不再加 <available_skills>(含临时目录路径)与「请调用 skill 工具」两条系统消息,agent.getSkill 仍按绑定版本加载。
Skill 切片 skill-slice.ts,规则常量 RECTIFICATION_SKILL_SECTIONS_BY_ACTION:证据轮按标题取「ConversationFocus 与意图承接」「批量证据与日期真实性」(10.0.x 的 §5、§7),前面加一句「本轮是证据轮…其余章节不适用」;开场、只读等其余动作发整份。按标题不按编号:9.0.0 的编号不同(它的 §5 是候选语言),找不齐标题就整份发,历史 Case 仍按绑定版本原文运行(BUG-621)。§6(全是 full_diagnostics 才有的字段:method_followup_plan、guided_collect_windows…)在证据轮不再发送。

红线测试清单

红线 测试
1 事实句不裁、流式终态 = 落库、不先出后撤 rectification-grounding-20260927.test.ts:repro 1(两句)、repro 1b(三句)、repro 2、「这次没有重新比较」、「候选比较这次没跑成」——都断言 states.at(-1) === persisted 且事实句一旦出现不再消失(assertNeverWithdrawn);rectification-grounding-route-20260927.test.ts 两条(路由级 shown === persisted)
2 时刻 / 区间 / 百分比 = 当轮事实 同文件:首句旧范围整句不落库、编造吻合率 90% 被丢 / 服务端 80% 保留、重算后范围与代表分钟可复述、白名单单元(整句门、不在句中删词)
3 9 个候选仍保留对话与证据 rectification-grounding-read-case-20260927.test.ts:真实引擎 9 候选 + 6 条对话;超限先砍候选明细;第一步只去 cluster_range
5 历史会话(BUG-621) rectification-grounding-prompt-20260927.test.ts:9.0.0 快照仍可 resolveExactSkillPackage,切片对它整份返回;Skill 未 bump,既有 skill-registry / rectification-history-open-20260909 测试照旧通过
D2 rectification-grounding-regenerate-20260927.test.ts(校正无重新生成、普通对话有、路由与代码已删、DB 函数仍在迁移里)
R1 rectification-grounding-projection-20260927.test.ts(offer / compare 模型视图,真实引擎响应)
R4 / 切片 / 体量 rectification-grounding-prompt-20260927.test.ts(真实 runV9AgentTurn + 真实 Agent + 记录提示词的假模型:无 <available_skills>、只含 §5/§7、每次调用固定开销 < 8K)

未改动基线上:rectification-grounding-20260927 的 8 条运行级用例全部失败;路由级 2 条在去掉修复后失败。

fixture:frontend/tests/fixtures/rectification-grounding-aa-30min.golden.json,诊断时本机真实引擎对公开 AA 盘(30 分钟窗、3 件公开事件)的 runV9CandidateScore 请求与响应(388KB,AGENTS §7.4)。

体量前后对比

测量方法:同一个脚本分别在 cd4dde9d(改前)与本分支(改后)上跑:真实 runV9AgentTurn(证据轮)+ 真实 getRectificationV9Agent(真实绑定 Skill 10.0.31、真实工具)+ 记录每次调用提示词的假模型;数据是上面的公开 AA 盘 fixture(9 个候选)。token 为估算:每个 CJK 字符(含全角标点)记 1,其余字符每 4 个记 1(estimateTokens,不是供应商分词器)。「固定开销」= 每次调用都会重发、与工具结果无关的部分:全部 system 消息 + 工具 schema。

每次调用 改前 改后
system 消息字符数 13,083 3,925
system 消息 token 8,211 2,680
工具 schema(第 1 步只有 read-case) 797 B / 229 797 B / 229
工具 schema(第 2 步起 14 个工具) 13,796 B / 3,791 13,796 B / 3,791
固定开销,第 1 步 8,440 2,909
固定开销,第 2 步起 12,002 6,471(目标 < 8K)
第 1 次调用合计输入 8,556 3,025
第 2 次调用合计输入(含 read-case 结果) 13,605 7,417

system 消息改前 = 系统提示 1.5K 字 + Skill 全文 10.7K 字 + Mastra 技能清单与「调用 skill 工具」约 0.95K 字;改后 = 系统提示 + Skill §5/§7。工具 schema 没动(按轮次收窄 schema 是可选项,本单未做,见让步)。

典型「read-case → compare → 写正文」三次调用合计(前两次实测,第三次 = 第二次 + compare 返回):改前约 8,556 + 13,605 + (13,605 + 10,488) ≈ 46.3K;改后约 3,025 + 7,417 + (7,417 + 8,129) ≈ 26.0K。

工具返回(给模型看的) 改前 改后
read-case,9 候选 + 2 条对话 5,639 B,truncated,对话 0 / 证据 0,候选明细 3,778 B 3,022 B,对话 2 / 证据 3,候选 754 B
read-case,9 候选 + 6 条对话 5,639 B,truncated,对话 0 / 证据 0 3,420 B,对话 6 / 证据 3
compare 38,026 B(约 10.5K token) 29,698 B(约 8.1K token)
offer 100,012 B 28,799 B

read-case 内容对比(9 候选):改前每个候选带 id、probability、posterior_score、rank、窗口定位五项(candidate_date / window_index / window_offset_minutes / segment_index / cluster_intervals),另有 candidate_summary.candidates 重复一份前 6 个;超限后对话与证据被清空。改后每个候选只有 time / score / status / cluster_range,无重复列表,对话与证据都在。

测试与构建(最终在 rebase 后的分支上跑)

项 基线 cd4dde9d 本分支 结论
tsc --noEmit — 0 错 通过
npm run lint 0 error / 127 warning 0 error / 127 warning(逐条同名单) 通过;中途多出的 3 条 unused-var 由 717a378b 改写法消掉
全量前端测试(Node 22.14,TAP 写文件避免 stdout 吞行) 4136 tests,fail 24 4152 tests,fail 24 失败名单逐条一致(24 条均为无 Docker 的 DB / 部署套件);新增 27 个测试名
消失的测试名 — 11 个 8 个是删掉的 rectification-v9-regenerate.test.ts;3 个改名:completed Agent replies restore feedback, copy and safe in-place regeneration actions、question gap: nothing is shown while busy, readonly, regenerating, or for non-resumable cases、reply regeneration is a separate Jyotisha agent with only read-case access(三栏见下表)
全量 Python 门禁集(gate-pytest-args.txt,含 test_api_server_growth_contract.py) 948 passed / 1 skipped 948 passed / 1 skipped 通过;未改 Python、scripts/jyotish_api_server.py 未动,伪造点与方法数不变
npm run build -- --webpack /、/chart、/ephemeris、/people ○ Static 同左 ○ Static 通过
rootMainFiles gzip(level 9,4 个文件) 131,145 B 130,933 B −0.16%,在 ±2% 内
首页 index.html 引用的 36 个 js/css gzip 661,107 B 660,413 B −0.10%
tests/*.py 读前端文本 — 已 grep 本单改动的文件名与文案(重新生成、有收窄就说进度、candidate_summary、消息组件等),无 Python 合同命中 —

失败名单对比方法:tsx --test --test-reporter=tap --test-reporter-destination=<file> tests/*.test.ts tests/*.test.tsx,取 ^(not )?ok N - 行排序后 comm;基线在 cd4dde9d 的一次性工作树上跑。构建后已删 frontend/frontend/。

改动的既有断言(原值 / 新值 / 原因)

每处在测试文件里都写了三栏注释,这里汇总:

文件 · 用例 原值 新值 原因
rectification-v9-agent · set-focus after spoken text… 第二句「范围收到 05:00 到 05:10。」 「这件事拿去和星盘对照了。」 P3:夹具无候选快照,手写时刻整句不落库;本用例锁 set-focus 之后不接第二段
rectification-v9-agent · agent-run strips question sentences… 首句「范围已经收到,收在 05:00–05:10。」 「2016 年 9 月去北京工作,这件事拿去和星盘对照。」 同上;剪问句断言不变
rectification-v9-agent · delivery persist keeps three sentences… 首句「目前范围 04:49–04:53。」,回执无吻合率 「代表分钟是 05:02。」,回执加 event_fit_rate 80% P3 只放行当轮事实;三句 / 四句裁三句的断言不变
rectification-v9-agent · reply regeneration is a separate… 断言重新生成 Agent 与只读工具 改名「rectification has no reply-regeneration agent…」,断言两者不存在 BUG-1056 删除
rectification-surface-state · question gap: nothing is shown while busy, readonly, regenerating… 含 regenerating: true → idle 去掉 regenerating 与该断言(改名) 输入已无此项
rectification-agentic-entry · completed Agent replies restore feedback, copy and safe in-place regeneration… 容器有重新生成请求、路由不计费且按绑定 Skill 改名「…feedback and copy, no regeneration」;断言校正不传 onRegenerate、无请求、路由文件不存在 BUG-1056
api-service-unavailable-20260904 路由清单含 regenerate 移出 文件已删
chat-composer-queue /if \(busy || regeneratingMessageKey\) \{/ /if \(busy\) \{/ 排队行为不变
rectification-answer-choice · attempt timeout stays under… 另断言 regenerate 路由 maxDuration 只断言 agent 路由 文件已删
rectification-delivery-ui-simplify · already_delivered… 断言 regenerate 路由 trigger: "regenerate" 删该断言 触发源随之删除
rectification-request-dossier-cache 路由清单含 regenerate 移出 文件已删
rectification-v9-stream · whole-run budget is less than both route maxDuration values 另断言 regenerate 路由 只断言 agent 路由(名称保留) 文件已删
rectification-settled-render-split 传 regenerating / canRegenerate / latestRegeneratableKey / regeneratingMessageKey 去掉 属性已删;渲染计数断言不变
11 个文件里的问题空档输入 regenerating: false,(delivery-vs-collect、open-collect-invite、post-adopt-verify、probe-pool-exhausted、replay-20260911、spoken-orphan、surface-state、targeted-card-live、targeted-spoken-focus-recovery、tiebreak-before-card、tied-first-fix)与 rectification-chat-run-fixtures 的 regeneratingMessageKey: null 有这一项 删除这一行 输入类型已无此项(tsc 会拒绝多余属性);各断言不变
rectification-v9-regenerate.test.ts 8 条重新生成用例 整份删除 测的是已删除的接口;替代为 rectification-grounding-regenerate-20260927.test.ts

BUG-588 原测试(rectification-probe-replay-loss-20260908)、BUG-593 原测试(rectification-delivery-report-facts)、BUG-615 源码合同(rectification-delivery-ui-simplify 的裁句顺序)未改、仍通过。

让步与未做

  • T6 可选项「按轮次收窄每步工具 schema」未做(第 2 步起 14 个工具 3,791 token 仍是固定开销的大头;固定开销已低于 8K,留下一单)。
  • compare 返回仍约 29.7KB(range_delivery.columns、window_scan、candidates、审计表等),本单只按任务书剥离了推断前数据、对照包、探针、重复 Markdown 与引擎宽度。
  • delivers_range_this_turn 是 batch 时刻按服务端交付判定(无下一问 + deliveryNarrationAllowed)给出的;finish 里 persistNextInterviewIfIdle 若再触发陈旧快照重算,最终是否交付以它为准(模型正文仍过白名单与三句上限)。
  • 证据轮切片不含 §9(候选语言);交付轮的边界句由系统提示第 5 条与白名单保证。若验收认为交付轮需要 §9,可在常量里加标题。
  • 数据库函数 regenerate_agentic_rectification_turn 待另开单退役。

环境缺口

  • 无 Docker:24 条 DB / 部署套件失败,与基线逐条一致(见上)。
  • 无真实模型凭据:真实模型下的旁白质量、delivers_range_this_turn 时的三句写法,需部署后按 docs/testing/rectification-grounding-20260927.md 走真机。

验收(Claude,2026-09-27)

基于 19376089 独立复跑(Node 22.14):tsc 0;lint 0 error;npm test 4152 / 24 fail / 28 skip,失败名单与 cd4dde9d 基线逐条一致;消失的 11 个测试名 = 随重试功能删除的 rectification-v9-regenerate.test.ts 8 条 + 3 条改名(有三栏说明);Python 门禁集退出 0;/、/chart、/ephemeris、/people ○ Static;rootMainFiles gzip(level 9)130933 B。未碰 .gitea/、vendor/、deploy/、迁移、任何 Python;删除的 3 个文件均为校正重试功能本身,源码里已无校正重试引用,普通对话重试按钮由 onRegenerate 控制仍在。数据库函数 regenerate_agentic_rectification_turn 保留,待另开单退役。

观察(不阻塞):数字白名单会把模型复述用户经历里自带的钟点(如「10:30 手术」)那一整句丢掉、由服务端 recap 顶上,措辞会略生硬;真机若常见再调。