fix(rectification): match the delivery inside the acknowledged reply (BUG-1241)
The answered choice persists the acknowledgement, a blank line, and the rewritten delivery. The exit fill-in required equality and wrote the deterministic delivery again. Check containment instead, and pin it on the real answer path. The suspected widen staleness regression did not hold; keep a guard test. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
6e16dc5e16
commit
e62d0a3257
+6
-6
@@ -16677,17 +16677,17 @@
|
||||
## BUG-1241 | 校正交付后同一轮又补出第二段回答
|
||||
|
||||
- 状态:resolved(分支 `codex/rectification-delivery-dup-adopt-20261006`)
|
||||
- 首次发现 / 最近更新:2026-10-06 / 2026-10-06
|
||||
- 来源:10-06 测试环境真机。任务书 `docs/tasks/TASK-rectification-delivery-dup-adopt-20261006.md`。
|
||||
- 首次发现 / 最近更新:2026-10-06 / 2026-10-06(验收修复单后)
|
||||
- 来源:10-06 测试环境真机。任务书 `docs/tasks/TASK-rectification-delivery-dup-adopt-20261006.md`;验收修复单 `docs/tasks/TASK-rectification-delivery-dup-adopt-fix-20261006.md`。
|
||||
- 影响面:生时校正答完最后一题后的交付。不改引擎,不改数据库。
|
||||
- 现象:同一轮出现两段助手回答。前一段是采用旁白的改写。后一段是确定性兜底,带分钟范围和「还剩几个候选」。刷新后两段都在。
|
||||
- 触发条件:答题路径已经把改写后的交付句写入本轮;收尾补位找不到已问轮次编号,又用子串去比对。改写句不是兜底原文的子串,于是再写一条。
|
||||
- 根因:BUG-1153 的防重只看「上一条是否包含确定性原文」。当时的测试把上一条写成已经包含这段原文,所以没覆盖「上一条是另一句改写」。60 秒内不重复记账的标记也没有拿来比对这句已经写下的交付。
|
||||
- 修复:答题回复在落库前记下这句交付原文。收尾补位若发现上一条助手原文(去掉空白后)就是这句,就不再写。子串比对保留。没有任何回复承载这句时,补位仍写 1 条。不把旁白函数传进收尾,也不改旁白提示词。选择题 JSON 在已有轮次编号时带上 `x-rectification-turn-id`,作为额外保险;真库验收不依赖这个头。
|
||||
- 验证:`frontend/tests/rectification-delivery-dup-adopt-20261006.test.ts` 覆盖「上一条是改写句则 0 条新写入」和「上一条不是这句则仍写 1 条」。BUG-1153 原测试仍通过。真 PostgreSQL 17 上,种好的并列卷宗走同一条收尾后,助手交付消息恰好 1 条。本机没有星历引擎,收尾里的刷新失败得很快,失败之后也没有写出第 2 条。
|
||||
- 防复发:上述测试。跳过条件是「上一条等于已记下的交付句」,不能只靠加长子串。
|
||||
- 修复:答题回复在落库前记下这句交付原文。收尾补位若发现上一条助手原文(去掉空白后)就是这句,就不再写。子串比对保留。没有任何回复承载这句时,补位仍写 1 条。**验收修复(b977f482 之后)**:首版要求上一条「等于」这句,但答题回复真实落库是「已记录,…」+ 空行 + 改写句(`persistApplied` 在 `skippedNextInterview` 时拼接答题回执),于是事故原样仍补出第二条;改为「上一条包含这句」(去空白比子串,不足 8 字不匹配)。不把旁白函数传进收尾,也不改旁白提示词。选择题 JSON 在已有轮次编号时带上 `x-rectification-turn-id`,作为额外保险;真库验收不依赖这个头。
|
||||
- 验证:`frontend/tests/rectification-delivery-dup-adopt-20261006.test.ts` 覆盖「上一条是改写句则 0 条新写入」和「上一条不是这句则仍写 1 条」;修复单补「回执前缀 + 改写」两种回执各 0 条、「只有回执」仍 1 条。`frontend/tests/rectification-answer-choice.test.ts`「a rewritten delivery reply is not followed by a second exit fill-in turn」走真实答题路径(`applyRectificationChoice` + 改写旁白)取实际落库文本再调收尾:b977f482 上 1 条(红),修后 0 条。BUG-1153 原测试仍通过。真 PostgreSQL 17 上,种好的并列卷宗走同一条收尾后,助手交付消息恰好 1 条。本机没有星历引擎,收尾里的刷新失败得很快,失败之后也没有写出第 2 条。
|
||||
- 防复发:上述测试。跳过条件是「上一条包含已记下的交付句」,比对对象必须取自真实答题路径的落库文本,不得手写只含改写句的夹具(首版回放就是这样漏掉回执前缀的)。
|
||||
- 相关记录:BUG-596、BUG-1149、BUG-1153。
|
||||
- 复发自:BUG-1153。旧测试的上一条已经包含确定性原文,改写句走不到那条断言。
|
||||
- 复发自:BUG-1153。旧测试的上一条已经包含确定性原文,改写句走不到那条断言。首版修复的测试与真库回放都只写了改写句、没有答题回执前缀。
|
||||
- 修复版本:待发布
|
||||
|
||||
## BUG-1242 | 点采用被拒,因为采用和读取没用答题时的出生日期
|
||||
|
||||
@@ -105,3 +105,48 @@ T2 做法:不把采用旁白函数传进收尾,也不改旁白提示词。
|
||||
`next build --webpack`。Turbopack 会拒绝工作树里指到主检出的 `node_modules` 联接,所以用 webpack。基线 `65b61ed5` 编译 2.9 分钟、构建自带的 TypeScript 100 秒,都通过,收集页面数据时停在 `/api/rectification/agent`。本分支编译 2.2 分钟、TypeScript 83 秒,都通过,收集页面数据时停在 `/api/birth-time-guide`。两边都是 Windows 不允许为 Skill 运行别名建符号链接(`EPERM`)。先碰到哪条路由取决于收集页面的工人,不是这次改动。路由表没有打出来,这台机器没能确认 `/` 仍是 Static,也没有首屏 gzip 数字。没有为了构建改符号链接,也没有在工作树里重装依赖。首页 `page.tsx` 没改;校正卡片组件有改,gzip 要等能建符号链接的环境再量。
|
||||
|
||||
本轮没有打开浏览器,也没有在手机上点采用。真机步骤在 `docs/testing/rectification-delivery-dup-adopt-20261006.md`。修复版本待发布,没有推送,没有部署。
|
||||
|
||||
## 验收修复单(Claude 直接执行,2026-10-06)
|
||||
|
||||
修复单:`docs/tasks/TASK-rectification-delivery-dup-adopt-fix-20261006.md`。在本分支 `b977f482` 之上追加。Linux,Node 22.14。
|
||||
|
||||
### F1 防重改为包含判断(BUG-1241 重开后修复)
|
||||
|
||||
- 事实:答题回复真实落库是「答题回执 + 空行 + 改写句」。`persistApplied` 在交付轮 `skippedNextInterview=true` 时拼 `${input.narration}\n\n${hostNarration}`。首版 `deliveryCarrierMatchesLastAssistant` 要求全等,事故原样仍补第二条。
|
||||
- 修:`delivery-turn-guard.ts` 改为 `text.includes(needle)`(去空白,不足 8 字不匹配)。
|
||||
- 新测试(均在 `b977f482` 上红、修后绿):
|
||||
- `rectification-delivery-dup-adopt-20261006.test.ts`:「已记录,范围没变。」+ 改写 → 0 条;「已记录,范围没变;…领先…落后。」+ 改写 → 0 条;只有回执 → 仍 1 条。
|
||||
- `rectification-answer-choice.test.ts`「a rewritten delivery reply is not followed by a second exit fill-in turn」:走 `applyRectificationChoice` 真实答题路径并注入改写旁白,取 `append_agentic_rectification_turn` 实际写入的文本(实测为「已记录,范围收到 …」+ 空行 + 改写),再调 `persistExhaustionGateTurn`:`b977f482` 上写入 1 条,修后 0 条。
|
||||
- 真库:本轮**没有**做 PG17 重放(修复单要求)。替代证据是上面这条:落库文本取自生产拼接代码,而不是手写夹具,首版恰恰是手写夹具漏掉了回执前缀。改动只是内存比较,不涉及 SQL。PG17 重放记为欠项。
|
||||
|
||||
### F2 撤回(修复单 §2.2 判断有误,BUG-1246 不成立,不入 Bug 历史)
|
||||
|
||||
修复单怀疑 `persistApplied` 删掉 `snapshotCurrent: input.snapshotCurrent` 后,放宽窗口失败时会按旧窗口出卡。复核结果:
|
||||
|
||||
- 放宽窗口(`mutateCaseForWidenWindow`)与选时段(`mutateCaseForBlockChoice`)进 `persistApplied` 时 `decisionState: null`。`decideAfterInferenceChange` 在 state 为空时 `candidateScores: []`,一律 `collect_evidence`,snapshotCurrent 不参与。
|
||||
- 唯一带 state 的调用(stop,`snapshotCurrent: rescored.snapshotCurrent`)里,这个值本来就是 `scoreableSnapshotCurrentFromDossier` 对同一卷宗算出来的。stored 与 current 只差证据指纹,revision 两边抵消,与函数内自算结果相同。
|
||||
- 结论:删那一行不改变行为。保留一条守护测试「answering widen A with a failed rescore never delivers or adopts on the old window's candidates」(对照组同一份存量快照按当前读是可采用的并列交付;放宽后 `can_adopt=false`、不出 `ready_to_adopt` / `complete_with_range`),修前修后都绿。BUG-1246 号码不使用。
|
||||
|
||||
### 本机测试(对照基线 `86cafe68`,同机同命令)
|
||||
|
||||
| | tests | pass | fail | cancelled |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| 基线 `86cafe68` | 4936 | 4871 | 24 | 0 |
|
||||
| `b977f482` | 4952 | 4886 | 25(多 1 条首页计时 flaky,单跑 17/17) | 0 |
|
||||
| 本轮 | 4957 | 4892 | 24 | 0 |
|
||||
|
||||
失败名与基线逐名相同(diff 为空)。新增 21 条,没有删除。`tsc --noEmit` 0;`npm run lint` 0 error,126 warning(与基线相同)。
|
||||
`next build`:本轮只改了服务端的 `delivery-turn-guard.ts` 和测试,没有重建。`b977f482` 的构建 Claude 已验:`/` 是 `○ Static`,首屏 gzip 636,730 对 636,565(+0.03%)。
|
||||
|
||||
既有断言:未改。
|
||||
|
||||
### 变基到 staging `da226db1` 后复验
|
||||
|
||||
staging 在这期间进了普通对话的代码(`da226db1`),本分支已变基到它之上。CHANGELOG、BUG_HISTORY、README 的冲突都是两边各追加,按「BUG 号升序」「两条 CHANGELOG 都留」解决。复验以 `da226db1` 为基线,同机同命令:
|
||||
|
||||
| | tests | fail | cancelled |
|
||||
| --- | --- | --- | --- |
|
||||
| 基线 `da226db1` | 4941 | 24 | 0 |
|
||||
| 本分支(变基后) | 4962 | 24 | 0 |
|
||||
|
||||
失败名逐名相同,测试名无消失,新增 21 条。tsc 0;lint 0 error。`next build`:`/` 是 `○ Static`;首屏 gzip 636,129 对 635,958(+171 B,+0.03%)。`tests/test_repo_privacy_markers.py` 通过。
|
||||
|
||||
@@ -423,5 +423,5 @@
|
||||
| `TASK-gate-log-volume-20261004.md` | `PROGRESS-gate-log-volume-20261004.md` | 门禁 validate 单步日志 4.4 万行 / 2 MB 网页打不开(run 3170):快速门只打摘要(失败给末 200 行 + 日志文件)、前端测试门禁上 dot + 失败汇总(本机仍 TAP)、拆分 validate 为 7 个 step(产品授权改 workflow,只限拆分与重定向);检查一项不少 | **已实现待验收**(BUG-1230;未推送、未部署;Gitea 各 step 页面是否打得开留待推送后由产品确认) | 分支 `codex/gate-log-volume-20261004` |
|
||||
| `TASK-consult-latency-quickwins-20261005.md` | `PROGRESS-consult-latency-quickwins-20261005.md` | 普通对话耗时两项快修:咨询链校正闸不再同步白等 VedAstro 官方快照子进程(每域约 4 s,复用顶层缓存 + 负缓存,不推翻 BUG-301);补分段计时(第 0/1 步耗时、推理 token、分类耗时进日志与 usage)。推理强度/精简说明待模型 key 另单(BUG-1231、BUG-1232) | 已验收(Claude 10-05 Linux:Python 62 / 前端 24 失败与基线逐名相同,Static、gzip 0%;装 SDK 实测每领域 4.47→1.66 s,剩余 1.33 s 是保留的顶层前台等待;/admin/usage 真机欠) | 分支 `codex/consult-latency-quickwins-20261005` |
|
||||
| `TASK-rectification-nadi-seconds-research-20261005.md` | `PROGRESS-rectification-nadi-seconds-research-20261005.md` | **「纳迪秒级校准」可证伪检验(离线)**:竞品宣传「问前事到天 → 秒级」。本仓主链只用三层小运,Sookshma / Prana 与 D150 从未进评价集;v5 真值 52/77 是整 5 分钟(秒级无真值可对)。N0 五层小运 + D150 底座(前三层与主链对账 0 差)、N1 拟合率 vs 安慰剂日期(核心)、N2 留一件预测、N3 六题后区间内再细分能否提头名、N4 岁差 / 坐标 / 时间扰动的噪声地板、N5 D150 结构层(原文比对 blocked)、N6 结论 + 对外口径草稿。规则先登记再跑;不改生产代码;不得重调 BUG-1091 已关的权重 | 已验收(Claude 10-05:N1/N2/N3/N5 不过门、N4 触发文案禁令;Linux 复跑内容 0 差,结果改 LF;补 Lahiri/KP/True Chitra 对照,月亮差 0.83′ 第 5 层仍换 94%)→ 不立实现单 | 分支 `codex/rectification-nadi-seconds-research-20261005`(BUG-1240 `closed_by_design`) |
|
||||
| `TASK-rectification-delivery-dup-adopt-20261006.md` | `PROGRESS-rectification-delivery-dup-adopt-20261006.md` | **校正交付一次两段回答 + 点「改用…盘解读」被拒 `adoption_not_allowed`**(10-06 真机):收尾补位按文字子串防重,遇采用旁白 Agent 改写版失效(BUG-1153 复发);accept/GET 调 `decideFromDossier` 不带出生日期,未成年探针复活致判定与出卡时相反(BUG-598/680 同族);卡片按钮不看公开 can_adopt、拒绝码直出英文。不放宽采用门,判定入参统一 + 结构防重 + 按钮判据 + 中文提示(BUG-1241~1243) | **验收未通过 → 修复单待领取**(Claude 10-06 验收 `b977f482`:T1 出生日期统一、T3 按钮/中文通过;T2 防重只认全等,事故原样「已记录,范围没变。」+改写仍补第二条;`persistApplied` 丢了放宽窗口/选时段后强制的 snapshotCurrent=false(BUG-1246)。tsc 0、lint 0 error、npm 失败与基线同名(+1 首页计时 flaky)、Static、gzip +0.03%) | 分支 `codex/rectification-delivery-dup-adopt-20261006`;修复单 `TASK-rectification-delivery-dup-adopt-fix-20261006.md` |
|
||||
| `TASK-rectification-delivery-dup-adopt-20261006.md` | `PROGRESS-rectification-delivery-dup-adopt-20261006.md` | **校正交付一次两段回答 + 点「改用…盘解读」被拒 `adoption_not_allowed`**(10-06 真机):收尾补位按文字子串防重,遇采用旁白 Agent 改写版失效(BUG-1153 复发);accept/GET 调 `decideFromDossier` 不带出生日期,未成年探针复活致判定与出卡时相反(BUG-598/680 同族);卡片按钮不看公开 can_adopt、拒绝码直出英文。不放宽采用门,判定入参统一 + 结构防重 + 按钮判据 + 中文提示(BUG-1241~1243) | **已验收,待推 staging**(Claude 10-06:首版 `b977f482` 验收 T1/T3 过、T2 未过;修复单 F1 由 Claude 直接执行——防重改为「上一条包含交付句」,真实答题路径测试修前红修后绿;F2 疑似 BUG-1246 复核不成立已撤回、留守护测试。全量 npm 失败名与基线逐名相同、tsc 0、lint 0 error;PG17 重放未做,见 PROGRESS) | 分支 `codex/rectification-delivery-dup-adopt-20261006`;修复单 `TASK-rectification-delivery-dup-adopt-fix-20261006.md` |
|
||||
| `TASK-consult-conversational-answer-20261006.md` | `PROGRESS-consult-conversational-answer-20261006.md` | **普通对话首轮去汇报骨架,改成聊天**(10-06 产品反馈「像机器在汇报」):首轮取消全部 `##`(三个固定节 + 按对象标题,多人改分段落点名)、行动不强制(盘上有具体指向才顺口一句,问「怎么办」再展开)、首轮三到六段;示范删可抄的行动句(实测被逐字照抄);思考栏去「这周可以做什么」。推翻 10-01 D2 与 D8 首轮部分(BUG-1244~1245) | 已验收 | 28113fa0 验收未过 → 修复单 `TASK-consult-conversational-answer-fix-20261006.md` 由 Claude 直接执行(含去掉助手头像),复验通过后推 staging;真机清单 10 步待产品 |
|
||||
|
||||
@@ -71,7 +71,9 @@ export function deliveryCarrierNarration(input: {
|
||||
}
|
||||
|
||||
/**
|
||||
* True when the latest assistant text is the stored delivery sentence.
|
||||
* True when the latest assistant text carries the stored delivery sentence.
|
||||
* The answered choice persists 「已记录,…」 + a blank line + that sentence,
|
||||
* so this is a containment check, not equality (BUG-1241 fix sheet).
|
||||
* Whitespace is ignored. A needle shorter than 8 characters does not match.
|
||||
*/
|
||||
export function deliveryCarrierMatchesLastAssistant(input: {
|
||||
@@ -91,7 +93,7 @@ export function deliveryCarrierMatchesLastAssistant(input: {
|
||||
if (turn?.role !== "assistant") continue;
|
||||
const text = typeof turn.text === "string" ? squash(turn.text) : "";
|
||||
if (!text) continue;
|
||||
return text === needle;
|
||||
return text.includes(needle);
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
@@ -2,7 +2,8 @@ import assert from "node:assert/strict";
|
||||
import { readFileSync } from "node:fs";
|
||||
import test from "node:test";
|
||||
|
||||
import { applyRectificationChoice } from "../src/lib/rectification-agentic/v9/answer-choice.ts";
|
||||
import { applyRectificationChoice, persistExhaustionGateTurn } from "../src/lib/rectification-agentic/v9/answer-choice.ts";
|
||||
import { resetDeliveryTurnGuardForTests } from "../src/lib/rectification-agentic/v9/delivery-turn-guard.ts";
|
||||
import { composeCollectSpokenAssistantText } from "../src/lib/rectification-agentic/v9/collect-prompt.ts";
|
||||
import { mergeTurnQuestions, type SnapshotTurnMessage } from "../src/lib/rectification-snapshot-messages.ts";
|
||||
import { choiceCardFromCaseDossier } from "../src/lib/rectification-agentic/v9/interview-state.ts";
|
||||
@@ -42,6 +43,7 @@ import {
|
||||
CASE_ID,
|
||||
EVIDENCE_ID,
|
||||
FOCUS_ID,
|
||||
RESULT_ID,
|
||||
SESSION_ID,
|
||||
TURN_ID,
|
||||
USER_ID,
|
||||
@@ -734,6 +736,53 @@ test("last structured choice emits the adoption range and persists the same narr
|
||||
assert.doesNotMatch(String(turn?.args.p_assistant_message ?? ""), /排盘用/);
|
||||
});
|
||||
|
||||
// BUG-1241 fix sheet: drive the real answer path with the adopt-narration
|
||||
// rewrite, take the assistant text it actually persists, and make sure the
|
||||
// exit fill-in does not add the deterministic delivery underneath it.
|
||||
test("a rewritten delivery reply is not followed by a second exit fill-in turn", async () => {
|
||||
resetDeliveryTurnGuardForTests();
|
||||
const rewrite = "因为再问下去也分不开这两个时间,这一轮到这里停住,下面是这次的结果。";
|
||||
const accounting = persistChoiceAccounting(adoptionDossier());
|
||||
const applied = await applyRectificationChoice(accounting.client, {
|
||||
userId: USER_ID,
|
||||
caseId: CASE_ID,
|
||||
sessionId: SESSION_ID,
|
||||
actionId: ACTION_ID,
|
||||
action: CHOICE_ACTION,
|
||||
focusId: FOCUS_ID,
|
||||
questionId: QUESTION_ID,
|
||||
probeId: RELOCATION_2015_PROBE.id,
|
||||
optionId: "A",
|
||||
expectedRevision: adoptionInference().revision,
|
||||
narrateAdopt: async () => rewrite,
|
||||
});
|
||||
assert.equal(applied.nextAction.can_adopt, true);
|
||||
const answerTurn = accounting.calls.find((call) => call.fn === "append_agentic_rectification_turn");
|
||||
const persisted = String(answerTurn?.args.p_assistant_message ?? "");
|
||||
assert.match(persisted, new RegExp(rewrite));
|
||||
|
||||
const deterministic = String(applied.narration).includes(rewrite) ? "" : String(applied.narration);
|
||||
const fillIn = deterministic || "目前范围 04:45–05:15。这只是代表性候选,不是已确认的唯一出生分钟。";
|
||||
const gate = fakeAccounting({
|
||||
get_agentic_rectification_case_dossier: () => dossierFixture({
|
||||
turns: [
|
||||
{ id: "t-user", role: "user", text: "A", status: "completed", created_at: "2026-10-06T00:00:00.000Z" },
|
||||
{ id: "t-answer", role: "assistant", text: persisted, status: "completed", created_at: "2026-10-06T00:00:01.000Z" },
|
||||
],
|
||||
}),
|
||||
append_agentic_rectification_turn: () => ({ turn_id: "44444444-4444-4444-8444-444444444444", idempotent: false }),
|
||||
});
|
||||
await persistExhaustionGateTurn({
|
||||
accounting: gate.client,
|
||||
userId: USER_ID,
|
||||
caseId: CASE_ID,
|
||||
hostNarration: fillIn,
|
||||
resultId: RESULT_ID,
|
||||
});
|
||||
assert.equal(gate.calls.filter((call) => call.fn === "append_agentic_rectification_turn").length, 0);
|
||||
resetDeliveryTurnGuardForTests();
|
||||
});
|
||||
|
||||
test("clicking A applies the choice without invoking a language model", async () => {
|
||||
const accounting = choiceAccounting();
|
||||
const applied = await applyRectificationChoice(accounting.client, {
|
||||
|
||||
@@ -14,7 +14,10 @@ import { RectificationRangeDelivery } from "../src/components/rectification-rang
|
||||
import { RectificationSegmentDelivery } from "../src/components/rectification-segment-delivery.tsx";
|
||||
import type { RectificationCandidateResult } from "../src/lib/rectification-candidate-result.ts";
|
||||
import type { RangeDeliveryProjection } from "../src/lib/rectification-agentic/v9/divergence-panel.ts";
|
||||
import { persistExhaustionGateTurn } from "../src/lib/rectification-agentic/v9/answer-choice.ts";
|
||||
import { applyRectificationChoice, persistExhaustionGateTurn } from "../src/lib/rectification-agentic/v9/answer-choice.ts";
|
||||
import { CHOICE_ACTION } from "../src/lib/rectification-agentic/v9/choice-action.ts";
|
||||
import { WIDEN_WINDOW_INTENT, WIDEN_WINDOW_KIND } from "../src/lib/rectification-agentic/v9/block-scan.ts";
|
||||
import { buildWidenWindowFrame } from "../src/lib/rectification-agentic/v9/window-widen.ts";
|
||||
import { dossierResponse } from "../src/lib/rectification-agentic/v9/case-dossier-response.ts";
|
||||
import { decisionInputsFromSnapshot } from "../src/lib/rectification-agentic/v9/decision-inputs.ts";
|
||||
import { decideFromDossier } from "../src/lib/rectification-agentic/v9/decision-from-dossier.ts";
|
||||
@@ -32,9 +35,12 @@ import {
|
||||
} from "../src/lib/rectification-agentic/v9/tool-service.ts";
|
||||
import { adoptionRefusalCopy, runRectificationCandidateAccept } from "../src/lib/rectification-chat-accept-run.ts";
|
||||
import {
|
||||
activeFocusFixture,
|
||||
CASE_DOSSIER_RPC,
|
||||
CASE_ID,
|
||||
candidateSnapshotFixture,
|
||||
FOCUS_ID,
|
||||
SESSION_ID,
|
||||
dossierFixture,
|
||||
fakeAccounting,
|
||||
RESULT_ID,
|
||||
@@ -135,7 +141,12 @@ function inference(withProbe: boolean, tied: boolean) {
|
||||
};
|
||||
}
|
||||
|
||||
function parsed(withProbe: boolean, tied: boolean, fingerprint: string | null): V9CaseDossier {
|
||||
function rawDossier(
|
||||
withProbe: boolean,
|
||||
tied: boolean,
|
||||
fingerprint: string | null,
|
||||
extra: { candidateRange?: { start_time: string; end_time: string }; activeFocus?: unknown } = {},
|
||||
) {
|
||||
const supportA = tied ? 16 : 40;
|
||||
const supportB = tied ? 16 : 8;
|
||||
const snapshot = candidateSnapshotFixture({
|
||||
@@ -162,10 +173,28 @@ function parsed(withProbe: boolean, tied: boolean, fingerprint: string | null):
|
||||
const rpc = dossierFixture({
|
||||
evidenceCount: evidence.length,
|
||||
evidence,
|
||||
candidateRange: { start_time: OPEN_START, end_time: OPEN_END },
|
||||
candidateRange: extra.candidateRange ?? { start_time: OPEN_START, end_time: OPEN_END },
|
||||
latestResult: snapshot,
|
||||
...(extra.activeFocus ? {
|
||||
conversationSummary: {
|
||||
confirmed_evidence_summary: [],
|
||||
pending_revisions: [],
|
||||
active_focus: extra.activeFocus,
|
||||
declined_skipped_topics: [],
|
||||
candidate_divergence_summary: null,
|
||||
missing_evidence_categories: [],
|
||||
last_result_policy: null,
|
||||
summary_version: 1,
|
||||
updated_at: "2026-10-06T10:00:00.000Z",
|
||||
},
|
||||
} : {}),
|
||||
});
|
||||
(rpc.case as { rectification_domain?: string }).rectification_domain = "general";
|
||||
return rpc;
|
||||
}
|
||||
|
||||
function parsed(withProbe: boolean, tied: boolean, fingerprint: string | null): V9CaseDossier {
|
||||
const rpc = rawDossier(withProbe, tied, fingerprint);
|
||||
const dossier = parseV9CaseDossier(rpc);
|
||||
if (!dossier) throw new Error("dossier did not parse");
|
||||
return dossier;
|
||||
@@ -547,3 +576,161 @@ test("a refused adoption refreshes the snapshot and says so in Chinese", async (
|
||||
assert.equal(adoptionRefusalCopy({ error: "adoption_not_allowed" }), "这次的结果还不能采用,已刷新到最新状态");
|
||||
assert.doesNotMatch(adoptionRefusalCopy({ code: "adoption_not_allowed" }), /adoption_not_allowed/);
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Fix sheet 2026-10-06 (F1, BUG-1241 reopened): the answered choice persists
|
||||
// 「已记录,…」 + blank line + the rewrite, never the rewrite alone.
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
for (const ack of ["已记录,范围没变。", "已记录,范围没变;04:53 仍领先,05:00 落后。"]) {
|
||||
test(`the exit fill-in does not write again when the reply is 「${ack}」 + the stored rewrite`, async () => {
|
||||
resetDeliveryTurnGuardForTests();
|
||||
markDeliveryTurn({ caseId: CASE_ID, resultId: RESULT_ID, narration: REWRITE });
|
||||
const accounting = gateAccounting(`${ack}\n\n${REWRITE}`);
|
||||
await persistExhaustionGateTurn({
|
||||
accounting: accounting.client,
|
||||
userId: USER_ID,
|
||||
caseId: CASE_ID,
|
||||
hostNarration: FALLBACK,
|
||||
resultId: RESULT_ID,
|
||||
});
|
||||
assert.equal(accounting.calls.filter((call) => call.fn === "append_agentic_rectification_turn").length, 0);
|
||||
resetDeliveryTurnGuardForTests();
|
||||
});
|
||||
}
|
||||
|
||||
test("the exit fill-in still writes once when the reply is only the acknowledgement", async () => {
|
||||
resetDeliveryTurnGuardForTests();
|
||||
markDeliveryTurn({ caseId: CASE_ID, resultId: RESULT_ID, narration: REWRITE });
|
||||
const accounting = gateAccounting("已记录,范围没变。");
|
||||
await persistExhaustionGateTurn({
|
||||
accounting: accounting.client,
|
||||
userId: USER_ID,
|
||||
caseId: CASE_ID,
|
||||
hostNarration: FALLBACK,
|
||||
resultId: RESULT_ID,
|
||||
});
|
||||
assert.equal(accounting.calls.filter((call) => call.fn === "append_agentic_rectification_turn").length, 1);
|
||||
resetDeliveryTurnGuardForTests();
|
||||
});
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Fix sheet 2026-10-06 (F2): guard only. Widening must never deliver or adopt
|
||||
// on the old window's tied candidates, even when the stored snapshot's
|
||||
// evidence fingerprint still matches and the rescore after widening fails.
|
||||
// (The suspected BUG-1246 did not hold: widen/block choices decide with no
|
||||
// inference state, so they always collect; this pins that.)
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
const NARROW = { start_time: RANGE_START, end_time: RANGE_END };
|
||||
const WIDE = { start_time: "04:15", end_time: "05:45" };
|
||||
|
||||
function widenFocus() {
|
||||
const frame = buildWidenWindowFrame({ reportedTime: TIME_B, currentWindow: NARROW });
|
||||
assert.ok(frame);
|
||||
return activeFocusFixture({
|
||||
intent: WIDEN_WINDOW_INTENT,
|
||||
targetDomain: null,
|
||||
targetKind: null,
|
||||
questionId: "widen_window:search_range:holdout",
|
||||
expectedAnswerSchema: {
|
||||
choice: {
|
||||
prompt: frame.prompt,
|
||||
option_a: frame.option_a_hint,
|
||||
option_b: frame.option_b_hint,
|
||||
option_c: frame.neither_label,
|
||||
option_d: frame.unsure_label,
|
||||
options: [
|
||||
{ key: "A", label: frame.option_a_hint, answer_class: frame.option_a_answer_class, role: "primary" },
|
||||
{ key: "B", label: frame.option_b_hint, answer_class: frame.option_b_answer_class, role: "primary" },
|
||||
{ key: "C", label: frame.neither_label, answer_class: frame.option_c_answer_class, role: "primary" },
|
||||
{ key: "D", label: frame.unsure_label, answer_class: frame.option_d_answer_class, role: "primary" },
|
||||
],
|
||||
},
|
||||
choice_kind: WIDEN_WINDOW_KIND,
|
||||
scoring: false,
|
||||
widen_windows: {
|
||||
A: { radius: 30, ...WIDE },
|
||||
B: { radius: 60, start_time: "04:00", end_time: "06:00" },
|
||||
},
|
||||
},
|
||||
});
|
||||
}
|
||||
|
||||
test("answering widen A with a failed rescore never delivers or adopts on the old window's candidates", async () => {
|
||||
const draft = parsed(false, true, null);
|
||||
const fingerprint = evidenceLedgerFingerprint(draft.evidence);
|
||||
// Control: the same stored snapshot, read as current, is a tied delivery that may be adopted.
|
||||
const control = publicDecisionFields(decideFromDossier(
|
||||
parsed(false, true, fingerprint),
|
||||
decisionInputsFromSnapshot(BIRTH, { currentEvidenceFingerprint: fingerprint }),
|
||||
));
|
||||
assert.equal(control.can_adopt, true);
|
||||
|
||||
const current = { value: rawDossier(false, true, fingerprint, { candidateRange: NARROW, activeFocus: widenFocus() }) };
|
||||
const accounting = fakeAccounting({
|
||||
[CASE_DOSSIER_RPC]: () => current.value,
|
||||
get_agentic_rectification_case_compute: () => ({
|
||||
case_id: CASE_ID,
|
||||
skill_version: "9.0.0",
|
||||
baseline_profile_fingerprint: "a".repeat(64),
|
||||
baseline_birth_snapshot: {
|
||||
birth_date: BIRTH,
|
||||
latitude: 31.23,
|
||||
longitude: 121.47,
|
||||
timezone_id: "Asia/Shanghai",
|
||||
timezone_offset: 8,
|
||||
birth_time_source: "approximate",
|
||||
reported_birth_time: TIME_B,
|
||||
},
|
||||
candidate_range: (current.value.case as { candidate_range: unknown }).candidate_range,
|
||||
}),
|
||||
append_agentic_rectification_turn: () => ({ turn_id: "44444444-4444-4444-8444-444444444444", idempotent: false }),
|
||||
apply_agentic_rectification_choice_action: (_fn: string, args: Record<string, unknown>) => ({
|
||||
action_id: args.p_action_id,
|
||||
status: "applied",
|
||||
idempotent: false,
|
||||
question_id: args.p_question_id,
|
||||
option_id: args.p_option_id,
|
||||
revision: 1,
|
||||
narration: args.p_narration,
|
||||
focus_status: args.p_focus_status,
|
||||
}),
|
||||
widen_agentic_rectification_dated_window: () => {
|
||||
current.value = rawDossier(false, true, fingerprint, { candidateRange: WIDE });
|
||||
return { case_id: CASE_ID, stage: "minute", candidate_range: WIDE };
|
||||
},
|
||||
});
|
||||
const savedEnv = {
|
||||
algorithm: process.env.RECTIFICATION_ALGORITHM_VERSION,
|
||||
policy: process.env.RECTIFICATION_DECISION_POLICY_VERSION,
|
||||
};
|
||||
process.env.RECTIFICATION_ALGORITHM_VERSION = ALGO;
|
||||
process.env.RECTIFICATION_DECISION_POLICY_VERSION = POLICY;
|
||||
let receipt: Awaited<ReturnType<typeof applyRectificationChoice>>;
|
||||
try {
|
||||
receipt = await applyRectificationChoice(accounting.client, {
|
||||
userId: USER_ID,
|
||||
caseId: CASE_ID,
|
||||
sessionId: SESSION_ID,
|
||||
actionId: "aaaaaaaa-aaaa-4aaa-8aaa-aaaaaaaaab46",
|
||||
action: CHOICE_ACTION,
|
||||
focusId: FOCUS_ID,
|
||||
optionId: "A",
|
||||
expectedRevision: 0,
|
||||
});
|
||||
} finally {
|
||||
for (const [key, value] of [
|
||||
["RECTIFICATION_ALGORITHM_VERSION", savedEnv.algorithm],
|
||||
["RECTIFICATION_DECISION_POLICY_VERSION", savedEnv.policy],
|
||||
] as const) {
|
||||
if (value === undefined) delete process.env[key];
|
||||
else process.env[key] = value;
|
||||
}
|
||||
}
|
||||
assert.ok(accounting.calls.some((call) => call.fn === "widen_agentic_rectification_dated_window"));
|
||||
assert.equal(receipt.snapshotCurrent, false);
|
||||
assert.equal(receipt.nextAction.can_adopt, false);
|
||||
assert.notEqual(receipt.nextAction.type, "ready_to_adopt");
|
||||
assert.notEqual(receipt.nextAction.type, "complete_with_range");
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user