Compare commits

...
Author SHA1 Message Date
Jesse_ChenandClaude Opus 5.5 c0f059125b fix(consult): progress timeline shows only real steps, not the answer outline (BUG-1145)
Independent Staging Quality Gate / validate (push) Successful in 14m43s
Independent Staging Quality Gate / publish (push) Successful in 3m46s
The answer outline (think.plan / thinking.section) was drawn as four think
rows and ticked done as soon as the chart calculation started. The reducer
now keeps the outline only as calculate-row detail; live and history views
share it. Product chose option B on 2026-10-01.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
2026-10-01 16:40:27 +08:00
Jesse_ChenandClaude Opus 5.5 f1a169cb93 feat(rectification): segment-gain order on by default up to 61 minutes after the health-focus fix; records and replay evidence (BUG-1143, BUG-1144)
Independent Staging Quality Gate / validate (push) Successful in 13m28s
Independent Staging Quality Gate / publish (push) Successful in 3m31s
The 77-case persisted replay after the BUG-1143 fix keeps every truth segment
at both 21 and 61 minutes with segment order on, so the default envelope moves
from 21 to 61 minutes. Bug numbers follow staging (BUG-1142 was taken).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-10-01 13:15:36 +08:00
Jesse_ChenandClaude Opus 5.5 564eede340 fix(rectification): health distinguish focuses persist again; segment-gain order on by default for windows <= 21 minutes (BUG-1142, BUG-1143)
BUG-1142: distinguish focuses wrote the planner domain verbatim, and
health_pressure is not in the focus table's target_domain check. The write
failed inside a swallowed retry and the session went straight to delivery
although tap cards remained. Every focus write now goes through
focusTargetDomain (persistableFocusDomain, BUG-672), and focus-to-probe
matching uses sameCollectDomain.

BUG-1143: segmentOrderEnabledFor turns segment-gain probe order on by default
when the scanned window is at most 21 minutes, where the 77-case persisted
replay gained head hits with no lost truth segment. RECTIFICATION_SEGMENT_ORDER
=off disables it; =on keeps the 61-minute research envelope.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-10-01 11:47:04 +08:00
18 changed files with 242522 additions and 51 deletions
+11
View File
@@ -1,5 +1,16 @@
# 印度占星 Skill 更新日志
## 2026-10-01 — 普通对话进度栏不再提前打勾
- 发问后进度栏不再一下子出现四行「先回答你问的这件事 / 把判断钉在盘上 / 时间窗口怎么看 / 这周可以做什么」并全部打勾(那时盘还没算完)。现在只显示真实发生的步骤:读取分析方法 → 计算本命盘 → 正在写(按正文实际写到的标题)→ 完成;打开历史对话也一样(BUG-1145)。
- Skill 不 bump;不改数据库结构;服务端事件不变。
## 2026-10-01 — 生时校正不再答两三题就提前结束;默认按「哪道题最能分开盘型」出题
- 修复:轮到「健康」方面的选择题时,这道题一直写不进去,系统就直接给出结果,后面还能问的题都不问了。修好后平均每次多问约 2 题(前后 10 分钟:2.8 → 4.6 题;前后 30 分钟:3.6 → 5.4 题),D9 / D10 判对的例数明显上升,没有一例把真实时间排除在外(BUG-1143)。
- 「按盘型挑题」默认开启(范围在前后 30 分钟以内时):优先问最能区分 D9 / D10 上升的题。77 例回放:前后 30 分钟 D9 判对 46 → 51 例、D10 52 → 55 例(BUG-1144)。
- Skill 版本不变;不改数据库结构。
## 2026-10-01 — 普通对话写回答最多可以等 5 分钟(原来 70 秒)
- 推理较重的模型(例如 deepseek-v4-pro)写长回答不再在最后一节被截断;5 分钟只用来防模型卡死,不限制回答长度(BUG-1142)。
+45
View File
@@ -15347,3 +15347,48 @@
- 相关记录:BUG-1051、BUG-1053、BUG-944。
- 复发自:无
- 修复版本:`codex/consult-answer-clock-20261001`
## BUG-1143 | 健康类区分选择题写不进焦点表,流程静默提前交付
- 状态:resolved(2026-10-01 Claude 直接执行;77 例真库持久回放修前修后对比;未部署时以部署后 health 为准)
- 首次发现 / 最近更新:2026-10-01 / 2026-10-01
- 影响面:`frontend/src/lib/rectification-agentic/v9/server-focus.ts`(区分类焦点写 `targetDomain`)、`answer-choice.ts::persistFocusAfterChoice`(写入失败被吞)、所有生时校正用户。
- 用户现象:答完两三道选择题就直接出交付卡,后面还有家人、感情、搬迁等可答的题却不再问;与「按盘型挑题」开关无关。
- 触发条件:任一轮的区分题落在健康领域(引擎领域名 `health_pressure`)。
- 根因:区分类焦点把计划层领域名原样写入 `agentic_rectification_conversation_focuses.target_domain`;该列检查约束只收 `education/career/relationship/relocation/finance/health/family/other`。写入抛 `target_domain_check` 违约,被 `persistFocusAfterChoice` 的 catch 吞掉(日志仅 `retry next focus failed reason=tool_failed`),流程落到 exhaustion 出口 `terminalNote` 交付。采集类焦点早经 `persistableFocusDomain` 映射,09-13 `530f260f`(BUG-661~663)的区分类写入分支漏掉。假数据库测试不校验约束,所以一直没测出来。
- 修复:新增 `focusTargetDomain`,所有意图的焦点写入都经 `persistableFocusDomain`;`method-followup.ts` 三处焦点与探针的领域比较改走 `sameCollectDomain`(BUG-672 规则);不改数据库约束。
- 验证:`frontend/tests/rectification-focus-target-domain-20261001.test.ts`(引擎全部领域名 × 四种意图映射后都在迁移的约束清单内;健康区分焦点存 `health` 且与 `health_pressure` 探针同域;源码合同禁止原样写领域名与 `probe.domain === focus.targetDomain`)。真库(PostgreSQL 17 替身)持久回放 77 例 × ±10/±30(`docs/research/varga_resolution_persisted_onoff_{before,after}_fix_2026_10_01.json`):平均问题数 ±10 2.82→4.56、±30 3.58→5.43;问满 6 题 21→50、30→69;默认顺序下 D9 / D10 头段命中 ±10 65→70 / 64→65、±30 38→46 / 39→52;真值段保留全部 76/76;修后日志 `retry next focus failed` 0 条。
- 防复发:焦点写入只能经 `focusTargetDomain`;引擎领域与表约束的对照由测试逐值锁定。
- 相关记录:BUG-672(同族:健康 / 职业别名)、BUG-586、BUG-661~663、BUG-1144。
- 复发自:BUG-672(别名未经归并,写入侧新分支漏掉)。
- 修复版本:`codex/rectification-segment-order-20261001`。
## BUG-1144 | 「按盘型挑题」默认关闭
- 状态:resolved(2026-10-01;默认开启窗口 ≤ 61 分钟)
- 首次发现 / 最近更新:2026-10-01 / 2026-10-01
- 影响面:`core/segment-probe-order.ts`、`v9/score-persist.ts`。
- 用户现象:盘型口径实现(BUG-1116)把按段选题做成环境变量 `RECTIFICATION_SEGMENT_ORDER=on` 才启用,生产默认按引擎信息增益出题,分盘头段命中低于研究「按段」列。
- 根因:实现单允许 T5 延后;执行方因真库持久回放未跑通而默认关闭。
- 修复:`segmentOrderEnabledFor(windowMinutes, flag)`:默认窗口 ≤ 61 分钟开启,`off` 关闭,`on` 为研究包络。
- 验证:修复 BUG-1143 后的 77 例真库持久回放,开 vs 关:±10 D9 70→71、D10 65→67;±30 D9 46→51、D10 52→55;真值段保留四格全部 76/76。修 BUG-1143 前 ±30 开启曾丢 3 例真值段(`…before_fix…json`),根因即 BUG-1143(开启后健康题被排到第一位,更早撞上写入失败)。`frontend/tests/rectification-segment-order-default-20261001.test.ts`。
- 防复发:放宽到 > 61 分钟前须用同一回放脚本在对应窗口重跑并逐格核对真值段保留。
- 相关记录:BUG-1116、BUG-1143、BUG-1105。
- 复发自:无
- 修复版本:`codex/rectification-segment-order-20261001`。
## BUG-1145 | 普通对话进度栏提前打勾:回答提纲在算盘前就显示为「已完成」
- 状态:resolved(2026-10-01 Claude 直接执行;单测复现修前失败、修后通过;未部署时以部署后 health 为准)
- 首次发现 / 最近更新:2026-10-01 / 2026-10-01
- 影响面:`frontend/src/lib/consultation-run-timeline.ts`(普通对话进度栏的实时与历史重建共用同一个 reducer);生时校正进度栏不用这个 reducer,不受影响。
- 用户现象:发问后,「先回答你问的这件事」「把判断钉在盘上」「时间窗口怎么看」「这周可以做什么」四行立刻全部打勾,而这时「正在计算本命盘…」还在转,正文一个字都没写。
- 触发条件:任何本命普通对话首轮(通用 / 每日 / 申报时段的计划节同理)。
- 根因:服务端在第一个流块就下发回答提纲(`think.plan` + 每节一条 `thinking.section`);reducer 把每节做成一行 live 的 think 行,而 `completeLiveThink` 在下一个事件(下一节、`tool.started`)到来时把所有 live think 行标成完成。提纲是「回答将怎么写」,不是已走过的步骤,却按步骤渲染,于是算盘一开始四行就假完成。自 09-18 进度栏设计(`94c1e81f`)起即存在;BUG-1132 之后回答按人分段、「盘上依据」「时间怎么看」可省,提纲与真实正文进一步对不上。
- 修复:产品 10-01 选 B——提纲不再成行。`think.plan` 不改状态;`thinking.section` 只把领域与对照项暂存为计算行的明细(`calculateHints`),在计算行出现时补上(实时与历史一致);正文开始(`answer.delta`)时计算行收口为完成;历史重建的思考文本改走 `thinking.delta` 行,不再挂在提纲行上。服务端仍下发 `think.plan` / `thinking.section`(公共事件合同与存档字段不变,供 `applyThinkingSectionProgress` 等使用),客户端只是不再把它们画成行。
- 验证:`frontend/tests/consultation-run-timeline.test.ts` 新增三条(算盘前提纲不成行、只有「读取分析方法」完成;整轮结束无提纲行、无 live 行;历史重建与实时同形且计算行仍有明细),三条在修前代码上失败、修后通过;改写一条既有断言(三栏见测试注释)。
- 防复发:进度栏只画真实发生的步骤(方法、计算、写作、模型推理);回答提纲若要再上屏,必须由正文标题驱动、不能用「下一事件到来即完成」的规则。
- 相关记录:BUG-942~944(进度 / 思考 / 正文分通道)、BUG-1074(发出即显示排队行)、BUG-1132(按人分段)。
- 复发自:无
- 修复版本:`codex/consult-timeline-no-plan-20261001`。
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,34 @@
# PROGRESS · 普通对话进度栏去掉回答提纲行(2026-10-01)
任务书:`docs/tasks/TASK-consult-timeline-no-plan-20261001.md`;BUG-1145。
## 实现
- `frontend/src/lib/consultation-run-timeline.ts`:
- `think.plan` 直接返回原状态(不再生成提纲行)。
- `thinking.section` 只把领域与对照项累积到新增的 `calculateHints`,经 `applyCalculateHints` 补到计算行;`tool.started` 建计算行时也补上(提纲在算盘前下发,原 `enrichCalculate` 在实时流里因计算行尚不存在而落空,历史重建里才有明细;现在两边一致)。
- `answer.delta` 同时把计算行收口为完成(原来由 `thinking.section` 顺手收口)。
- `consultationTimelineFromSettled`:有思考文本就发一条 `thinking.delta`,删掉「把思考文本挂到最后一条提纲行」的补丁。
- 服务端 `think.plan` / `thinking.section` 照发(公共事件合同、`consultation-agentic-runtime` 测试锁定 `think.plan` 存在;客户端已不再用它画行)。`thinkPlanSteps` 仍被服务端与测试使用,未删。
## 改动的断言(原值 / 新值 / 原因)
| 测试 | 原值 | 新值 | 原因 |
| --- | --- | --- | --- |
| `consultation-run-timeline.test.ts` 首条(改名为 “…then the first write (no outline row)”) | 行种类 `["method","calculate","think","write"]`,第 3 行「先回答你问的这件事」已完成 | 行种类 `["method","calculate","write"]`,不存在该标签行;写作行下标 3→2 | BUG-1145:提纲不是步骤,产品选 B 去掉 |
新增三条:算盘前提纲不成行且只有方法行完成;整轮结束无 think 行、无 live 行;历史重建同形且计算行有明细。三条与改写的一条在修前代码上均失败(已实测:4 fail),修后通过。
## 门禁(Node 22.14)
| 项 | 基线 `f1a169cb` | 分支 |
| --- | --- | --- |
| tsc | — | 0 |
| lint | — | 0 error / 126 warning |
| npm test | 4,851 条,fail 24 | 4,854 条,fail 24(同名) |
| build `/` | — | Static |
| `.next/static` js gzip 合计 | 1,632,680 B | 1,632,662 B(−18 B) |
## 环境缺口
- 浏览器级(真机看进度栏)留给 `docs/testing/consult-plain-answer-20261001.md` 第 10 步。
@@ -0,0 +1,49 @@
# PROGRESS · 生时校正:健康区分题写入失败 + 按盘型挑题默认开启(2026-10-01)
执行:Claude 直接执行(产品负责人 2026-10-01「同意」:前后 10 分钟先开、修复后重测前后 30 分钟,过红线再放开)。分支 `codex/rectification-segment-order-20261001`。无单独任务书;背景见 `TASK-rectification-varga-resolution-20260930.md` T5 与 `docs/research/rectification_varga_resolution_2026_09_30.md` §4。BUG-1143、BUG-1144。
## 经过
1. 77 例 × ±10 / ±30 真库持久回放(PostgreSQL 17 替身;`frontend/tests/research/rectification-segment-persisted-replay.ts` 经 `scripts/research/varga_resolution_persisted_replay.py`)比较按段选题关 / 开:±10 开启头段多命中 6 例、真值段 0 丢失;±30 开启丢 3 例真值段(`docs/research/varga_resolution_persisted_onoff_before_fix_2026_10_01.json`)。
2. 逐例插桩(临时,未提交):丢失不是排序删题。区分类健康题写焦点时 `target_domain=health_pressure` 违反表约束,写入失败被吞,流程从 exhaustion 出口提前交付;开启排序只是把健康题排到第一位、更早撞上。默认顺序下同样发生(BUG-1143,BUG-672 同族复发)。
3. 修 BUG-1143 后重跑(`…after_fix_2026_10_01.json`)。
## 77 例真库持久回放(修前 → 修后)
| 窗口 / 顺序 | 平均问题数 | 问满 6 题 | D9 头段命中 | D10 头段命中 | 真值段保留 D9 / D10 |
| --- | --- | --- | --- | --- | --- |
| ±10 默认 | 2.82 → 4.56 | 21 → 50 | 65 → 70 | 64 → 65 | 76/76 → 76/76 |
| ±10 按段 | 2.87 → 4.61 | 22 → 52 | 68 → 71 | 67 → 67 | 76/76 → 76/76 |
| ±30 默认 | 3.58 → 5.43 | 30 → 69 | 38 → 46 | 39 → 52 | 76/76 → 76/76 |
| ±30 按段 | 3.53 → 5.43 | 29 → 69 | 41 → 51 | 40 → 55 | 74/75 → 76/76 |
修后日志 `retry next focus failed`:0 条。±60 未回放(排序包络上限 61 分钟,>61 不启用)。
## 改动
| 文件 | 改动 |
| --- | --- |
| `v9/server-focus.ts` | 新增 `focusTargetDomain`,所有意图焦点写入经 `persistableFocusDomain` |
| `v9/method-followup.ts` | 三处焦点 ↔ 探针领域比较改 `sameCollectDomain` |
| `core/segment-probe-order.ts` | `segmentOrderEnabledFor(window, flag)`:默认 ≤ 61 分钟开、`off` 关、`on` 研究包络 |
| `v9/score-persist.ts` | 用上面函数代替 `RECTIFICATION_SEGMENT_ORDER === "on"` |
| 测试 | `rectification-focus-target-domain-20261001.test.ts`(3)、`rectification-segment-order-default-20261001.test.ts`(2) |
既有断言改动:无。
## 门禁(对 `origin/staging`,替身真库除外均为本机)
| 检查 | 基线 | 本分支 |
| --- | --- | --- |
| tsc / lint | 0 / 0 error、126 warning | 0 / 0 error、126 warning |
| npm test | 4846 / fail 24 | 4851 / fail 24;失败清单逐名相同,基线名 0 丢失,新增 5 |
| 按镜像 COPY 清单构建 | — | exit 0 |
| `/` / 首屏 gzip | Static / JS 555,622、CSS 68,813 | Static / 相同 |
| 快速门 | — | Python 1036 passed / 1 skipped / 0 failed |
默认窗口由 21 放宽到 61 后:受影响 6 个测试文件 322 / 322、tsc 0。
## 缺口
- 真库证据为 PostgreSQL 17 替身;本次不改数据库结构。
- 真机:部署后在 staging 走一次前后 30 分钟窗口的校正,确认能问到健康题、能问满几题。
+2
View File
@@ -264,6 +264,7 @@
| `TASK-rectification-message-cleanup-20261001.md` | `PROGRESS-rectification-message-cleanup-20261001.md` | **校正消息清理**:题干两遍(作废旧焦点仍打印题干 + 独立问题块再打印一次)→ 一屏一题干;证据轮正文改为「记下了 N 件事」+ 服务端清单可展开;校正结算后不显示「已完成 N 步」与「1.」泄漏。推翻 BUG-917 作废题留题干、VOICE 逐条复述 | 已验收(Claude 10-01:全套门禁对 staging 209352eb 一致;「1.」来源为推断——收起的有序列表复制补号,结算后整块不渲染已消除;真机待产品) | BUG-1135~1137 |
| `TASK-rectification-chart-tier-caveats-20261001.md` | — | **报告与聊天标注判断不了的分盘**:采用后读 `active_birth_provenance`(segment-v1)→ 结果 `segment_summary` 档位,blocked/indistinct 的 D9/D10 在聊天 answer_policy 与一句说明、报告章节顶部提示 + 主题降级;旧资料逐字节不变;不改库、不改 consult/route.ts | 待领取 | BUG-1138~1139 |
| `TASK-rectification-angle-timing-research-20261001.md` | `PROGRESS-rectification-angle-timing-research-20261001.md` | **角度类时间技法研究(离线)**:行运土木罗压本命上升 / 天顶度数、次限推运角、太阳弧角(1° ≈ 4 分钟出生时间),77 例 v5;乱序日期对照 ≥ 50 次必做、留一法、不得挤出真值;显著才叠加到六题后验比头段命中与区间宽度 | 已实现待验收(0/9 预登记检验显著,BUG-1141 closed_by_design;分支 codex/rectification-angle-timing-research-20261001,未推送) | BUG-1141 |
| (直接执行,无任务书) | `PROGRESS-rectification-segment-order-20261001.md` | **健康区分题写入失败(BUG-1143)+ 按盘型挑题默认开启(BUG-1144)**:区分类焦点把 `health_pressure` 原样写入、违反表约束、写入被吞后提前交付(BUG-672 同族);修后 77 例真库回放平均问题数 ±10 2.8→4.6、±30 3.6→5.4,D10 头段 ±30 39→52;按段选题 ≤61 分钟默认开,±30 D9 46→51、D10 52→55,真值段 76/76 | 已验收(Claude 10-01 直接执行) | 分支 codex/rectification-segment-order-20261001 |
| `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 已验收关单(Claude 10-01:三类技法 9 项检验 0 显著,真值平均排位 0.49~0.59;Claude 补正对照——埋入信号排位 0.077、挪年 0.502——流程有效;BUG-1141 closed_by_design) | BUG-1088、1089 |
| `TASK-report-reader-polish-20260929.md` | `PROGRESS-report-reader-polish-20260929.md` | **报告页打磨**:生成入口挪进页面主体(删标题栏按钮)、详情页到底部按钮(懒渲染一次到底)、导出按钮带文字、目录一级/二级分层、去掉「字段」与 RL/NL 缩写表头、状态列同义重复去重、报告表格淡底色、分块导出显示真实文件大小 | Claude 直接执行并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +7 B、快速门 Python 1000 passed),已部署 staging `d7772011`,真机欠 | BUG-1092~1094 |
| `TASK-report-english-edition-20260929.md` | `PROGRESS-report-english-edition-20260929.md` | **报告中英两版**:同一次引擎计算渲染 zh/en 两遍(不用模型翻译),瑜伽库 477 条补英文、模板与前端表头英文化;中文逐字节不变、英文零汉字、两版数字序列一致;阅读页 `?lang=en` 切换,导出当前语言 + 「问 AI 建议导出英文版」提示;旧报告不补英文 | Claude 直接执行(含两个 fork 子代理)并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +0.06%、Python 定向全绿),已部署 staging `d7772011`,真机欠 | — |
@@ -399,3 +400,4 @@
- **实现合入 `staging` 的同一次推送里,必须同时把状态板那一行改掉。** 2026-09-16 的对账发现 10 份早已合入的单仍写着「待领取 / 待验收」,会导致重复派活。
| `TASK-consult-plain-answer-20261001.md` | `PROGRESS-consult-plain-answer-20261001.md` | 普通对话「还是废话」:09-17 四步开场形状(格局名 → 谁推谁修 → 扮演哪个象)逼出谜语,问父母时爸妈被揉成一段;产品授权推翻该形状,改成先答 + 按问题里的对象分段 + 人话自检 + 空宫不单独下结论 + 父母卡标 mother/father(BUG-1132~1134) | **已实现,待 Claude 验收**(fork 子代理直接执行);未推 staging、未部署;模型对比为环境缺口,真机清单 `docs/testing/consult-plain-answer-20261001.md` | 分支 `codex/consult-plain-answer-20261001` |
| `TASK-consult-answer-clock-20261001.md` | `PROGRESS-consult-answer-clock-20261001.md` | 普通对话答题时钟 70 s 容不下推理模型(v4-pro 77 s),无出生分钟路线只受 110 s 工具钟管;产品定 5 分钟防卡死(BUG-1142) | **已实现,待 Claude 验收**(直接执行) | 分支 `codex/consult-answer-clock-20261001` |
| `TASK-consult-timeline-no-plan-20261001.md` | `PROGRESS-consult-timeline-no-plan-20261001.md` | 普通对话进度栏提前打勾:回答提纲在算盘前就作为四行步骤显示并全部打勾(BUG-1145);产品选 B 去掉提纲行,只显示真实步骤 | **实现完成,待 Claude 验收**(直接执行) | 分支 `codex/consult-timeline-no-plan-20261001` |
@@ -0,0 +1,31 @@
# TASK · 普通对话进度栏去掉回答提纲行(2026-10-01)
- 基线:`origin/staging` `f1a169cb`
- 分支 / worktree:`codex/consult-timeline-no-plan-20261001` / `.worktrees/consult-timeline-no-plan-20261001`
- 模式:直接执行(产品 10-01 选 B 后明确要求直接做),Claude 派子代理实现、Claude 独立验收
- BUG:BUG-1145
## 事故实证
产品 10-01 真机截图:问「我跟我父母的关系怎么样」,进度栏「读取分析方法」之后立刻出现「先回答你问的这件事 / 把判断钉在盘上 / 时间窗口怎么看 / 这周可以做什么」四行且全部打勾,最下面「正在计算本命盘…」仍在转。
## 根因
`frontend/src/lib/stream-agent-response.ts` `flushThinkingPlan` 在第一个流块就下发回答提纲(`think.plan` + 每节 `thinking.section`);`frontend/src/lib/consultation-run-timeline.ts` `reduceConsultationTimeline` 把每节做成 live think 行,`completeLiveThink` 在下一个事件到来时把它们全标完成。提纲不是走过的步骤,却按步骤渲染。
## 决策记录
产品 10-01 在两案中选 **B:去掉提纲行,只显示真实步骤**(读取分析方法 → 计算本命盘 → 正在写 → 完成)。未选 A(提纲等算完再出现、按正文标题逐条打勾),原因:BUG-1132 之后回答按人分段、两节可省,提纲与正文对不上;产品偏好「多余入口宁可删除」。
## 硬红线
1. 只改渲染层(timeline reducer);服务端事件、`thinkingPlan`、`applyThinkingSectionProgress`、存档字段、API / DB 结构不动。
2. 生时校正进度栏不动(它不用这个 reducer,只共用行类型与组件)。
3. `page.tsx` 不增长;不加 spinner / 骨架。
## 任务与验收
- T1 reducer:`think.plan` 不改状态;`thinking.section` 只暂存计算行明细并在计算行出现时补上;正文开始时计算行收口。验收:新增回归测试在修前失败、修后通过。
- T2 历史重建与实时同形(`consultationTimelineFromSettled` 走同一 reducer;思考文本改走 `thinking.delta` 行)。验收:历史重建测试。
- T3 记录:BUG_HISTORY、CHANGELOG、DESIGN.md、真机清单第 10 步、本任务书与进度记录、索引。
- 门禁:tsc 0;lint 0 error;`npm test` 失败名单与基线同名;`next build` `/` Static;首屏 gzip ±2%。
@@ -13,3 +13,4 @@
| 7 | 第 1 步的回答如果明显被截断(最后一节没写完就停了) | 记下时间和模型名报给 Claude——这是取消字数上限后要盯的风险(见进度记录「长度副作用核对」) |
| 8 | (BUG-1142,答题时钟改 5 分钟后)选 `deepseek-v4-pro` 新开对话,问「我和父母关系如何,他们怎么对待我」 | 回答完整写到「这周可以做的一件事」,不出现「回答未完成」;记下从发出到写完大约多少秒 |
| 9 | 用没填出生分钟的人物档案、选 `deepseek-v4-pro` 问一个长问题(例如「完整讲讲我今年的事业和感情」) | 回答写完,不出现「回答未完成」 |
| 10 | (BUG-1145)新开对话问任一问题,盯住回答上方的进度栏 | 只出现「读取分析方法」「计算本命盘」「正在写…」这类真实步骤;不再出现「先回答你问的这件事 / 把判断钉在盘上 / 时间窗口怎么看 / 这周可以做什么」四行,也没有在算盘时就打上的勾;刷新后打开这条历史对话,进度栏一样 |
+1 -1
View File
@@ -482,7 +482,7 @@ The birth-time rectification session is the consultation transcript plus a house
#### Streaming states
Every assistant reply moves through the same states on both chat surfaces, and each state has exactly one visual. The step timeline (`ConsultationRunTimeline`) is the only activity surface. Consultation thinking is complete `think.step` items (v1 `thinking.delta` still renders for one release). Rectification never shows thinking text: provider reasoning is dropped at the public boundary.
Every assistant reply moves through the same states on both chat surfaces, and each state has exactly one visual. The step timeline (`ConsultationRunTimeline`) is the only activity surface. Its rows are only steps the run has actually taken — reading the method, calculating the chart, writing (one row per heading the answer has written), and model reasoning when present. The answer outline the server plans (`think.plan` / `thinking.section`) is never a row: it only details the calculate row (2026-10-01, BUG-1145; it used to show four outline rows that ticked done the moment the calculation started). Consultation thinking is complete `think.step` items (v1 `thinking.delta` still renders for one release). Rectification never shows thinking text: provider reasoning is dropped at the public boundary.
| State | When | Visible | Transition in |
|---|---|---|---|
+38 -35
View File
@@ -30,6 +30,12 @@ export type ConsultationTimelineRow = Readonly<{
export type ConsultationTimelineState = Readonly<{
rows: readonly ConsultationTimelineRow[];
answer: string;
/**
* What the run's planned sections say the calculation looks at, held until
* the calculate row exists: the plan is sent before the tool starts, and it
* only ever details that row (BUG-1145).
*/
calculateHints?: Readonly<{ queries: readonly string[]; sources: readonly string[] }>;
}>;
const METHOD_ID = "method";
@@ -55,12 +61,13 @@ export function reduceConsultationTimeline(
return completeRow(state, METHOD_ID, CONSULTATION_DONE_SKILL_LABEL);
}
if (event.type === "tool.started") {
return upsertRow(completeLiveThink(state), {
const created = upsertRow(completeLiveThink(state), {
id: CALCULATE_ID,
kind: "calculate",
status: "live",
label: event.label || CONSULTATION_CHART_CALCULATION_LABEL,
});
return applyCalculateHints(created);
}
if (event.type === "activity") {
if (event.phase === "chart-calculation") {
@@ -88,18 +95,12 @@ export function reduceConsultationTimeline(
if (event.type === "tool.failed") {
return completeRow(state, CALCULATE_ID, CONSULTATION_DONE_CHART_LABEL);
}
if (event.type === "think.plan") {
let next = completeLiveThink(state);
for (const step of event.steps) {
next = upsertRow(next, {
id: `think-${step.id}`,
kind: "think",
status: "live",
label: consultationThinkTitle(step.title),
});
}
return next;
}
// The answer outline (think.plan / thinking.section) is what the answer
// will be written as, not a step the run has taken, so it never becomes a
// row: it was shown up front and ticked done as soon as the chart calculation
// started, before a word of the answer existed (BUG-1145). The headings the
// answer actually writes show as write rows below.
if (event.type === "think.plan") return state;
if (event.type === "think.step") {
const id = `think-${event.id}`;
const existing = state.rows.find((row) => row.id === id);
@@ -127,14 +128,7 @@ export function reduceConsultationTimeline(
heading: event.heading,
steps: event.steps,
};
const withCalc = enrichCalculate(completeRow(state, CALCULATE_ID, CONSULTATION_DONE_CHART_LABEL), section);
const withoutOpen = completeLiveThink(withCalc);
return upsertRow(withoutOpen, {
id: `think-${section.id}`,
kind: "think",
status: "live",
label: consultationThinkTitle(section.title),
});
return applyCalculateHints(addCalculateHints(state, section));
}
if (event.type === "thinking.delta") {
const think = lastRow(state.rows, (row) => row.kind === "think" && row.status === "live");
@@ -152,7 +146,11 @@ export function reduceConsultationTimeline(
});
}
if (event.type === "answer.delta") {
const next = { ...completeLiveThink(state), answer: event.replace ? event.text : `${state.answer}${event.text}` };
// Answer text only comes after the calculation, so the calculate row is
// done by now even if its own completion event has not been seen (BUG-1145:
// the planned sections used to close it).
const calculated = completeRow(completeLiveThink(state), CALCULATE_ID, CONSULTATION_DONE_CHART_LABEL);
const next = { ...calculated, answer: event.replace ? event.text : `${state.answer}${event.text}` };
return syncWriteRows(next, false);
}
if (event.type === "run.completed" || event.type === "run.failed") {
@@ -209,7 +207,7 @@ export function consultationTimelineFromSettled(input: {
for (const section of sections) {
events.push({ type: "thinking.section", ...section });
}
if (input.thinkingText?.trim() && sections.length === 0) {
if (input.thinkingText?.trim()) {
events.push({ type: "thinking.delta", text: input.thinkingText });
}
if (text) events.push({ type: "answer.delta", text });
@@ -228,14 +226,7 @@ export function consultationTimelineFromSettled(input: {
workflow: { route: "settled", status: "ready", preciseTiming: "blocked", missingLayers: [] },
},
});
let state = reduceConsultationTimelineEvents(events);
if (input.thinkingText?.trim() && sections.length > 0) {
const lastThink = [...state.rows].reverse().find((row) => row.kind === "think");
if (lastThink) {
state = upsertRow(state, { ...lastThink, thinkingText: input.thinkingText.slice(0, 4_000) });
}
}
return state.rows;
return reduceConsultationTimelineEvents(events).rows;
}
export function consultationThinkTitle(title: string): string {
@@ -390,13 +381,25 @@ function syncWriteRows(state: ConsultationTimelineState, settled: boolean): Cons
return next;
}
function enrichCalculate(
function addCalculateHints(
state: ConsultationTimelineState,
section: PublicThinkingSection,
): ConsultationTimelineState {
const hints = state.calculateHints;
return {
...state,
calculateHints: {
queries: uniqueLabels([...(hints?.queries ?? []), ...queriesFromSection(section)]).slice(0, 8),
sources: uniqueLabels([...(hints?.sources ?? []), ...sourcesFromSection(section)]).slice(0, 8),
},
};
}
function applyCalculateHints(state: ConsultationTimelineState): ConsultationTimelineState {
const calculate = state.rows.find((row) => row.id === CALCULATE_ID);
if (!calculate) return state;
const queries = uniqueLabels([...(calculate.queries ?? []), ...queriesFromSection(section)]).slice(0, 8);
const sources = uniqueLabels([...(calculate.sources ?? []), ...sourcesFromSection(section)]).slice(0, 8);
const hints = state.calculateHints;
if (!calculate || !hints) return state;
const queries = uniqueLabels([...(calculate.queries ?? []), ...hints.queries]).slice(0, 8);
const sources = uniqueLabels([...(calculate.sources ?? []), ...hints.sources]).slice(0, 8);
return upsertRow(state, { ...calculate, queries, sources });
}
@@ -62,3 +62,19 @@ export function orderProbesBySegment(input: {
return input.probes.map((probe, index) => ({ probe, index, gain: segmentInformationGain(probe, weights, input.segments, input.offsets) }))
.sort((a, b) => b.gain - a.gain || a.index - b.index).map((row) => row.probe);
}
/**
* Production default for segment-gain probe order (BUG-1144). 77-case persisted replay
* (2026-10-01, after the health-focus write fix BUG-1143): at 21 and 61 minutes it gained
* D9/D10 head hits with every truth segment kept (76/76). Before that fix the 61-minute cell lost
* truth segments, which is why the first cut stopped at 21. Wider windows were not replayed.
* "off" disables; "on" is the research envelope used by the persisted replay harness.
*/
export const SEGMENT_ORDER_DEFAULT_MAX_WINDOW = 61;
export const SEGMENT_ORDER_RESEARCH_MAX_WINDOW = 61;
export function segmentOrderEnabledFor(windowMinutes: number | null | undefined, flag: string | undefined): boolean {
if (flag === "off" || !windowMinutes || windowMinutes <= 0) return false;
const limit = flag === "on" ? SEGMENT_ORDER_RESEARCH_MAX_WINDOW : SEGMENT_ORDER_DEFAULT_MAX_WINDOW;
return windowMinutes <= limit;
}
@@ -821,14 +821,14 @@ function liveDistinguishProbe(
isValidDistinguishProbe({ ...probe, role: "distinguish" })
&& (schemaKey
? probe.semantic_key === schemaKey
: (!focus.targetDomain || probe.domain === focus.targetDomain))
: (!focus.targetDomain || sameCollectDomain(probe.domain, focus.targetDomain)))
);
const fromEvents = (eventProbes ?? []).find(matchEvent) ?? null;
if (fromEvents) return fromEvents;
const matchContrast = (probe: CandidateDiscriminatorProbe) => (
schemaKey
? probe.semanticKey === schemaKey
: (!focus.targetDomain || probe.domain === focus.targetDomain)
: (!focus.targetDomain || sameCollectDomain(probe.domain, focus.targetDomain))
);
const fromContrast = (contrastProbes ?? []).find(matchContrast);
if (!fromContrast) return null;
@@ -2580,7 +2580,7 @@ export function buildMethodFollowupPlan(input: {
: themeFromDomain ?? (focus.intent === "out_of_sample_check" ? "oos_blind" : "education_style")
) as MethodFollowup["ask_theme"];
const matchingVerify = (input.eventProbes ?? []).find((probe) => (
probe.domain === focus.targetDomain
sameCollectDomain(probe.domain, focus.targetDomain)
|| REVERSE_VERIFY_THEME[probe.domain as keyof typeof REVERSE_VERIFY_THEME] === keepAskTheme
));
const schema = focus.expectedAnswerSchema;
@@ -5,6 +5,7 @@
* write a fresh inference_state, but with an empty answer ledger.
*/
import { shouldForceMinuteAfterSubBlocks } from "./block-scan.ts";
import { segmentOrderEnabledFor } from "../core/segment-probe-order.ts";
import { scanCaseSegments } from "./segment-scan.ts";
import { targetChartsForDomain } from "../core/segment-summary.ts";
import { buildProductCaseInferenceState } from "./product-inference.ts";
@@ -450,7 +451,7 @@ export async function scoreAndPersistCurrentEvidence(input: {
: undefined;
const inference = buildProductCaseInferenceState({
segmentMinutes, segmentTargets, segmentScanComplete: Boolean(segmentMinutes),
segmentOrderEnabled: process.env.RECTIFICATION_SEGMENT_ORDER === "on",
segmentOrderEnabled: segmentOrderEnabledFor(segmentMinutes?.length, process.env.RECTIFICATION_SEGMENT_ORDER),
range: candidateRange,
candidates: score.candidates,
evidence: scorable,
@@ -272,6 +272,20 @@ export function persistableFocusDomain(domain: string | null | undefined): strin
return null;
}
/**
* The domain a focus row may store. Every intent goes through persistableFocusDomain:
* the table only accepts education/career/relationship/relocation/finance/health/family/other,
* while planner domains include health_pressure and occupation (BUG-672). A raw
* health_pressure on a distinguish focus failed the check constraint, the write was
* swallowed and the session ended early (BUG-1143).
*/
export function focusTargetDomain(followup: Pick<MethodFollowup, "intent" | "domain">): string | null {
return persistableFocusDomain(followup.domain)
?? (followup.intent === "collect_method_evidence"
? persistableFocusDomain(collectQuestionDomain(followup.domain))
: null);
}
function clampFocusTargetKind(kind: string | null | undefined): string | null {
const value = kind?.trim() || "";
if (!value) return null;
@@ -672,10 +686,7 @@ async function persistServerOwnedFocusCore(input: {
questionId,
intent: followup.intent,
targetEvidenceId: followup.date_reliability_evidence_id ?? null,
targetDomain: followup.intent === "collect_method_evidence"
? persistableFocusDomain(followup.domain)
?? persistableFocusDomain(collectQuestionDomain(followup.domain))
: followup.domain,
targetDomain: focusTargetDomain(followup),
targetKind: followup.intent === "collect_method_evidence"
? collectFocusTargetKind(followup)
: null,
@@ -9,7 +9,7 @@ import {
reduceConsultationTimelineEvents,
} from "../src/lib/consultation-run-timeline.ts";
test("career and wealth events append method, calculate 1/2, think, then the first write", () => {
test("career and wealth events append method, calculate 1/2, then the first write (no outline row)", () => {
const [question] = natalConsultationThinkingPlan({ domains: ["career", "wealth"] });
assert.ok(question);
const afterProgress = reduceConsultationTimelineEvents([
@@ -39,13 +39,15 @@ test("career and wealth events append method, calculate 1/2, think, then the fir
{ type: "answer.delta", text: "外松内紧。\n## 盘里支持这个判断的地方\n事业宫被土星压着。\n" },
], afterProgress);
assert.deepEqual(state.rows.map((row) => row.kind), ["method", "calculate", "think", "write"]);
// 原值: 行种类 ["method", "calculate", "think", "write"],第 3 行是计划节「先回答你问的这件事」且已勾
// 新值: 行种类 ["method", "calculate", "write"],计划节不成行;计算行由正文开始收口
// 原因: TASK-consult-timeline-no-plan-20261001(BUG-1145):回答提纲不是已走过的步骤,产品 10-01 选 B 去掉提纲行
assert.deepEqual(state.rows.map((row) => row.kind), ["method", "calculate", "write"]);
assert.equal(state.rows[1]?.status, "done");
assert.equal(state.rows[2]?.label, "先回答你问的这件事");
assert.equal(state.rows[2]?.status, "done");
assert.equal(state.rows[3]?.id, "write-盘里支持这个判断的地方");
assert.equal(state.rows[3]?.status, "live");
assert.match(state.rows[3]?.label ?? "", /正在写盘里支持这个判断的地方/);
assert.equal(state.rows.some((row) => row.label === "先回答你问的这件事"), false);
assert.equal(state.rows[2]?.id, "write-盘里支持这个判断的地方");
assert.equal(state.rows[2]?.status, "live");
assert.match(state.rows[2]?.label ?? "", /正在写盘里支持这个判断的地方/);
});
test("thinking after a tool starts a new think row instead of reopening the first", () => {
@@ -107,3 +109,62 @@ test("settled natal messages hydrate a calculate row from domain sections", () =
const calculate = rows.find((row) => row.kind === "calculate");
assert.ok(calculate?.queries?.some((item) => item.includes("事业") || item.includes("D10") || item.includes("财富")));
});
// BUG-1145 (TASK-consult-timeline-no-plan-20261001): the answer outline was
// shown as four think rows and ticked done the moment the calculation started.
test("the answer outline sent before the calculation adds no row and ticks nothing", () => {
const plan = natalConsultationThinkingPlan({ domains: ["parents"] });
const state = reduceConsultationTimelineEvents([
{ type: "run.started", runId: "run-1", requestId: "req-1" },
{ type: "skill.started", name: "jyotish-vedic-astrology" },
{ type: "skill.completed", name: "jyotish-vedic-astrology" },
{ type: "think.plan", steps: plan.map((section) => ({ id: section.id, title: section.title })) },
...plan.map((section) => ({ type: "thinking.section" as const, ...section })),
{ type: "tool.started", callId: "tool-1", tool: "run-jyotish-consultation", label: "正在计算本命盘…" },
]);
assert.deepEqual(state.rows.map((row) => row.kind), ["method", "calculate"]);
assert.deepEqual(state.rows.filter((row) => row.status === "done").map((row) => row.kind), ["method"]);
assert.equal(state.rows[1]?.status, "live");
for (const section of plan) {
assert.equal(state.rows.some((row) => row.label === section.title), false, section.title);
}
// The outline still details what the calculation looks at.
assert.ok((state.rows[1]?.queries?.length ?? 0) + (state.rows[1]?.sources?.length ?? 0) > 0);
});
test("a finished run leaves no outline row and nothing live", () => {
const plan = natalConsultationThinkingPlan({ domains: ["parents"] });
const state = reduceConsultationTimelineEvents([
{ type: "skill.started", name: "jyotish-vedic-astrology" },
{ type: "skill.completed", name: "jyotish-vedic-astrology" },
{ type: "think.plan", steps: plan.map((section) => ({ id: section.id, title: section.title })) },
...plan.map((section) => ({ type: "thinking.section" as const, ...section })),
{ type: "tool.started", callId: "tool-1", tool: "run-jyotish-consultation", label: "正在计算本命盘…" },
{ type: "answer.delta", text: "你和父母不疏远。\n## 你和妈妈\n她操心多。\n## 这周可以做的一件事\n- 先问她最近怎样。\n" },
{
type: "run.completed",
receipt: {
runId: "run-1",
runtime: "mastra-agentic",
skill: { name: "jyotish-vedic-astrology", loaded: true, referenceReads: 0, methodologySections: 0 },
steps: [],
workflow: { route: "parents", status: "ready", preciseTiming: "blocked", missingLayers: [] },
},
},
]);
assert.deepEqual(state.rows.map((row) => row.kind), ["method", "calculate", "write", "write"]);
assert.equal(state.rows.every((row) => row.status === "done"), true);
assert.equal(state.rows.some((row) => row.kind === "think"), false);
});
test("a stored natal answer rebuilds the same rows as live: no outline rows", () => {
const plan = natalConsultationThinkingPlan({ domains: ["parents"] });
const rows = consultationTimelineFromSettled({
text: "你和父母不疏远。\n## 你和妈妈\n她操心多。\n## 你和爸爸\n他话少。\n",
thinkingSections: plan,
});
assert.deepEqual(rows.map((row) => row.kind), ["method", "calculate", "write", "write"]);
assert.equal(rows.some((row) => row.kind === "think"), false);
assert.equal(rows.every((row) => row.status === "done"), true);
assert.ok((rows[1]?.queries?.length ?? 0) + (rows[1]?.sources?.length ?? 0) > 0);
});
@@ -0,0 +1,43 @@
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import test from "node:test";
import { sameCollectDomain } from "../src/lib/rectification-agentic/v9/domain-alias.ts";
import { focusTargetDomain } from "../src/lib/rectification-agentic/v9/server-focus.ts";
// BUG-1143: a distinguish focus wrote the planner domain verbatim. health_pressure is not in the
// focus table's target_domain check, the write failed inside a swallowed retry, and the session
// went straight to delivery although tap cards remained (BUG-672 alias family).
const migration = readFileSync(new URL("../supabase/migrations/20260814020000_rectification_v10_runtime.sql", import.meta.url), "utf8");
const allowedBlock = /target_domain is null or target_domain in \(([^)]*)\)/.exec(migration);
assert.ok(allowedBlock, "focus table target_domain check is readable");
const allowed = new Set([...allowedBlock[1]!.matchAll(/'([a-z_]+)'/g)].map((match) => match[1]!));
// Every domain the planner, engine catalog or collect pool can put on a followup.
const plannerDomains = [
"education", "relocation", "relationship", "career", "finance", "health_pressure", "family",
"appearance", "occupation", "health", "other", "horary", "unknown", "active_focus",
];
test("every followup domain maps to a value the focus table accepts, for every intent", () => {
for (const intent of ["distinguish_candidates", "collect_method_evidence", "reverse_verify", "out_of_sample_check"] as const) {
for (const domain of plannerDomains) {
const stored = focusTargetDomain({ intent, domain } as never);
assert.ok(stored === null || allowed.has(stored), `${intent}/${domain} -> ${stored}`);
}
}
});
test("a health distinguish focus stores health and still matches health_pressure probes", () => {
assert.equal(focusTargetDomain({ intent: "distinguish_candidates", domain: "health_pressure" } as never), "health");
assert.equal(focusTargetDomain({ intent: "distinguish_candidates", domain: "occupation" } as never), "other");
assert.equal(focusTargetDomain({ intent: "distinguish_candidates", domain: "relationship" } as never), "relationship");
assert.equal(sameCollectDomain("health_pressure", "health"), true);
});
test("focus writes and focus-to-probe matching never use a raw planner domain", () => {
const serverFocus = readFileSync(new URL("../src/lib/rectification-agentic/v9/server-focus.ts", import.meta.url), "utf8");
assert.doesNotMatch(serverFocus, /:\s*followup\.domain,\s*\n/, "no raw followup.domain passed as targetDomain");
assert.match(serverFocus, /targetDomain: focusTargetDomain\(followup\)/);
const followup = readFileSync(new URL("../src/lib/rectification-agentic/v9/method-followup.ts", import.meta.url), "utf8");
assert.doesNotMatch(followup, /probe\.domain === focus\.targetDomain/);
});
@@ -0,0 +1,21 @@
import assert from "node:assert/strict";
import test from "node:test";
import { segmentOrderEnabledFor, SEGMENT_ORDER_DEFAULT_MAX_WINDOW } from "../src/lib/rectification-agentic/core/segment-probe-order.ts";
// BUG-1144: segment-gain order is on by default only where the 77-case persisted replay lost no truth segment.
test("segment order is on by default for windows up to 61 minutes and off beyond", () => {
assert.equal(SEGMENT_ORDER_DEFAULT_MAX_WINDOW, 61);
assert.equal(segmentOrderEnabledFor(21, undefined), true);
assert.equal(segmentOrderEnabledFor(11, ""), true);
assert.equal(segmentOrderEnabledFor(61, undefined), true);
assert.equal(segmentOrderEnabledFor(62, undefined), false);
assert.equal(segmentOrderEnabledFor(121, undefined), false);
});
test("off disables everywhere; on keeps the research envelope; no scan means off", () => {
assert.equal(segmentOrderEnabledFor(21, "off"), false);
assert.equal(segmentOrderEnabledFor(61, "on"), true);
assert.equal(segmentOrderEnabledFor(121, "on"), false);
assert.equal(segmentOrderEnabledFor(undefined, undefined), false);
assert.equal(segmentOrderEnabledFor(0, "on"), false);
});