fix(consult): five-minute answer clock as a hang guard; general mode runs on it from the start (BUG-1142)
Independent Staging Quality Gate / validate (push) Successful in 13m22s
Independent Staging Quality Gate / publish (push) Successful in 3m46s

The 70s answer clock was sized for a non-reasoning writer; thinking tokens
come out of the same clock and deepseek-v4-pro took 77s on a parents answer.
Product 2026-10-01 chose five minutes. The general / no-birth-minute loop has
no tools, so it now starts on the answer clock instead of the 110s tool clock.
Tool phase and domain budget unchanged; maxDuration 240 -> 480.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
This commit is contained in:
Jesse_Chen
2026-10-01 11:12:01 +08:00
co-authored by Claude Opus 5.5
parent 8912ba685a
commit 04c09cbf4d
11 changed files with 185 additions and 18 deletions
@@ -0,0 +1,35 @@
# PROGRESS · 普通对话答题时钟 5 分钟(2026-10-01)
## 改动
- `frontend/src/mastra/consultation-tools.ts`:`CONSULTATION_ANSWER_TIMEOUT_MS` 70_000 → 300_000,注释改写;`createConsultationRunClock` 新增 `answerFromStart`(创建时即启动答题时钟,工具钟到点因 `if (answer) return` 不再掐循环)。
- `frontend/src/app/api/consult/route.ts`:`maxDuration` 240 → 480,注释改 410 s;run clock 传 `answerFromStart: usesPublicDailyGeneralAgent(consultationMode, generalDailyContext)`。
- 申报时段路线:有窗口工具,拿到窗口计算结果时经 `onAnswerPhase` 切答题时钟(原有行为不变);预计算在 warmup 里用工具钟。
- 新测试 `frontend/tests/consult-answer-clock-20261001.test.ts`(4 条)。
## 改动的既有断言
| 文件 | 原值 | 新值 | 原因 |
|---|---|---|---|
| `tests/consult-answer-truncation-20260926.test.ts` | `CONSULTATION_ANSWER_TIMEOUT_MS = 70_000`;总和 ≤ 180_000 | `= 300_000`;总和 ≤ 410_000 | 产品 10-01 定 5 分钟防卡死 |
| `tests/consult-single-pass-answer-20260927.test.ts` | 总和 ≤ 180_000 | 总和 ≤ 410_000 | 同上 |
`tests/consult-evidence-lookup-20260927.test.ts` 的测试名与注释里仍写「70 s」,测试本身用自己的短时钟、不读常量,未改(描述性文字,记在此处)。
## 门禁(Node 22.14)
| 项 | 基线 `8912ba68` | 本分支 |
|---|---|---|
| tsc | — | 0 |
| lint | — | 0 error / 126 warning |
| npm test | 4,842 条,fail 24 | 4,846 条,fail 24,与基线逐条同名 |
| next build | — | `/` Static |
## 下游时限核对(T5,只报告)
- Caddy(`deploy/Caddyfile*`):反代未设置读写超时,默认不限响应时长。
- Next standalone / Node:无自定义 server;Node `requestTimeout` 300 s 管的是接收请求体,不管流式响应。
- Provider(undici 默认):`bodyTimeout` 300 s 是两块数据之间的空闲时间;DeepSeek 推理内容也是流式下发(10-01 实测推理 5,200 token 在 30 s 内持续到达),正常不会空闲到 300 s。
- 浏览器:无停滞检测;断线后服务端继续生成,客户端每 1.75 s 轮询状态。
- 预留扣点:`/api/consult/status` 15 分钟租约后才懒取消,> 410 s。
- 输出 token 上限 16,384 + 思考 8,192(`agent-generation-settings.ts`,与校正共用,未动)。
+1
View File
@@ -398,3 +398,4 @@
- 任务书里的行号会随代码漂移,定位以符号名为准。
- **实现合入 `staging` 的同一次推送里,必须同时把状态板那一行改掉。** 2026-09-16 的对账发现 10 份早已合入的单仍写着「待领取 / 待验收」,会导致重复派活。
| `TASK-consult-plain-answer-20261001.md` | `PROGRESS-consult-plain-answer-20261001.md` | 普通对话「还是废话」:09-17 四步开场形状(格局名 → 谁推谁修 → 扮演哪个象)逼出谜语,问父母时爸妈被揉成一段;产品授权推翻该形状,改成先答 + 按问题里的对象分段 + 人话自检 + 空宫不单独下结论 + 父母卡标 mother/father(BUG-1132~1134) | **已实现,待 Claude 验收**(fork 子代理直接执行);未推 staging、未部署;模型对比为环境缺口,真机清单 `docs/testing/consult-plain-answer-20261001.md` | 分支 `codex/consult-plain-answer-20261001` |
| `TASK-consult-answer-clock-20261001.md` | `PROGRESS-consult-answer-clock-20261001.md` | 普通对话答题时钟 70 s 容不下推理模型(v4-pro 77 s),无出生分钟路线只受 110 s 工具钟管;产品定 5 分钟防卡死(BUG-1142) | **已实现,待 Claude 验收**(直接执行) | 分支 `codex/consult-answer-clock-20261001` |
@@ -0,0 +1,31 @@
# TASK · 普通对话答题时钟 70 s → 5 分钟(2026-10-01)
- 基线:`origin/staging` `8912ba68`
- 分支 / worktree:`codex/consult-answer-clock-20261001` / `.worktrees/consult-answer-clock-20261001`
- 模式:直接执行(产品 10-01 选择「直接执行」),Claude 派子代理实现、独立验收
- BUG:BUG-1142
## 事故实证
10-01 真实模型对比(`docs/testing/consult-plain-answer-20261001-model-runs.md`):`deepseek-v4-pro` 父母题 77 s 写完,超过 `CONSULTATION_ANSWER_TIMEOUT_MS`(`frontend/src/mastra/consultation-tools.ts`)的 70 s,线上会以 `answer_truncated` 截断。时限全链路排查:从浏览器、Caddy、Node 到 provider,能在写作中截断回答的只有本应用的两只钟(工具 110 s、答题 70 s),其余层都 ≥ 240 s 或不按总时长计。另发现无出生分钟路线(`usesPublicDailyGeneralAgent`)不切答题时钟,整段循环受 110 s 工具钟管。
## 决策记录
- 产品 2026-10-01:答题时长「可以延长、别定 70 秒这个上限」;在「3 / 5 / 10 分钟」中选 **5 分钟**,定位为防卡死,不是长度限制。推翻 BUG-1051 / 1053 记录中「110 + 70 = 180 s,产品接受三分钟」的口径,新上限 110 + 300 = 410 s。
- 无出生分钟路线一并走答题时钟。
- 工具阶段 110 s 与领域预算(2 个领域)不变。
## 硬红线
不动生时校正(`rectification-*`、`RECTIFICATION_RUN_BUDGET_MS`)、`agent-generation-settings.ts`、Caddy / deploy / workflow;截断仍记 `answer_truncated` 且不扣点。
## 任务与验收
| # | 内容 | 验收 |
|---|---|---|
| T1 | 答题时钟 300_000,注释写清原因 | 单测锁 300 s,工具钟仍 110 s |
| T2 | 领域预算不随之变化 | 单测锁 65 s / 2 个领域,公式不读答题时钟 |
| T3 | 无出生分钟路线从第一步起走答题时钟(`answerFromStart`) | 单测:工具钟到点不掐循环,答题钟到点仍掐;route 接线断言 |
| T4 | `maxDuration` 240 → 480 | 单测:110 + 300 + 60 s 以内 |
| T5 | 下游时限核对(只报告) | 见 PROGRESS |
| T6 | 记录:BUG-1142、CHANGELOG、README、真机清单加一步 | — |