fix(consult): five-minute answer clock as a hang guard; general mode runs on it from the start (BUG-1142)
Independent Staging Quality Gate / validate (push) Successful in 13m22s
Independent Staging Quality Gate / publish (push) Successful in 3m46s

The 70s answer clock was sized for a non-reasoning writer; thinking tokens
come out of the same clock and deepseek-v4-pro took 77s on a parents answer.
Product 2026-10-01 chose five minutes. The general / no-birth-minute loop has
no tools, so it now starts on the answer clock instead of the 110s tool clock.
Tool phase and domain budget unchanged; maxDuration 240 -> 480.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
This commit is contained in:
Jesse_Chen
2026-10-01 11:12:01 +08:00
co-authored by Claude Opus 5.5
parent 8912ba685a
commit 04c09cbf4d
11 changed files with 185 additions and 18 deletions
+6 -3
View File
@@ -115,14 +115,14 @@ import { generateSessionTitle, shouldGenerateSessionTitle } from "@/lib/session-
import { z } from "zod";
export const runtime = "nodejs";
export const maxDuration = 240;
export const maxDuration = 480;
// The step budget, the wall-clock budget and the domain cap all bound this same
// run, so they are declared as one group in @/mastra/consultation-tools with the
// reasoning that ties them together. maxDuration above is the ceiling they must
// stay under: the tool phase (AGENT_TIMEOUT_MS) plus the answer phase's own
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), 180s, plus setup and settlement
// (BUG-1051, BUG-1053). Self-hosted `node server.js` does not enforce it; it
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), 410s, plus setup and settlement
// (BUG-1051, BUG-1053, TASK-consult-answer-clock-20261001). Self-hosted `node server.js` does not enforce it; it
// documents the ceiling.
const chatRequestMetadataSchema = z.object({
@@ -1103,10 +1103,13 @@ export async function POST(request: Request) {
// that deadline until the calculation result is in hand, then on the answer
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), because its next step writes the
// answer. Continuation and answer retries share the same answer clock.
// The general / no-birth-minute agent has no tools, so its loop is the
// answer from the first step and runs on the answer clock from the start.
const runClock = createConsultationRunClock({
toolPhaseMs: AGENT_TIMEOUT_MS,
answerMs: CONSULTATION_ANSWER_TIMEOUT_MS,
answerReady: () => state.consultationToolCompleted,
answerFromStart: usesPublicDailyGeneralAgent(consultationMode, generalDailyContext),
});
const agentAbortSignal = runClock.toolSignal;
const streamOptions = {
+24 -12
View File
@@ -78,19 +78,24 @@ export const AGENT_TIMEOUT_MS = 110_000;
* covers the rest of the loop plus any length continuation or empty-answer
* retry.
*
* 70s is measured, not guessed: staging writer calls on the default model
* produced 560-1240 output tokens in 5.7-13.6s end to end (about 90-100 tok/s
* including first-token latency, PROGRESS-report-writer-failure-20260902). 70s
* therefore holds about 6,300-7,000 output tokens, three to four times a
* typical four-heading answer; at half that throughput it still holds about
* 3,000 tokens. The answer step keeps provider thinking on (it is the loop's
* step), so thinking tokens now come out of the same 70s; the loop's step
* after the tool result used to think and write inside the 110s as well, and
* its text was thrown away. The calculation must finish inside the 110s tool
* phase, so the worst case is 110 + 70 = 180s, the three minutes the product
* accepted. The route's maxDuration must stay above the sum.
* 300s is a hang guard, not a length limit (product decision 2026-10-01,
* TASK-consult-answer-clock-20261001). The old 70s was sized for a
* non-reasoning writer (about 90-100 tok/s, PROGRESS-report-writer-failure-20260902).
* The answer step keeps provider thinking on, so thinking tokens come out of
* this clock too, and a reasoning model spends most of its time there: in the
* 10-01 real-model run (docs/testing/consult-plain-answer-20261001-model-runs.md)
* a parents answer on deepseek-v4-pro took 77s end to end and would have been
* cut, while deepseek-flash took 30-41s with about 85% of its tokens in
* thinking. The answer shape no longer has a word cap either, so the clock
* only has to stop a provider that hangs. A run that does hit it still ends
* as answer_truncated and is not charged (BUG-1051).
*
* The calculation must finish inside the 110s tool phase, so the worst case is
* 110 + 300 = 410s; the route's maxDuration must stay above the sum. The
* domain budget below is derived from the tool phase and its own reserve, not
* from this clock, so raising it does not change how many domains run.
*/
export const CONSULTATION_ANSWER_TIMEOUT_MS = 70_000;
export const CONSULTATION_ANSWER_TIMEOUT_MS = 300_000;
export type ConsultationRunClock = Readonly<{
/** The tool phase's deadline. Tools and precompute run under it, unchanged. */
@@ -114,11 +119,17 @@ export type ConsultationRunClock = Readonly<{
* where the calculation settles just before the tool-phase timer fires but
* the stream consumer has not seen the result yet: the loop is then handed to
* the answer clock instead of being cut.
*
* `answerFromStart` is for a loop with no calculation to wait for (the
* general / no-birth-minute agent has no tools): its first step already writes
* the answer, so the loop runs on the answer clock from the start instead of
* being cut by the tool phase's timer (TASK-consult-answer-clock-20261001).
*/
export function createConsultationRunClock(options: {
toolPhaseMs?: number;
answerMs?: number;
answerReady?: () => boolean;
answerFromStart?: boolean;
} = {}): ConsultationRunClock {
const toolPhaseMs = options.toolPhaseMs ?? AGENT_TIMEOUT_MS;
const answerMs = options.answerMs ?? CONSULTATION_ANSWER_TIMEOUT_MS;
@@ -141,6 +152,7 @@ export function createConsultationRunClock(options: {
}
loop.abort(toolSignal.reason);
}, { once: true });
if (options.answerFromStart) answerSignal();
return Object.freeze({ toolSignal, loopSignal: loop.signal, answerSignal });
}
export const CONSULTATION_NATAL_CALC_TOOL_ID = "run-jyotish-consultation";
@@ -0,0 +1,55 @@
// TASK-consult-answer-clock-20261001: the answer clock is a five-minute hang
// guard, the general / no-birth-minute loop runs on it from the start, and the
// domain budget does not move with it.
import assert from "node:assert/strict";
import { readFileSync } from "node:fs";
import test from "node:test";
import {
AGENT_TIMEOUT_MS,
CONSULTATION_ANSWER_TIMEOUT_MS,
CONSULTATION_DOMAIN_WALL_CLOCK_MS,
MAX_CONSULTATION_DOMAINS,
createConsultationRunClock,
} from "../src/mastra/consultation-tools.ts";
const route = readFileSync(new URL("../src/app/api/consult/route.ts", import.meta.url), "utf8");
const tools = readFileSync(new URL("../src/mastra/consultation-tools.ts", import.meta.url), "utf8");
test("the answer clock is five minutes and the tool phase stays 110s", () => {
assert.equal(CONSULTATION_ANSWER_TIMEOUT_MS, 300_000);
assert.equal(AGENT_TIMEOUT_MS, 110_000);
const maxDuration = Number(route.match(/export const maxDuration = (\d+);/)?.[1]);
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS + 60_000 <= maxDuration * 1000, "setup and settlement still fit");
});
test("the domain budget is unchanged by the answer clock", () => {
// Same numbers as before the change: 110s tool phase minus a 45s reserve,
// 31s per domain, so two domains run.
assert.equal(CONSULTATION_DOMAIN_WALL_CLOCK_MS, 65_000);
assert.equal(MAX_CONSULTATION_DOMAINS, 2);
const formula = tools.match(/export const CONSULTATION_DOMAIN_WALL_CLOCK_MS = [^;]+;/)?.[0] ?? "";
assert.match(formula, /AGENT_TIMEOUT_MS - CONSULTATION_ANSWER_RESERVE_MS/);
assert.doesNotMatch(formula, /CONSULTATION_ANSWER_TIMEOUT_MS/);
});
test("a loop with no calculation runs on the answer clock from the start", async () => {
const clock = createConsultationRunClock({ toolPhaseMs: 30, answerMs: 200, answerFromStart: true });
await new Promise((resolve) => setTimeout(resolve, 60));
assert.equal(clock.toolSignal.aborted, true, "the tool phase's timer still fires");
assert.equal(clock.loopSignal.aborted, false, "but it does not cut a loop that is already writing the answer");
await new Promise((resolve) => setTimeout(resolve, 220));
assert.equal(clock.loopSignal.aborted, true, "the answer clock still bounds it");
assert.equal(clock.answerSignal(), clock.answerSignal(), "retries and continuations share the same clock");
});
test("the consult route starts the general / no-birth-minute loop on the answer clock", () => {
assert.match(
route,
/const runClock = createConsultationRunClock\(\{[\s\S]*?answerFromStart: usesPublicDailyGeneralAgent\(consultationMode, generalDailyContext\),\s*\}\);/,
);
// That branch still streams on the shared loop signal and keeps its traces.
const general = route.match(/if \(usesPublicDailyGeneralAgent\(consultationMode, generalDailyContext\)\) \{[\s\S]*?\n {4}\}\n/)?.[0] ?? "";
assert.match(general, /await streamWithOverflowRetry\(agent\)/);
assert.match(general, /onError: \(error\) => settleRun\(\s*cancel,/);
});
@@ -356,7 +356,11 @@ test("the consult route gives the answer phase its own clock inside maxDuration"
// 原因: BUG-1053 删除 compose 流;「写回答有自己的时钟、不与工具阶段共用闸刀」这一性质不变
const route = readFileSync(new URL("../src/app/api/consult/route.ts", import.meta.url), "utf8");
const tools = readFileSync(new URL("../src/mastra/consultation-tools.ts", import.meta.url), "utf8");
assert.match(tools, /export const CONSULTATION_ANSWER_TIMEOUT_MS = 70_000;/);
// 原值: CONSULTATION_ANSWER_TIMEOUT_MS = 70_000;总上限 AGENT_TIMEOUT_MS + 答题时钟 <= 180_000
// 新值: CONSULTATION_ANSWER_TIMEOUT_MS = 300_000;总上限 <= 410_000(110s 工具 + 300s 答题)
// 原因: 产品 2026-10-01 决定答题时钟只防卡死、不限长度(TASK-consult-answer-clock-20261001);
// 推理模型 deepseek-v4-pro 父母题 77s 写完,70s 会截断
assert.match(tools, /export const CONSULTATION_ANSWER_TIMEOUT_MS = 300_000;/);
assert.match(route, /const runClock = createConsultationRunClock\(\{\s+toolPhaseMs: AGENT_TIMEOUT_MS,\s+answerMs: CONSULTATION_ANSWER_TIMEOUT_MS,/);
assert.match(route, /const answerPhaseSignal = runClock\.answerSignal;/);
assert.equal(route.match(/onAnswerPhase: startAnswerPhase,/g)?.length, 2, "natal and window hand the loop over");
@@ -371,5 +375,5 @@ test("the consult route gives the answer phase its own clock inside maxDuration"
assert.match(route, /const agentAbortSignal = runClock\.toolSignal;/);
const maxDuration = Number(route.match(/export const maxDuration = (\d+);/)?.[1]);
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS < maxDuration * 1000);
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS <= 180_000, "product accepted about three minutes");
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS <= 410_000, "product accepted a five-minute answer guard (2026-10-01)");
});
@@ -446,5 +446,8 @@ test("the consult route has no separate compose stream and wires the single-pass
assert.match(route, /const answerPhaseSignal = runClock\.answerSignal;/);
const maxDuration = Number(route.match(/export const maxDuration = (\d+);/)?.[1]);
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS < maxDuration * 1000);
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS <= 180_000, "product accepted about three minutes");
// 原值: AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS <= 180_000(110s + 70s)
// 新值: <= 410_000(110s + 300s)
// 原因: 产品 2026-10-01 把答题时钟改成 5 分钟防卡死(TASK-consult-answer-clock-20261001)
assert.ok(AGENT_TIMEOUT_MS + CONSULTATION_ANSWER_TIMEOUT_MS <= 410_000, "product accepted a five-minute answer guard (2026-10-01)");
});