fix(consult): five-minute answer clock as a hang guard; general mode runs on it from the start (BUG-1142)
The 70s answer clock was sized for a non-reasoning writer; thinking tokens come out of the same clock and deepseek-v4-pro took 77s on a parents answer. Product 2026-10-01 chose five minutes. The general / no-birth-minute loop has no tools, so it now starts on the answer clock instead of the 110s tool clock. Tool phase and domain budget unchanged; maxDuration 240 -> 480. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
8912ba685a
commit
04c09cbf4d
@@ -115,14 +115,14 @@ import { generateSessionTitle, shouldGenerateSessionTitle } from "@/lib/session-
|
||||
import { z } from "zod";
|
||||
|
||||
export const runtime = "nodejs";
|
||||
export const maxDuration = 240;
|
||||
export const maxDuration = 480;
|
||||
|
||||
// The step budget, the wall-clock budget and the domain cap all bound this same
|
||||
// run, so they are declared as one group in @/mastra/consultation-tools with the
|
||||
// reasoning that ties them together. maxDuration above is the ceiling they must
|
||||
// stay under: the tool phase (AGENT_TIMEOUT_MS) plus the answer phase's own
|
||||
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), 180s, plus setup and settlement
|
||||
// (BUG-1051, BUG-1053). Self-hosted `node server.js` does not enforce it; it
|
||||
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), 410s, plus setup and settlement
|
||||
// (BUG-1051, BUG-1053, TASK-consult-answer-clock-20261001). Self-hosted `node server.js` does not enforce it; it
|
||||
// documents the ceiling.
|
||||
|
||||
const chatRequestMetadataSchema = z.object({
|
||||
@@ -1103,10 +1103,13 @@ export async function POST(request: Request) {
|
||||
// that deadline until the calculation result is in hand, then on the answer
|
||||
// clock (CONSULTATION_ANSWER_TIMEOUT_MS), because its next step writes the
|
||||
// answer. Continuation and answer retries share the same answer clock.
|
||||
// The general / no-birth-minute agent has no tools, so its loop is the
|
||||
// answer from the first step and runs on the answer clock from the start.
|
||||
const runClock = createConsultationRunClock({
|
||||
toolPhaseMs: AGENT_TIMEOUT_MS,
|
||||
answerMs: CONSULTATION_ANSWER_TIMEOUT_MS,
|
||||
answerReady: () => state.consultationToolCompleted,
|
||||
answerFromStart: usesPublicDailyGeneralAgent(consultationMode, generalDailyContext),
|
||||
});
|
||||
const agentAbortSignal = runClock.toolSignal;
|
||||
const streamOptions = {
|
||||
|
||||
@@ -78,19 +78,24 @@ export const AGENT_TIMEOUT_MS = 110_000;
|
||||
* covers the rest of the loop plus any length continuation or empty-answer
|
||||
* retry.
|
||||
*
|
||||
* 70s is measured, not guessed: staging writer calls on the default model
|
||||
* produced 560-1240 output tokens in 5.7-13.6s end to end (about 90-100 tok/s
|
||||
* including first-token latency, PROGRESS-report-writer-failure-20260902). 70s
|
||||
* therefore holds about 6,300-7,000 output tokens, three to four times a
|
||||
* typical four-heading answer; at half that throughput it still holds about
|
||||
* 3,000 tokens. The answer step keeps provider thinking on (it is the loop's
|
||||
* step), so thinking tokens now come out of the same 70s; the loop's step
|
||||
* after the tool result used to think and write inside the 110s as well, and
|
||||
* its text was thrown away. The calculation must finish inside the 110s tool
|
||||
* phase, so the worst case is 110 + 70 = 180s, the three minutes the product
|
||||
* accepted. The route's maxDuration must stay above the sum.
|
||||
* 300s is a hang guard, not a length limit (product decision 2026-10-01,
|
||||
* TASK-consult-answer-clock-20261001). The old 70s was sized for a
|
||||
* non-reasoning writer (about 90-100 tok/s, PROGRESS-report-writer-failure-20260902).
|
||||
* The answer step keeps provider thinking on, so thinking tokens come out of
|
||||
* this clock too, and a reasoning model spends most of its time there: in the
|
||||
* 10-01 real-model run (docs/testing/consult-plain-answer-20261001-model-runs.md)
|
||||
* a parents answer on deepseek-v4-pro took 77s end to end and would have been
|
||||
* cut, while deepseek-flash took 30-41s with about 85% of its tokens in
|
||||
* thinking. The answer shape no longer has a word cap either, so the clock
|
||||
* only has to stop a provider that hangs. A run that does hit it still ends
|
||||
* as answer_truncated and is not charged (BUG-1051).
|
||||
*
|
||||
* The calculation must finish inside the 110s tool phase, so the worst case is
|
||||
* 110 + 300 = 410s; the route's maxDuration must stay above the sum. The
|
||||
* domain budget below is derived from the tool phase and its own reserve, not
|
||||
* from this clock, so raising it does not change how many domains run.
|
||||
*/
|
||||
export const CONSULTATION_ANSWER_TIMEOUT_MS = 70_000;
|
||||
export const CONSULTATION_ANSWER_TIMEOUT_MS = 300_000;
|
||||
|
||||
export type ConsultationRunClock = Readonly<{
|
||||
/** The tool phase's deadline. Tools and precompute run under it, unchanged. */
|
||||
@@ -114,11 +119,17 @@ export type ConsultationRunClock = Readonly<{
|
||||
* where the calculation settles just before the tool-phase timer fires but
|
||||
* the stream consumer has not seen the result yet: the loop is then handed to
|
||||
* the answer clock instead of being cut.
|
||||
*
|
||||
* `answerFromStart` is for a loop with no calculation to wait for (the
|
||||
* general / no-birth-minute agent has no tools): its first step already writes
|
||||
* the answer, so the loop runs on the answer clock from the start instead of
|
||||
* being cut by the tool phase's timer (TASK-consult-answer-clock-20261001).
|
||||
*/
|
||||
export function createConsultationRunClock(options: {
|
||||
toolPhaseMs?: number;
|
||||
answerMs?: number;
|
||||
answerReady?: () => boolean;
|
||||
answerFromStart?: boolean;
|
||||
} = {}): ConsultationRunClock {
|
||||
const toolPhaseMs = options.toolPhaseMs ?? AGENT_TIMEOUT_MS;
|
||||
const answerMs = options.answerMs ?? CONSULTATION_ANSWER_TIMEOUT_MS;
|
||||
@@ -141,6 +152,7 @@ export function createConsultationRunClock(options: {
|
||||
}
|
||||
loop.abort(toolSignal.reason);
|
||||
}, { once: true });
|
||||
if (options.answerFromStart) answerSignal();
|
||||
return Object.freeze({ toolSignal, loopSignal: loop.signal, answerSignal });
|
||||
}
|
||||
export const CONSULTATION_NATAL_CALC_TOOL_ID = "run-jyotish-consultation";
|
||||
|
||||
Reference in New Issue
Block a user