T2 of TASK-rectification-grounding-20260927 (product decision P1).
The rectification 「重试(重新生成)」 was a blind agent.generate rewrite that
persisted unchecked text (invented range / fit rate reproduced). Removed:
the regenerate route, regenerate-turn.ts, the regeneration agent and its
read-only tool set, the client regenerate action/state and the
regenerating/canRegenerate props. ChatMessageActions renders the regenerate
button only when onRegenerate is passed; ordinary consultation is unchanged.
The DB function regenerate_agentic_rectification_turn is kept (AGENTS §7.6,
retire in a later round). DESIGN.md / VOICE.md updated; source-contract
tests follow with 原值/新值/原因 notes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T1 of TASK-rectification-grounding-20260927 (recurrence of BUG-588).
- The attempt no longer streams range-changed / rescore-skipped /
compare-failed sentences; the finish whitelists and trims the model body,
then joins the server facts, and emits one final replace equal to the
persisted text.
- P3 whitelist (spoken-grounding.ts): a model sentence with a clock, clock
range or percentage that is not this turn's server fact is dropped whole;
the batch recap stands in when nothing is left.
- record-evidence-batch returns range_after_rescore (post-rescore
credible_range, representative minute, fit percent, delivers_range_this_turn);
the receipt fingerprint stays over the old shape.
- System prompt: range is said by the server; the delivery three sentences
only when the batch says this turn delivers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- BUG-1049 (recurrence of BUG-504): the opening stem carries six examples and
an example answer again and is server-owned on the zero-evidence opening;
the body is two plain sentences (no 大运/盘面/代表分钟/精确到秒, no year, no
question). Stem de-dup compares whole sentences / near-equality instead of a
12-char prefix, which had deleted the body's examples sentence since
aa7ccb30 (BUG-604) + dd8f35f7 (BUG-648).
- BUG-1050: plain step labels; a finished step label shows once and
「已完成 N 步」counts shown rows; failed rows read 「…未完成」 from the
in-progress wording.
- Skill 10.0.30 -> 10.0.31 (OpeningPolicy); 10.0.30 kept as deprecated.
- VOICE / DESIGN / CHANGELOG / BUG_HISTORY / PROGRESS / real-device checklist
and screenshots.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Unify minute and block cache identity, keep unverifiable historical results read-only across server tools and write entrypoints, and aggregate completed receipt sources chronologically through a compatible function migration.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Hide the button once both personality cards are answered, persist the second card in the same opt-in round, and freeze Skill 10.0.26.
Co-authored-by: Cursor <cursoragent@cursor.com>
Dated-pool empty now yields existence cards per remaining layer, leftover-candidate copy, and an optional D9/D10 tie-break on the range card. Skill 10.0.25.
Co-authored-by: Cursor <cursoragent@cursor.com>
Dated-choice exhaustion is not convergence. Refresh probes from remaining
active candidates, then ask a targeted collect, then deliver. Skill 10.0.24.
Co-authored-by: Cursor <cursoragent@cursor.com>
Occupation collect was wiping year-month into an unscored note, and idle gap copy never joined the evidence turn.
Co-authored-by: Cursor <cursoragent@cursor.com>
Stop domain-wheel collecting and age-band years in prompts. Ask until the training gate, then discriminate until convergence, then deliver a range plus a concrete follow-up. Reserve holdout only with four dated events.
Co-authored-by: Cursor <cursoragent@cursor.com>
Ask finance and health like other domains, put year-cued collect questions first, and let the classifier distinguish decline vs skip. Skill 10.0.22.
Co-authored-by: Cursor <cursoragent@cursor.com>
The staging gate failed because that case claimed 记下了 without a write tool, which the unwritten-evidence guard now rejects.
Co-authored-by: Cursor <cursoragent@cursor.com>
Ask dated dasha probes first; D9/D10 and nakshatra wait until that pool is empty, score at half weight, and never eliminate. Skill 10.0.21.
Co-authored-by: Cursor <cursoragent@cursor.com>
Hour windows no longer drop later signature clusters. Credible range uses cluster coverage, and a mid-session spoken birth window gets a fixed reply without calling the model.
Co-authored-by: Cursor <cursoragent@cursor.com>
Engine by_time ledgers now include the full candidate set so posterior columns can look up fit and windows. Delivery turns keep up to three sentences after idle persist; the card subtracts repeated traits and folds the board chart.
Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the minute-row delivery card with up to three compare columns so users can pick the time that fits, and apply the adult-year floor on inspect fallbacks.
Co-authored-by: Cursor <cursoragent@cursor.com>
Make the range card row-select, drop the composer 先这样 control and adopt status bar, keep delivery copy to three sentences with a folded verification report, and skip a second no-message agent run after a terminal delivery turn.
Co-authored-by: Cursor <cursoragent@cursor.com>
The verification template was still quoting the engine's pre-inference span and eliminated dasha tops, and D9/D10 signs were left for the model to guess from transition clocks.
Co-authored-by: Cursor <cursoragent@cursor.com>
Set-focus must not append a second narration, and a server-owned stem must not also appear as a question in the assistant body.
Co-authored-by: Cursor <cursoragent@cursor.com>
Stop/idle reuse the scored evidence path so existing answers stay on the range. Delivery uses one range card (Skill 10.0.15) instead of four minute cards.
Co-authored-by: Cursor <cursoragent@cursor.com>
Holdout stays closed until the training gate opens and skips declined domains.
Dated collect order is shared with exhaustion; spoken prompts must name the domain; progress copy uses collection_progress only.
BUG-527 through BUG-530.
Co-authored-by: Cursor <cursoragent@cursor.com>
Focuses now carry asked_turn_id so GET rebuilds stem and options on the
same turn. Agent writes spokenPrompt; the live question slot is gone.
Co-authored-by: Cursor <cursoragent@cursor.com>
The live question slot disappeared on refresh because it was never written to assistant_message. Attach the current collect_spoken prompt to that turn before finalize so chat history keeps it.
Co-authored-by: Cursor <cursoragent@cursor.com>
Persist the interview slot before the opening turn, keep the stem in the question slot, and hide that slot while a run is busy so the same ask does not appear twice.
Co-authored-by: Cursor <cursoragent@cursor.com>
Natal answers were opening on parameter tables, and rectification turns were one-sentence legal copy. Centralize user-facing strings, keep representative-minute and question-slot red lines, and stop duplicating the opening collect prompt as a second assistant message.
Co-authored-by: Cursor <cursoragent@cursor.com>
Silent unrenderable discriminators, a missing question-contract golden, and a always-on tool table were hiding fail-closed drops behind the prompt wall.
Co-authored-by: Cursor <cursoragent@cursor.com>
Family and occupation method layers were blocking discrimination even when
training events were complete and a discriminator probe existed, so the agent
only acknowledged evidence and stopped.
Co-authored-by: Cursor <cursoragent@cursor.com>
Three collected events with a reserved holdout were stalling because the discriminator door counted holdout. Public selection_allowed still had snapshot fallbacks, and health only proved the image SHA.
Co-authored-by: Cursor <cursoragent@cursor.com>
Exam-quality cards may still jump ahead of adoption, but career years stay on method rotation. Server stamps only period and family; spoken questions remain model-authored.
Co-authored-by: Cursor <cursoragent@cursor.com>
Recorded-year quality probes were spoken-only, so the interview had no
choice card. Compare also re-scored after batch until the 105s attempt
aborted the turn.
Co-authored-by: Cursor <cursoragent@cursor.com>
Conflict probes were jumping after one dated event, so the interview asked
another domain before method collection. Spoken replies now follow the
stamped choice prompt instead of a topic denylist.
Co-authored-by: Cursor <cursoragent@cursor.com>
Evidence writes now return the persisted open_question so the model asks that stem instead of a second education probe, and the jump-to-latest chip is centered again.
Co-authored-by: Cursor <cursoragent@cursor.com>
Coverage-complete ties stayed in discrimination because whole-window D9/D24 follow-ups were treated as probes, and restated dates inserted duplicate evidence. Skip encoded remaining layers, ask leftover D4 or offer a provisional range, and dedupe dated rows by kind and date.
Co-authored-by: Cursor <cursoragent@cursor.com>
Give the interview a hidden reasoning channel so planning leaves the spoken reply, publish terminal text-delta as-is, and stop regex or Case templates from replacing the model.
Co-authored-by: Cursor <cursoragent@cursor.com>
Aborting the turn on a second identical public tool-call failed staging
after evidence and compare had already succeeded. Skip the duplicate
receipt instead and let maxSteps bound real loops.
Co-authored-by: Cursor <cursoragent@cursor.com>
Mastra intermediate text-delta was published as answer.delta, then set-focus domain errors reset the attempt and replayed evidence. Publish only the terminal no-tool step, persist the next probe on the server, and ground batch quotes in the source turn.
Co-authored-by: Cursor <cursoragent@cursor.com>
Clicking A/B/C/D or stop must persist the answer, close the probe, and
update posteriors in one idempotent transaction instead of sending the
option text as a chat message.
Co-authored-by: Cursor <cursoragent@cursor.com>
Coverage complete only unlocks discrimination. A 34/33/33 window plus an
occupation note must ask a D9/D10 contrast probe instead of offering a
stale winner card.
Co-authored-by: Cursor <cursoragent@cursor.com>
Engine result rows stay immutable. Choice answers append transitions, and reads overlay the latest revision instead of patching the cached receipt.
Co-authored-by: Cursor <cursoragent@cursor.com>
Keep provider thinking on a separate channel so process talk is not billed as the spoken reply (BUG-359).
Co-authored-by: Cursor <cursoragent@cursor.com>