About 30 fixed rectification lines lose the receipt tone: no 「已记录。」
openings (收到 when the range holds, 好 when it moves), 麻烦 instead of 请,
时间 instead of 候选 in what the user reads, and no 「新建对话」 after
adoption. Numbers, clocks, ranges and the honesty boundaries are unchanged.
CHOICE_ACK_RE also recognises the new acknowledgements, so a model sentence
that repeats one is still dropped; the delivery guard matches the delivery
sentence by containment and does not depend on the opening.
Full suite 4973 / fail 24, identical to 61d6f248; build keeps / Static.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Both rectification prompts now share RECTIFICATION_VOICE_HEAD, the same warm
voice as the consultation answer (no mind-reading here: these turns collect
facts). Opening and read-only turns used to send the whole bound Skill (11,812
characters for 10.0.33); they now send the sections they act on, plus the
product-identity section when the bound Skill has it. 9.0.0 still goes out
whole (BUG-621). Rule bodies, scoring, cards and adoption are unchanged.
Opening input 7,433 -> 3,892 tokens in a DeepSeek replay. Full suite 4970 /
fail 24, identical to c76fdda7; build keeps / Static.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
The answered choice persists the acknowledgement, a blank line, and the
rewritten delivery. The exit fill-in required equality and wrote the
deterministic delivery again. Check containment instead, and pin it on the
real answer path. The suspected widen staleness regression did not hold;
keep a guard test.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
Accept and GET now decide with the same birth snapshot as the answer, so an under-age probe cannot flip can_adopt. The exit skips a second write when this reply already delivered. The card stays without an adopt button when adoption is not allowed, and a refusal shows the Chinese sentence.
Product decision 2026-10-02 (TASK-upstream-sync5 R1): a time range is offered
only with at least 4 dated, primary-scoreable events covering 3 domains,
counted on all of them (training + reserved holdout). Was 3 training events /
2 domains in three TS copies and the Python acceptance gate while the policy
file already said 4/3.
- One definition: references/rectification_policy.v1.json
(minConfirmationEvents / minConfirmationDomains). TS core/types MIN_DATED_*,
rectification-decision MIN_STANDALONE_*, evidence-model MIN_ACCEPTANCE_*,
the convergence evaluator and the post-inference trainingGateOpen all read
it; Python decision_policy MIN_ACCEPTANCE_* alias MIN_CONFIRMATION_*.
- Python receipt counts all scoreable events / domains for event_quality and
domain_diversity; decision policy identity v3 -> v4 (candidate UUIDs carry
it). Candidate scores unchanged (77 v5 cases A/B identical), so the
algorithm stays rectification-v5-matrix-scoring-10.
- Memoization golden v3 written by write_golden; v2 frozen by sha256 with a
test that its scores equal v3 and only the receipt policy moved.
- Collect gap copy names the exact gap ("再来两件……其中至少一件不是……")
instead of always "再来一件"; VOICE.md updated. Legacy life-events form copy
4/3 as well.
- 30 frontend test files, 4 Python tests: fixtures extended to the same
scenario at 4/3, or assertions changed with 原值/新值/原因 notes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
Functional roles feed the *_functional_*_auxiliary rules, so 57782aea changes
candidate scores for identical input (memoization fixture 12:00: 8.6274 ->
8.4977). Per the "scoring semantics change => bump ALGORITHM_VERSION"
precedent (scoring-7 -> 8 -> 9), the identity moves to scoring-10; policy v3,
input contract v5 and Skill versions are unchanged, history is not relabeled.
Five frontend sites and one SQL guard tested `=== "...scoring-9"` for the
dated candidate-window contract; they now use isDatedScoringAlgorithmVersion /
a generation regex (>= 9). Migration 20261002010000 only recreates
validate_dated_rectification_candidate (one-line guard change).
Memoization golden v2 written by the test's own write_golden; v1 (scoring-8)
frozen by sha256. Real-engine scoring-10 cross-midnight golden added. Research
records re-frozen per ERR-110 (label functional_v2_2026_10_02) and
scripts/functional_benefics.py added to the frozen production identity
(ERR-114: 57782aea changed scores without tripping it).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
The 77-case persisted replay after the BUG-1143 fix keeps every truth segment
at both 21 and 61 minutes with segment order on, so the default envelope moves
from 21 to 61 minutes. Bug numbers follow staging (BUG-1142 was taken).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
BUG-1142: distinguish focuses wrote the planner domain verbatim, and
health_pressure is not in the focus table's target_domain check. The write
failed inside a swallowed retry and the session went straight to delivery
although tap cards remained. Every focus write now goes through
focusTargetDomain (persistableFocusDomain, BUG-672), and focus-to-probe
matching uses sameCollectDomain.
BUG-1143: segmentOrderEnabledFor turns segment-gain probe order on by default
when the scanned window is at most 21 minutes, where the 77-case persisted
replay gained head hits with no lost truth segment. RECTIFICATION_SEGMENT_ORDER
=off disables it; =on keeps the 61-minute research envelope.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- Evidence turns whose batch accepted items: the body is the server recap
built from those items (two or fewer listed inline, more as a count);
the model no longer restates them (both prompt sets updated).
- The case snapshot carries each assistant turn's recorded evidence from
the ledger (rejected/superseded excluded); the message shows it in a
collapsed <details> list when more than two.
- Chinese date labels read straight into the event phrase (2016年入学);
ISO labels keep their space.
- Assertions updated with original/new/reason notes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- GET attach: a live hanging focus replaces a superseded-unanswered focus
linked to the last assistant turn instead of falling to a standalone copy.
- Client placement: a dead question on the latest message is replaceable,
and the stem the body still ends with is stripped (stripQuestionSentences).
- Chat view: a dead question no longer turns the gap into 「没有拿到下一个问题」
while the case has a different live question.
- Message entry: a superseded or replaced unanswered choice draws nothing
(no stem without options); the current question keeps its stem while busy.
- Copy text skips a dead question.
- Three DOM tests reproduce 「stem, stem」 and 「dead stem + unavailable」, red then green.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Review-only snapshot for BUG-1115 through BUG-1117; not merge-ready. New opening and append-turn PostgreSQL permission failures remain blocked. Persisted joint replay has zero completed questions; segment ordering remains off by default. The existing offline replay JSON is retained stale and unchanged after a denied overwrite, including its CRLF line endings. Browser/provider validation and final serial gates remain pending. No deployment, role permission changes, or staging/main push.
Co-Authored-By: Claude Code <noreply@anthropic.com>
The China city table loads on demand; old code-only profiles load it before
the people list settles, so they resolve exactly as before. The gate keeps
the six holdout fields it reads, pinned to the JSON by a test. The
onboarding shell is preloaded before reveal only when the profile is
incomplete.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
Once the discriminator training gate is open, only choice cards are asked
and the range card goes out when they are exhausted; targeted lines, their
re-ask and guided windows no longer hold the card or invite more events.
Delivery body says how many choice questions were used instead of the event
fit percent; narration names an excluded cluster instead of "range
unchanged"; a delivered turn no longer carries a collect question.
Offline replay (v4, 3 radii x 2 directions): truth in range 20/20 in every
cell; guided-window injections give the same width in truth and opposite
directions, so red line 1 was revised by product to truth-in-range only.
Skill 10.0.31 -> 10.0.32 (10.0.31 kept as deprecated for pinned cases).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T6 of TASK-rectification-grounding-20260927.
- skill-slice.ts: per-action slice rule as a code constant, keyed on heading
titles (evidence: 「ConversationFocus 与意图承接」「批量证据与日期真实性」, i.e.
§5/§7 of the 10.0.x layout); other actions and any bound Skill without those
headings (9.0.0) get the whole body, so historical Cases still run with their
exact bound Skill (BUG-621).
- The rectification Agent declares providesSkillDiscovery "on-demand"
(rectificationSkillBoundProcessor): no <available_skills> block with a temp
path and no "call the skill tool" system message; getSkill still loads the
bound package.
- Measured on a real Agent + recording model (public AA case, estimate = CJK
chars + other chars / 4): fixed overhead per call 12,002 → 6,471 tokens
(step 0: 8,440 → 2,909). Skill text unchanged; no version bump.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T3 of TASK-rectification-grounding-20260927 (red line 3).
The turn-decision read-case is the model's only conversation memory; with
about 7+ candidates (9 in the public AA case) it exceeded 6 KB and cleared
recent_turns and relevant_evidence_summary first. Candidates in the
model-visible inference now carry time / score / status / cluster_range only
(the last round names candidates by time), candidate_summary.candidates (a
duplicate) is gone, and over budget the order is: drop cluster ranges → keep
the best six candidates → shorten turns/evidence → clear them. Stored
inference and receipts are untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T2 of TASK-rectification-grounding-20260927 (product decision P1).
The rectification 「重试(重新生成)」 was a blind agent.generate rewrite that
persisted unchecked text (invented range / fit rate reproduced). Removed:
the regenerate route, regenerate-turn.ts, the regeneration agent and its
read-only tool set, the client regenerate action/state and the
regenerating/canRegenerate props. ChatMessageActions renders the regenerate
button only when onRegenerate is passed; ordinary consultation is unchanged.
The DB function regenerate_agentic_rectification_turn is kept (AGENTS §7.6,
retire in a later round). DESIGN.md / VOICE.md updated; source-contract
tests follow with 原值/新值/原因 notes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T4 of TASK-rectification-grounding-20260927 (BUG-588 / BUG-569 family).
The deferFollowup path persisted the choice before the agent turn, so the
agent turn's prepare read the post-answer range and the deferred choice
narration was never persisted: nobody said the range moved. The typed-message
preflight now hands the pre-answer credible range to the agent turn
(rangeBeforeTurn); the turn's one server range sentence covers the answer and
the dated event. Route-level regression with the real handler.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T1 of TASK-rectification-grounding-20260927 (recurrence of BUG-588).
- The attempt no longer streams range-changed / rescore-skipped /
compare-failed sentences; the finish whitelists and trims the model body,
then joins the server facts, and emits one final replace equal to the
persisted text.
- P3 whitelist (spoken-grounding.ts): a model sentence with a clock, clock
range or percentage that is not this turn's server fact is dropped whole;
the batch recap stands in when nothing is left.
- record-evidence-batch returns range_after_rescore (post-rescore
credible_range, representative minute, fit percent, delivers_range_this_turn);
the receipt fingerprint stays over the old shape.
- System prompt: range is said by the server; the delivery three sentences
only when the batch says this turn delivers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- BUG-1049 (recurrence of BUG-504): the opening stem carries six examples and
an example answer again and is server-owned on the zero-evidence opening;
the body is two plain sentences (no 大运/盘面/代表分钟/精确到秒, no year, no
question). Stem de-dup compares whole sentences / near-equality instead of a
12-char prefix, which had deleted the body's examples sentence since
aa7ccb30 (BUG-604) + dd8f35f7 (BUG-648).
- BUG-1050: plain step labels; a finished step label shows once and
「已完成 N 步」counts shown rows; failed rows read 「…未完成」 from the
in-progress wording.
- Skill 10.0.30 -> 10.0.31 (OpeningPolicy); 10.0.30 kept as deprecated.
- VOICE / DESIGN / CHANGELOG / BUG_HISTORY / PROGRESS / real-device checklist
and screenshots.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Move-only (TASK-rectification-code-split-20260926 T3/T4). No behavior,
receipt, billing, Skill or scoring change.
- agent-run.ts: 1400 -> 135 lines; runV9AgentTurn 1046 -> 25 lines, no
nested functions. It keeps the public types, the budget constants and the
phase order: agent-run-prepare.ts (reliability, session, year re-ask, Skill
identity, delivered guard, reserve, turn row / replay),
agent-run-retry.ts (attempt loop), agent-run-attempt.ts (the old nested
streamAttempt), agent-run-finish.ts (billing settle, receipts, interview,
finalize), agent-run-support.ts (outcome types, retry classification, turn
receipt writers), agent-run-messages.ts (buildAgentMessages /
buildOpeningBrief, re-exported).
- One token changed with the move: the attempt passes its own
`previousErrorCode` argument to buildAgentMessages instead of reading the
enclosing `lastAttemptError`; the loop passes that same value (clears the
old unused-parameter warning).
- Tests: whole-source contracts read tests/rectification-agent-run-surface.ts;
behavioral runV9AgentTurn tests unchanged. agent-run growth caps added.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Move-only (TASK-rectification-code-split-20260926 T2/T4). No behavior, API,
copy, DB, Skill, billing or scoring change.
- route.ts: 1077 -> 152 lines; POST 927 -> 114. It builds the Supabase
clients (so the setup-failure mapping stays here) and assembles:
agent-route-request.ts (auth / product / schema / flag / Case-Session
binding), agent-route-typed-message.ts (declared-window reply, typed-answer
preflight, unfocused classification), agent-route-structured-choice.ts,
agent-route-agent-turn.ts (opening / read-only / typed agent stream and its
exit gate), agent-route-billing.ts, agent-route-support.ts (schema, one-shot
NDJSON reply, context types).
- The four mutable preflight lets (expectedWrite, collectIntent,
writeClassified, classifierDiagnostic) that crossed branches are one
turnState object; values and flow unchanged.
- Tests: whole-source contracts read tests/rectification-agent-route-surface.ts;
the typed fast-path slice is rebuilt from the new files; billing,
declared-window and structured-choice slices now call the handlers
(billing in a child process because feature-pricing imports server-only),
each with 原值/新值/原因. Route growth caps added.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
One row per rectification Case, written once when the range card is first
delivered (GET /api/rectification/cases/[caseId], fire-and-forget after the
response is built). Numbers and closed enums only: no user / case / session
id, birth data, names, text or timestamps finer than the ISO week. Dedupe via
a separate case_id ledger that cascades with the Case (and account deletion).
Migration 20260926010000 is additive: two RLS tables with no runtime table
grants, SECURITY DEFINER write (service_role), purge (service_role) and
aggregate-only summary (admin_runtime) functions; 180-day retention.
Admin: 「校正统计」 page + GET /api/admin/rectification-telemetry
(admin.customers.read), aggregates only, no per-row view or export.
TASK-rectification-telemetry-20260926. test:db not run locally (no Docker).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- D2: each intent-classifier attempt is capped at 10 s (same session model,
thinking untouched); a hang takes the existing retry -> classifier_unavailable
path. Success is timed too (RectificationClassifierDiagnostic).
- D3: a typed message builds the NDJSON stream first; the first line is
turn.progress "received", then the classifier and deterministic replies run
inside the stream. Preflight rejections become turn.rejected (old status,
code, message) and the client handles them like the old HTTP rejection.
Stage lines 收到,正在对照你的档案… / 正在记下这件事… / 正在重新对照盘面… /
正在准备下一个问题… are driven by existing tool events and engine calls, are
transient (live row only) and never persisted. VOICE / DESIGN updated.
- D4: RectificationRunDiagnostic records per-step start/end, provider token
usage incl. reasoning tokens, classifier timing and per-engine-call
durations (AsyncLocalStorage scope per turn); RectificationTurnDiagnostic
for deterministic turns. No user text, birth data or model text.
- D5: /v5/versions memo (30 s, complete identities only, per transport); the
exit gate skips its second persistNextInterviewIfIdle when the run's own
call found the next focus already active (provably identical).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
BUG-1045 (recurrence of BUG-585 via BUG-969): the snapshot merge now drops
the attached stem from the streamed "ack + stem" text with the same
stripQuestionSentences the GET route uses; server text unchanged.
BUG-1046: a failed choice submit (409 / network) or typed send withdraws
the local answered mark, remounts the card, re-reads the Case and shows
"这次没提交上,请再点一次。"; the persisted question hangs on the latest
settled assistant message with other copies of the same focus removed
(standalone block only when nothing can carry it); willContinue and the
send() settle merge unseen assistant turns (same gap as BUG-685).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Carry explicit local date intervals instead of inferring the day from clock
order. Cluster width, delivery, adoption, and reports keep the actual civil
date; adopted date is stored separately from the reported birth_date.
Algorithm identity is scoring-9 / spec-v5. Scoring weights, confirmation
thresholds, and Skill version are unchanged. Isolated Linux final-3 gates
passed; four pre-existing Python failures remain. This is not a production
release.
Unify minute and block cache identity, keep unverifiable historical results read-only across server tools and write entrypoints, and aggregate completed receipt sources chronologically through a compatible function migration.
Co-Authored-By: Claude Code <noreply@anthropic.com>
Add date-isolated caches and regression coverage, align scoring identity, and freeze full research reruns while preserving historical artifacts. Record unresolved cache/receipt identity and end-to-end acceptance gaps for branch review only.
Co-Authored-By: Claude Code <noreply@anthropic.com>