The earlier report had no previous-turn rows. This run measures V0, V1, V2, and Flash on the re-extracted corpus and records that the real sample is not representative of the simulated set.
Co-authored-by: Cursor <cursoragent@cursor.com>
Research brief reopening the 09-19 Jev evaluation: same model
(jev-1.13.0), only the state changes (previous turn + previous
decision + continuation Noul, per Magpie v0.1.142). Re-extract
source B with case_id; thresholds unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199rbQDTsUbCVw84wc8BTFe
Adds rules previously kept only in Claude's local memory so Codex/Cursor
follow them too: path-scoped docs commits (ERR-112), PostgREST compat
layer needs real queries (BUG-990), grep source-text contract tests
before moving code (BUG-933/934/939/992/1014), privacy marker test for
docs with measurements, quick gate for any scripts/tests change, Node
`# cancelled` check (BUG-995), and history-open check on Skill bumps
(BUG-621).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199rbQDTsUbCVw84wc8BTFe
T7 of TASK-rectification-grounding-20260927. Skill not bumped (10.0.31 text
unchanged, sliced per turn). regenerate_agentic_rectification_turn kept, to be
retired in a later brief.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T6 of TASK-rectification-grounding-20260927.
- skill-slice.ts: per-action slice rule as a code constant, keyed on heading
titles (evidence: 「ConversationFocus 与意图承接」「批量证据与日期真实性」, i.e.
§5/§7 of the 10.0.x layout); other actions and any bound Skill without those
headings (9.0.0) get the whole body, so historical Cases still run with their
exact bound Skill (BUG-621).
- The rectification Agent declares providesSkillDiscovery "on-demand"
(rectificationSkillBoundProcessor): no <available_skills> block with a temp
path and no "call the skill tool" system message; getSkill still loads the
bound package.
- Measured on a real Agent + recording model (public AA case, estimate = CJK
chars + other chars / 4): fixed overhead per call 12,002 → 6,471 tokens
(step 0: 8,440 → 2,909). Skill text unchanged; no version bump.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T5 of TASK-rectification-grounding-20260927.
- rectification-offer-candidates returned the full projection (~100 KB on the
public AA case: *_pre_inference, 31 KB contrast packet, probe lists). It now
returns the compare stripping (offerModelProjection); 100,012 → 28,799 bytes.
- agentVisibleLatestProjection also drops engine_indistinguishable_width_minutes
(audit-only since BUG-593), keeps the verification Markdown once
(skill_verification_report; range_delivery.verification_markdown was a
byte-identical copy) and shows the slim candidate list.
- Receipt fingerprints stay over the full payloads; the case API projection
(latestResultToolProjection) is unchanged.
Contract tests on a real local engine response (public AA chart).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T3 of TASK-rectification-grounding-20260927 (red line 3).
The turn-decision read-case is the model's only conversation memory; with
about 7+ candidates (9 in the public AA case) it exceeded 6 KB and cleared
recent_turns and relevant_evidence_summary first. Candidates in the
model-visible inference now carry time / score / status / cluster_range only
(the last round names candidates by time), candidate_summary.candidates (a
duplicate) is gone, and over budget the order is: drop cluster ranges → keep
the best six candidates → shorten turns/evidence → clear them. Stored
inference and receipts are untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T2 of TASK-rectification-grounding-20260927 (product decision P1).
The rectification 「重试(重新生成)」 was a blind agent.generate rewrite that
persisted unchecked text (invented range / fit rate reproduced). Removed:
the regenerate route, regenerate-turn.ts, the regeneration agent and its
read-only tool set, the client regenerate action/state and the
regenerating/canRegenerate props. ChatMessageActions renders the regenerate
button only when onRegenerate is passed; ordinary consultation is unchanged.
The DB function regenerate_agentic_rectification_turn is kept (AGENTS §7.6,
retire in a later round). DESIGN.md / VOICE.md updated; source-contract
tests follow with 原值/新值/原因 notes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T4 of TASK-rectification-grounding-20260927 (BUG-588 / BUG-569 family).
The deferFollowup path persisted the choice before the agent turn, so the
agent turn's prepare read the post-answer range and the deferred choice
narration was never persisted: nobody said the range moved. The typed-message
preflight now hands the pre-answer credible range to the agent turn
(rangeBeforeTurn); the turn's one server range sentence covers the answer and
the dated event. Route-level regression with the real handler.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
T1 of TASK-rectification-grounding-20260927 (recurrence of BUG-588).
- The attempt no longer streams range-changed / rescore-skipped /
compare-failed sentences; the finish whitelists and trims the model body,
then joins the server facts, and emits one final replace equal to the
persisted text.
- P3 whitelist (spoken-grounding.ts): a model sentence with a clock, clock
range or percentage that is not this turn's server fact is dropped whole;
the batch recap stands in when nothing is left.
- record-evidence-batch returns range_after_rescore (post-rescore
credible_range, representative minute, fit percent, delivers_range_this_turn);
the receipt fingerprint stays over the old shape.
- System prompt: range is said by the server; the delivery three sentences
only when the batch says this turn delivers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
PROGRESS with the final domain -> technique table, before/after
model-visible sizes over 3 public AA charts x 10 questions, answer-contract
keys kept/removed, red-line test list, lookup and clock design, telemetry,
assertion changes and test/build/gzip numbers. Device checklist for
parents/children/marriage/career, follow-up domain carry-over, D60 lookup
and waiting time. Task index row set to 待验收.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
After each natal turn the route logs a separate [agent-observability] event
{runId, agentVersion, evidenceCard: {domains, cardVersion, cardChars,
cardTokenEstimate, citedFieldIds, feedback}}. Cited field ids follow the
research R5 rule (ISO date, degree, planet-in-sign phrase of 8+ chars
appearing verbatim in a finished answer); matched text is discarded. Thumbs
are client state only today, so feedback is "none"; storing thumbs per turn
needs a table and is left to a follow-up. The strict schema has no user,
session, question or answer field.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
read-consultation-evidence returns one closed-enum section (a formal varga,
research/extended vargas, a Western layer, yogas, Ashtakavarga, Shadbala,
transits, Chara Dasha, arudha, karakas, KP, gulika, kakshya, mahadashas,
thematic evidence) from this request's finished calculation, never
recalculates, answers unavailable on a cache miss and refuses a second call.
The receipt records the step and the write row shows 「正在多看一眼:…」.
A lookup after answer text went out keeps the released text whole: a verbatim
restart is dropped as it arrives (40-char confirmation), a continuation is
kept, and settlement still reads the step that wrote the answer; the lookup
runs on the answer clock without resetting it. A length continuation carries
the lookup result with the card. BUG-1059: the visible-text transformer's open
clause is settled at each tool call, so unpunctuated narration no longer
leaks into the answer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
The natal system prompt now says: the card in claim_cards is the chart
evidence for this answer, quote it as given, look up one further section of
this calculation with read-consultation-evidence before writing, and read
evidence_card.backstage as a confidence cap. The "use every executed layer"
and must_use_layers sentences are replaced; local_layers paths now point at
the card. SKILL.md 关联技法完整调取 / 0.0.1 and router 0.7 state that the
full result stays in receipts, the 本轮技法 panel and reports while web chat
answers from the card; computing the full spectrum is unchanged. Skill and
package version 6.9.16 -> 6.9.17 (tests/run_all.py 三栏).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
New frontend/src/lib/consultation-evidence-card.ts ports the research
CARD_SPECS: base section (ascendant, house signs, placements with degrees,
functional benefics/malefics with lordship, Vimshottari MD/AD/PD with dates,
Narayana md/ad/pd) plus a section per domain; values copied verbatim from
the projection or engine context, gaps listed, never filled. The tool now
returns toModelEvidenceView: status, evidence_contract (policy, blockers,
layers, limitation), claim_cards = the card, evidence_card meta (D4 line,
supplementable sections), rectification, methodology, domains; the audit
table, spectra, must_use_layers and presentation stay server-side and the
single-domain consultations copy is gone. 「本轮技法」 rows are unchanged.
Golden tests over three public AA engine captures.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
parents (D12, 4/9 houses, Sun/Moon) and children (D7, 5th house, Jupiter,
PK) are plan/card domains with Chinese labels and the brief's aliases. They
send the family route contract to the engine and are rejected as a stored
session theme, so neither Python nor the chat_sessions.theme CHECK changes.
Methodology reports no strict checklist for them; must-use layers follow the
plan domain. Tests with 原值/新值/原因 where assertions changed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
timingKeys now admits current_dasha.md/ad/pd (sign, lord, years,
start_age, end_age), remaining_years and pratyantar_dasha_timeline. Depth
and item caps unchanged. Golden regression over three public AA engine
captures asserts values, not key presence.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Gate run 2958 failed tests/test_api_server_growth_contract.py: the
evidence-card research script added two JyotishAPIHandler forgery sites.
Reuse capture_report_blocked_repairs_golden._handler(); output unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Measure the model-visible consultation payload on three public charts, draft per-domain cards, and record the Narayana/pratyantar projection gap as BUG-1054. No runtime behavior change.
The natal loop's step after run-jyotish-consultation saw the evidence, but
its text was drained and a second, blind compose stream (history + question
only, empty findings) wrote the user-visible answer. Remove compose,
interpret and the drain; keep the loop's own final-step text.
- stepScopedAnswer: per-step holding; text of a step that calls a tool is
dropped, so narration around tool calls never reaches the answer
- writing shape (opener + four headings) moves into the user turn
- length continuation receives the calculation result; Pass 4 whole-answer
reject retries through retryForAnswer with the rewrite hint
- createConsultationRunClock: tools keep the 110s tool phase; the loop is
handed to the 70s answer clock when the calculation result arrives
- settlement judges the step that wrote the answer (BUG-1051 kept)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Product decision 2026-09-27 (D1-D4), overriding the BUG-1038 login-return
stash and the "latest session" default landing:
- A full load of / without ?c= (login, typed address, bookmark, refresh of
bare /) lands on the current person's blank starter home; an existing
empty draft of that person is reused, otherwise one is created locally.
- ?c= and ?new=1 are unchanged; a refresh inside a conversation keeps its ?c=.
- The login-return stash is removed: redirectToLogin and sidebar links no
longer write it, bootstrap no longer reads it and clears a leftover value
once. replace-selected, the lookup origin/other-subject branch and the warm
resumeRectification dependency go with it.
- A background answer being recovered no longer takes over the blank home
(same rule BUG-1015 set for ?new=1).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
The tool loop and the answer-writing stream shared one 110s AbortSignal.
Mastra 1.50 does not throw on abort: it emits an abort chunk and
finish(tripwire) and closes normally, so a half-written answer reached
onComplete, was charged and persisted as completed.
- Compose, length continuation and answer retry run on a 70s answer clock
started on first use (worst case 110s + 70s = 180s; maxDuration 240).
- Settlement requires finish=stop from the stream that wrote the answer;
abort/tripwire, content-filter, tool-calls, other/unknown/error or a
missing finish with visible text ends as answer_truncated (cancel, no
charge). The abort chunk records an abort runtime step; a cut stream no
longer flushes its dangling Pass 4 sentence. length still continues.
- [agent-observability] gains composeFinishReason, composeAborted and
answerVisibleChars (enum/boolean/count only).
- Regression tests use a real Mastra Agent over a fake model; the BUG-305
hand-thrown DOMException fixture is kept with a three-column note, and
eleven fixtures gain the finish(stop) chunk real streams always carry.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
- BUG-1049 (recurrence of BUG-504): the opening stem carries six examples and
an example answer again and is server-owned on the zero-evidence opening;
the body is two plain sentences (no 大运/盘面/代表分钟/精确到秒, no year, no
question). Stem de-dup compares whole sentences / near-equality instead of a
12-char prefix, which had deleted the body's examples sentence since
aa7ccb30 (BUG-604) + dd8f35f7 (BUG-648).
- BUG-1050: plain step labels; a finished step label shows once and
「已完成 N 步」counts shown rows; failed rows read 「…未完成」 from the
in-progress wording.
- Skill 10.0.30 -> 10.0.31 (OpeningPolicy); 10.0.30 kept as deprecated.
- VOICE / DESIGN / CHANGELOG / BUG_HISTORY / PROGRESS / real-device checklist
and screenshots.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Move-only (TASK-rectification-code-split-20260926 T3/T4). No behavior,
receipt, billing, Skill or scoring change.
- agent-run.ts: 1400 -> 135 lines; runV9AgentTurn 1046 -> 25 lines, no
nested functions. It keeps the public types, the budget constants and the
phase order: agent-run-prepare.ts (reliability, session, year re-ask, Skill
identity, delivered guard, reserve, turn row / replay),
agent-run-retry.ts (attempt loop), agent-run-attempt.ts (the old nested
streamAttempt), agent-run-finish.ts (billing settle, receipts, interview,
finalize), agent-run-support.ts (outcome types, retry classification, turn
receipt writers), agent-run-messages.ts (buildAgentMessages /
buildOpeningBrief, re-exported).
- One token changed with the move: the attempt passes its own
`previousErrorCode` argument to buildAgentMessages instead of reading the
enclosing `lastAttemptError`; the loop passes that same value (clears the
old unused-parameter warning).
- Tests: whole-source contracts read tests/rectification-agent-run-surface.ts;
behavioral runV9AgentTurn tests unchanged. agent-run growth caps added.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Move-only (TASK-rectification-code-split-20260926 T2/T4). No behavior, API,
copy, DB, Skill, billing or scoring change.
- route.ts: 1077 -> 152 lines; POST 927 -> 114. It builds the Supabase
clients (so the setup-failure mapping stays here) and assembles:
agent-route-request.ts (auth / product / schema / flag / Case-Session
binding), agent-route-typed-message.ts (declared-window reply, typed-answer
preflight, unfocused classification), agent-route-structured-choice.ts,
agent-route-agent-turn.ts (opening / read-only / typed agent stream and its
exit gate), agent-route-billing.ts, agent-route-support.ts (schema, one-shot
NDJSON reply, context types).
- The four mutable preflight lets (expectedWrite, collectIntent,
writeClassified, classifierDiagnostic) that crossed branches are one
turnState object; values and flow unchanged.
- Tests: whole-source contracts read tests/rectification-agent-route-surface.ts;
the typed fast-path slice is rebuilt from the new files; billing,
declared-window and structured-choice slices now call the handlers
(billing in a child process because feature-pricing imports server-only),
each with 原值/新值/原因. Route growth caps added.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
Move-only (TASK-rectification-code-split-20260926 T1/T4). No behavior, API,
copy, DB, Skill or scoring change.
- rectification-agentic-chat.tsx: 2043 -> 737 lines; component body
1625 -> 630; useState 35 -> 12, useRef 16 -> 9, effects 9 -> 5,
useCallback 12 -> 6.
- Snapshot sync -> hooks/use-rectification-case-snapshot.ts; board layout
and live label clock -> two small hooks.
- send / submitStructuredChoice / acceptCandidate bodies ->
lib/rectification-chat-{turn,choice,accept}-run.ts parameter functions
(bodies verbatim, deps destructured to the same names); copy/regenerate,
question repair, pure transcript and snapshot helpers and the per-render
view derivation -> lib/rectification-chat-*.ts (no React hooks).
- Question-gap blocks and the read-only range line ->
components/rectification-question-gap-notices.tsx.
- Dependency arrays kept exactly as before (lint warnings +3, listed in
PROGRESS); no dep added to appease the linter.
- Tests: whole-source contracts read tests/rectification-chat-surface.ts
(container + split files, like home-surface.ts); slices of moved code now
call the extracted functions or render the extracted block, each with a
原值/新值/原因 comment. New tests/rectification-growth-contract.test.ts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8