Commit Graph
2620 Commits
Author SHA1 Message Date
Jesse_ChenandClaude Opus 5.5 48fb160fc9 feat(consult): six-field evidence-card record in agent observability (log only)
After each natal turn the route logs a separate [agent-observability] event
{runId, agentVersion, evidenceCard: {domains, cardVersion, cardChars,
cardTokenEstimate, citedFieldIds, feedback}}. Cited field ids follow the
research R5 rule (ISO date, degree, planet-in-sign phrase of 8+ chars
appearing verbatim in a finished answer); matched text is discarded. Thumbs
are client state only today, so feedback is "none"; storing thumbs per turn
needs a table and is left to a follow-up. The strict schema has no user,
session, question or answer field.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 cb3ee55837 feat(consult): one read-only evidence lookup per turn; settle open clauses at tool calls (BUG-1059)
read-consultation-evidence returns one closed-enum section (a formal varga,
research/extended vargas, a Western layer, yogas, Ashtakavarga, Shadbala,
transits, Chara Dasha, arudha, karakas, KP, gulika, kakshya, mahadashas,
thematic evidence) from this request's finished calculation, never
recalculates, answers unavailable on a cache miss and refuses a second call.
The receipt records the step and the write row shows 「正在多看一眼:…」.
A lookup after answer text went out keeps the released text whole: a verbatim
restart is dropped as it arrives (40-char confirmation), a continuation is
kept, and settlement still reads the step that wrote the answer; the lookup
runs on the answer clock without resetting it. A length continuation carries
the lookup result with the card. BUG-1059: the visible-text transformer's open
clause is settled at each tool call, so unpunctuated narration no longer
leaks into the answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 147ebc1789 feat(consult): prompt and Skill read from the evidence card (Skill 6.9.17)
The natal system prompt now says: the card in claim_cards is the chart
evidence for this answer, quote it as given, look up one further section of
this calculation with read-consultation-evidence before writing, and read
evidence_card.backstage as a confidence cap. The "use every executed layer"
and must_use_layers sentences are replaced; local_layers paths now point at
the card. SKILL.md 关联技法完整调取 / 0.0.1 and router 0.7 state that the
full result stays in receipts, the 本轮技法 panel and reports while web chat
answers from the card; computing the full spectrum is unchanged. Skill and
package version 6.9.16 -> 6.9.17 (tests/run_all.py 三栏).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 850f18605b feat(consult): the answer model reads the answer contract plus an evidence card
New frontend/src/lib/consultation-evidence-card.ts ports the research
CARD_SPECS: base section (ascendant, house signs, placements with degrees,
functional benefics/malefics with lordship, Vimshottari MD/AD/PD with dates,
Narayana md/ad/pd) plus a section per domain; values copied verbatim from
the projection or engine context, gaps listed, never filled. The tool now
returns toModelEvidenceView: status, evidence_contract (policy, blockers,
layers, limitation), claim_cards = the card, evidence_card meta (D4 line,
supplementable sections), rectification, methodology, domains; the audit
table, spectra, must_use_layers and presentation stay server-side and the
single-domain consultations copy is gone. 「本轮技法」 rows are unchanged.
Golden tests over three public AA engine captures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 5c65834dda feat(consult): split family into parents and children domains (engine still family)
parents (D12, 4/9 houses, Sun/Moon) and children (D7, 5th house, Jupiter,
PK) are plan/card domains with Chinese labels and the brief's aliases. They
send the family route contract to the engine and are rejected as a stored
session theme, so neither Python nor the chat_sessions.theme CHECK changes.
Methodology reports no strict checklist for them; must-use layers follow the
plan domain. Tests with 原值/新值/原因 where assertions changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 75a1844cdf fix(consult): project the running Narayana period and pratyantar dates to the model (BUG-1054)
timingKeys now admits current_dasha.md/ad/pd (sign, lord, years,
start_age, end_age), remaining_years and pratyantar_dasha_timeline. Depth
and item caps unchanged. Golden regression over three public AA engine
captures asserts values, not key presence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 c33701a4bd docs(bugs): BUG-1052/1053 deployed to staging 560d4fc2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:20:51 +08:00
Jesse_ChenandClaude Opus 5.5 a98d3b2091 docs(tasks): rectification grounding brief; reserve BUG-1055..1058
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:19:54 +08:00
Jesse_ChenandClaude Opus 5.5 599fe7a94b docs(tasks): consult evidence-card implementation brief
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:52:07 +08:00
Jesse_ChenandClaude Opus 5.5 5bafdfdc8b docs(tasks): evidence-card research item-by-item acceptance
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:43:02 +08:00
Jesse_ChenandClaude Opus 5.5 560d4fc2dd fix(research): reuse existing handler instead of new __new__ forgeries
Independent Staging Quality Gate / validate (push) Successful in 11m55s
Independent Staging Quality Gate / publish (push) Successful in 3m28s
Gate run 2958 failed tests/test_api_server_growth_contract.py: the
evidence-card research script added two JyotishAPIHandler forgery sites.
Reuse capture_report_blocked_repairs_golden._handler(); output unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:40:47 +08:00
Jesse_ChenandClaude Opus 5.5 45ed054f09 docs(tasks): evidence-card research acceptance
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:31:59 +08:00
jesse-ux d8d03b5a69 research: consultation evidence-card inventory and draft
Independent Staging Quality Gate / validate (push) Failing after 6m40s
Independent Staging Quality Gate / publish (push) Skipped
Measure the model-visible consultation payload on three public charts, draft per-domain cards, and record the Narayana/pratyantar projection gap as BUG-1054. No runtime behavior change.
2026-09-27 01:27:57 +08:00
Jesse_ChenandClaude Opus 5.5 3b67d8e945 docs(tasks): consult single-pass acceptance
Independent Staging Quality Gate / validate (push) Successful in 11m39s
Independent Staging Quality Gate / publish (push) Canceled after 3m31s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:12:51 +08:00
Jesse_ChenandClaude Opus 5.5 3f8e180238 docs: BUG-1053 record, progress, real-device checklist, changelog and board row
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:04:27 +08:00
Jesse_ChenandClaude Opus 5.5 eef0cb7486 fix(consult): write the answer in the step that saw the chart (BUG-1053)
The natal loop's step after run-jyotish-consultation saw the evidence, but
its text was drained and a second, blind compose stream (history + question
only, empty findings) wrote the user-visible answer. Remove compose,
interpret and the drain; keep the loop's own final-step text.

- stepScopedAnswer: per-step holding; text of a step that calls a tool is
  dropped, so narration around tool calls never reaches the answer
- writing shape (opener + four headings) moves into the user turn
- length continuation receives the calculation result; Pass 4 whole-answer
  reject retries through retryForAnswer with the rewrite hint
- createConsultationRunClock: tools keep the 110s tool phase; the loop is
  handed to the 70s answer clock when the calculation result arrives
- settlement judges the step that wrote the answer (BUG-1051 kept)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:01:38 +08:00
Jesse_ChenandClaude Opus 5.5 03c7c0a97e docs(tasks): home-landing-blank acceptance
Independent Staging Quality Gate / validate (push) Successful in 11m42s
Independent Staging Quality Gate / publish (push) Canceled after 23s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:00:38 +08:00
Jesse_ChenandClaude Opus 5.5 d2551437c4 fix(home): land on the blank starter home after login and on bare / (BUG-1052)
Product decision 2026-09-27 (D1-D4), overriding the BUG-1038 login-return
stash and the "latest session" default landing:

- A full load of / without ?c= (login, typed address, bookmark, refresh of
  bare /) lands on the current person's blank starter home; an existing
  empty draft of that person is reused, otherwise one is created locally.
- ?c= and ?new=1 are unchanged; a refresh inside a conversation keeps its ?c=.
- The login-return stash is removed: redirectToLogin and sidebar links no
  longer write it, bootstrap no longer reads it and clears a leftover value
  once. replace-selected, the lookup origin/other-subject branch and the warm
  resumeRectification dependency go with it.
- A background answer being recovered no longer takes over the blank home
  (same rule BUG-1015 set for ?new=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 00:51:51 +08:00
Jesse_ChenandClaude Opus 5.5 76924e3362 docs(tasks): consult evidence-card research brief
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 00:50:04 +08:00
Jesse_ChenandClaude Opus 5.5 5f2e007f64 docs(bugs): BUG-1049/1050/1051 deployed to staging e8195783
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:49:35 +08:00
Jesse_ChenandClaude Opus 5.5 e8195783f8 docs(tasks): consult truncation acceptance
Independent Staging Quality Gate / validate (push) Successful in 12m0s
Independent Staging Quality Gate / publish (push) Successful in 3m39s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:28:34 +08:00
Jesse_ChenandClaude Opus 5.5 1dd51f2d42 docs: BUG-1051 record, progress, real-device checklist, changelog and board row
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:20:56 +08:00
Jesse_ChenandClaude Opus 5.5 1530a0dd63 fix(consult): give compose its own clock and never settle a cut answer (BUG-1051)
The tool loop and the answer-writing stream shared one 110s AbortSignal.
Mastra 1.50 does not throw on abort: it emits an abort chunk and
finish(tripwire) and closes normally, so a half-written answer reached
onComplete, was charged and persisted as completed.

- Compose, length continuation and answer retry run on a 70s answer clock
  started on first use (worst case 110s + 70s = 180s; maxDuration 240).
- Settlement requires finish=stop from the stream that wrote the answer;
  abort/tripwire, content-filter, tool-calls, other/unknown/error or a
  missing finish with visible text ends as answer_truncated (cancel, no
  charge). The abort chunk records an abort runtime step; a cut stream no
  longer flushes its dangling Pass 4 sentence. length still continues.
- [agent-observability] gains composeFinishReason, composeAborted and
  answerVisibleChars (enum/boolean/count only).
- Regression tests use a real Mastra Agent over a fake model; the BUG-305
  hand-thrown DOMException fixture is kept with a three-column note, and
  eleven fixtures gain the finish(stop) chunk real streams always carry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:17:36 +08:00
Jesse_ChenandClaude Opus 5.5 16f4600d5e docs(tasks): opening-plain acceptance
Independent Staging Quality Gate / validate (push) Successful in 13m52s
Independent Staging Quality Gate / publish (push) Canceled after 1m31s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:12:56 +08:00
Jesse_ChenandClaude Opus 5.5 8fc1a5bded fix(rectification): plain-word opening with examples in the stem, plain de-duplicated step receipt (BUG-1049, BUG-1050)
- BUG-1049 (recurrence of BUG-504): the opening stem carries six examples and
  an example answer again and is server-owned on the zero-evidence opening;
  the body is two plain sentences (no 大运/盘面/代表分钟/精确到秒, no year, no
  question). Stem de-dup compares whole sentences / near-equality instead of a
  12-char prefix, which had deleted the body's examples sentence since
  aa7ccb30 (BUG-604) + dd8f35f7 (BUG-648).
- BUG-1050: plain step labels; a finished step label shows once and
  「已完成 N 步」counts shown rows; failed rows read 「…未完成」 from the
  in-progress wording.
- Skill 10.0.30 -> 10.0.31 (OpeningPolicy); 10.0.30 kept as deprecated.
- VOICE / DESIGN / CHANGELOG / BUG_HISTORY / PROGRESS / real-device checklist
  and screenshots.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 22:04:56 +08:00
Jesse_ChenandClaude Opus 5.5 e53052a258 fix: restore vendored qizheng CLI and upload-artifact action deleted by 12cbe6f8
Independent Staging Quality Gate / validate (push) Successful in 12m26s
Independent Staging Quality Gate / publish (push) Successful in 6m7s
12cbe6f8 (docs-only intent) swept two pre-existing working-tree deletions in
via commit -a. Restore both byte-for-byte from e801fcf5; log ERR-112.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 20:36:08 +08:00
Jesse_ChenandClaude Opus 5.5 e0e2b38210 docs(tasks): rectification code split acceptance
Independent Staging Quality Gate / validate (push) Failing after 6m56s
Independent Staging Quality Gate / publish (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 20:25:19 +08:00
Jesse_ChenandClaude Opus 5.5 c1592c21ad docs(tasks): rectification code split progress, changelog, index row
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 20:14:14 +08:00
Jesse_ChenandClaude Opus 5.5 a00c40d4b8 refactor(rectification): split runV9AgentTurn into prepare / attempt / retry / finish
Move-only (TASK-rectification-code-split-20260926 T3/T4). No behavior,
receipt, billing, Skill or scoring change.

- agent-run.ts: 1400 -> 135 lines; runV9AgentTurn 1046 -> 25 lines, no
  nested functions. It keeps the public types, the budget constants and the
  phase order: agent-run-prepare.ts (reliability, session, year re-ask, Skill
  identity, delivered guard, reserve, turn row / replay),
  agent-run-retry.ts (attempt loop), agent-run-attempt.ts (the old nested
  streamAttempt), agent-run-finish.ts (billing settle, receipts, interview,
  finalize), agent-run-support.ts (outcome types, retry classification, turn
  receipt writers), agent-run-messages.ts (buildAgentMessages /
  buildOpeningBrief, re-exported).
- One token changed with the move: the attempt passes its own
  `previousErrorCode` argument to buildAgentMessages instead of reading the
  enclosing `lastAttemptError`; the loop passes that same value (clears the
  old unused-parameter warning).
- Tests: whole-source contracts read tests/rectification-agent-run-surface.ts;
  behavioral runV9AgentTurn tests unchanged. agent-run growth caps added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 19:52:51 +08:00
Jesse_ChenandClaude Opus 5.5 eb773f1299 refactor(rectification): split POST /api/rectification/agent into branch handlers
Move-only (TASK-rectification-code-split-20260926 T2/T4). No behavior, API,
copy, DB, Skill, billing or scoring change.

- route.ts: 1077 -> 152 lines; POST 927 -> 114. It builds the Supabase
  clients (so the setup-failure mapping stays here) and assembles:
  agent-route-request.ts (auth / product / schema / flag / Case-Session
  binding), agent-route-typed-message.ts (declared-window reply, typed-answer
  preflight, unfocused classification), agent-route-structured-choice.ts,
  agent-route-agent-turn.ts (opening / read-only / typed agent stream and its
  exit gate), agent-route-billing.ts, agent-route-support.ts (schema, one-shot
  NDJSON reply, context types).
- The four mutable preflight lets (expectedWrite, collectIntent,
  writeClassified, classifierDiagnostic) that crossed branches are one
  turnState object; values and flow unchanged.
- Tests: whole-source contracts read tests/rectification-agent-route-surface.ts;
  the typed fast-path slice is rebuilt from the new files; billing,
  declared-window and structured-choice slices now call the handlers
  (billing in a child process because feature-pricing imports server-only),
  each with 原值/新值/原因. Route growth caps added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 19:40:07 +08:00
Jesse_ChenandClaude Opus 5.5 3a66c39f48 refactor(rectification): split the chat component into hooks, parameter functions and blocks
Move-only (TASK-rectification-code-split-20260926 T1/T4). No behavior, API,
copy, DB, Skill or scoring change.

- rectification-agentic-chat.tsx: 2043 -> 737 lines; component body
  1625 -> 630; useState 35 -> 12, useRef 16 -> 9, effects 9 -> 5,
  useCallback 12 -> 6.
- Snapshot sync -> hooks/use-rectification-case-snapshot.ts; board layout
  and live label clock -> two small hooks.
- send / submitStructuredChoice / acceptCandidate bodies ->
  lib/rectification-chat-{turn,choice,accept}-run.ts parameter functions
  (bodies verbatim, deps destructured to the same names); copy/regenerate,
  question repair, pure transcript and snapshot helpers and the per-render
  view derivation -> lib/rectification-chat-*.ts (no React hooks).
- Question-gap blocks and the read-only range line ->
  components/rectification-question-gap-notices.tsx.
- Dependency arrays kept exactly as before (lint warnings +3, listed in
  PROGRESS); no dep added to appease the linter.
- Tests: whole-source contracts read tests/rectification-chat-surface.ts
  (container + split files, like home-surface.ts); slices of moved code now
  call the extracted functions or render the extracted block, each with a
  原值/新值/原因 comment. New tests/rectification-growth-contract.test.ts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 19:23:52 +08:00
Jesse_ChenandClaude Opus 5.5 12cbe6f868 docs(tasks): telemetry deployed, test:db passed in gate run 2938
Independent Staging Quality Gate / validate (push) Failing after 6m44s
Independent Staging Quality Gate / publish (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 18:56:15 +08:00
Jesse_Chen e801fcf53b fix(ci): preserve npm install diagnostics and classify forced timeout
Independent Staging Quality Gate / validate (push) Successful in 10m17s
Independent Staging Quality Gate / publish (push) Successful in 21m1s
2026-09-26 17:52:37 +08:00
Jesse_ChenandClaude Opus 5.5 40d7930ab4 docs(tasks): telemetry acceptance record
Independent Staging Quality Gate / validate (push) Failing after 16m43s
Independent Staging Quality Gate / publish (push) Skipped
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 15:46:25 +08:00
Jesse_ChenandClaude Opus 5.5 d0bfc1fc3d feat(rectification): anonymous aggregate telemetry + admin summary page
One row per rectification Case, written once when the range card is first
delivered (GET /api/rectification/cases/[caseId], fire-and-forget after the
response is built). Numbers and closed enums only: no user / case / session
id, birth data, names, text or timestamps finer than the ISO week. Dedupe via
a separate case_id ledger that cascades with the Case (and account deletion).

Migration 20260926010000 is additive: two RLS tables with no runtime table
grants, SECURITY DEFINER write (service_role), purge (service_role) and
aggregate-only summary (admin_runtime) functions; 180-day retention.

Admin: 「校正统计」 page + GET /api/admin/rectification-telemetry
(admin.customers.read), aggregates only, no per-row view or export.

TASK-rectification-telemetry-20260926. test:db not run locally (no Docker).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 15:36:29 +08:00
Jesse_ChenandClaude Opus 5.5 f74825a27c docs(tasks): close fewer-probes-card and offline-research acceptance
Independent Staging Quality Gate / validate (push) Successful in 11m51s
Independent Staging Quality Gate / publish (push) Successful in 3m48s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 15:04:05 +08:00
Jesse_ChenandClaude Opus 5.5 0f5442cea2 research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
  a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
  (two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
  identity at ±30/±60 by construction and leave six-question metrics
  unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
  pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
  gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
  (all-pairs adds no dated probes); the bottleneck is monthly evaluation.
  New finding recorded as BUG-1048 (investigating): _boundary_windows
  year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
  research docs; README board row -> 待验收. No production code, scoring,
  thresholds, gates or Skill changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:47:48 +08:00
Jesse_ChenandClaude Opus 5.5 7203d94e9c fix(rectification): 区分不开时不显示「最可能」副标题(验收补改)
Independent Staging Quality Gate / validate (push) Successful in 11m49s
Independent Staging Quality Gate / publish (push) Successful in 3m44s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:45:11 +08:00
Jesse_ChenandClaude Opus 5.5 4e6e8d87ef feat(rectification): 少问几道就出卡,交付卡以范围为主(D1–D4)
- D1 引导窗口题每个校正最多 2 道(前端 GUIDED_WINDOW_CASE_LIMIT;引擎
  GUIDED_COLLECT_LIMIT 未动:event_probes.py 属冻结评分身份,改它需重新冻结)
- D2 七条定向线与跳过线重问问完即出卡,没问到的引导窗口不再挡卡,出卡后也不再挂窗口题
- D3 卡头加副标题「最可能 HH:MM」
- D4 前两列相差 ≥5 个百分点才显示相对可能性,否则一句「这几个时刻目前区分不开……」
- 离线回放 scripts/research/fewer_probes_card_replay.py:真值不降、宽度中位 ±1 分钟、提问 11.4→7.8
- Skill 10.0.29 → 10.0.30;DESIGN / VOICE / CHANGELOG / PROGRESS / 真机清单

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:27:12 +08:00
Jesse_ChenandClaude Opus 5.5 7ddce2c96a docs: reconcile stale board rows, BUG records and BLOCKED items (docs-reconcile 20260926)
Board: 59 rectification-related rows re-checked against origin/staging
ancestry, Gitea deploy-staging runs (current 8a409434, run 2930) and
acceptance records; unaccepted merges marked 已合入, real-device debt kept.
BUG_HISTORY: 085/086 closed_obsolete (V5 retired in f3946eaa); 981/984
fix-version lines record deployed commits/runs, status stays investigating;
1038-1047 fix-version lines corrected; 743 left as is (style probes still
score via tie_break path). BLOCKED: undeployed notes struck with evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:05:01 +08:00
Jesse_ChenandClaude Opus 5.5 abb05b675c docs(tasks): rectification round briefs (fewer probes/card, telemetry, research, split, reconcile)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 13:46:21 +08:00
Jesse_ChenandClaude Opus 5.5 3ce11ee04d docs: BUG-1047 deployed (8a409434)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 12:19:31 +08:00
Jesse_ChenandClaude Opus 5.5 8a40943455 docs(tasks): record Claude acceptance of BUG-1047
Independent Staging Quality Gate / validate (push) Successful in 12m8s
Independent Staging Quality Gate / publish (push) Successful in 3m30s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 12:00:35 +08:00
Jesse_ChenandClaude Opus 5.5 86ff9a40d1 docs(rectification): BUG-1047 records, progress, device checklist, ledger ERR-110/111
- BUG_HISTORY BUG-1047 (investigating), PROGRESS with timing estimates, A/B
  evidence, diagnostics guide, test diffs and env gaps.
- CHANGELOG (Skill version not bumped), docs/testing checklist + headless
  screenshots, tasks README row -> 待验收, BLOCKED entry.
- pre-work ledger: ERR-110 (byte-identical engine perf changes still trip the
  frozen-scoring integrity gate), ERR-111 (Shadbala fingerprint depends on
  PYTHONHASHSEED).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 11:53:27 +08:00
Jesse_ChenandClaude Opus 5.5 ab0f01a9cf fix(rectification): stream-first typed turns, 10 s classifier cap, stage progress, run timings (BUG-1047)
- D2: each intent-classifier attempt is capped at 10 s (same session model,
  thinking untouched); a hang takes the existing retry -> classifier_unavailable
  path. Success is timed too (RectificationClassifierDiagnostic).
- D3: a typed message builds the NDJSON stream first; the first line is
  turn.progress "received", then the classifier and deterministic replies run
  inside the stream. Preflight rejections become turn.rejected (old status,
  code, message) and the client handles them like the old HTTP rejection.
  Stage lines 收到,正在对照你的档案… / 正在记下这件事… / 正在重新对照盘面… /
  正在准备下一个问题… are driven by existing tool events and engine calls, are
  transient (live row only) and never persisted. VOICE / DESIGN updated.
- D4: RectificationRunDiagnostic records per-step start/end, provider token
  usage incl. reasoning tokens, classifier timing and per-engine-call
  durations (AsyncLocalStorage scope per turn); RectificationTurnDiagnostic
  for deterministic turns. No user text, birth data or model text.
- D5: /v5/versions memo (30 s, complete identities only, per transport); the
  exit gate skips its second persistNextInterviewIfIdle when the run's own
  call found the next focus already active (provably identical).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 11:45:47 +08:00
Jesse_ChenandClaude Opus 5.5 fe76d54795 docs: BUG-1043..1046 deployed (f1d16405)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 11:13:11 +08:00
Jesse_ChenandClaude Opus 5.5 f1d1640572 docs(tasks): record Claude acceptance of BUG-1045/1046
Independent Staging Quality Gate / validate (push) Successful in 11m53s
Independent Staging Quality Gate / publish (push) Successful in 3m27s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 10:54:46 +08:00
Jesse_ChenandClaude Opus 5.5 e4c1c7a342 fix(rectification): one stem per turn, one card per focus (BUG-1045/1046)
BUG-1045 (recurrence of BUG-585 via BUG-969): the snapshot merge now drops
the attached stem from the streamed "ack + stem" text with the same
stripQuestionSentences the GET route uses; server text unchanged.

BUG-1046: a failed choice submit (409 / network) or typed send withdraws
the local answered mark, remounts the card, re-reads the Case and shows
"这次没提交上,请再点一次。"; the persisted question hangs on the latest
settled assistant message with other copies of the same focus removed
(standalone block only when nothing can carry it); willContinue and the
send() settle merge unseen assistant turns (same gap as BUG-685).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 10:49:25 +08:00
Jesse_ChenandClaude Opus 5.5 07998972e0 docs: record Claude acceptance of BUG-1043/1044
Independent Staging Quality Gate / publish (push) Canceled after 0s
Independent Staging Quality Gate / validate (push) Canceled after 6m33s
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 10:48:20 +08:00
Jesse_ChenandClaude Opus 5.5 da2613ffd9 fix(chat): attach scroll anchor when the scroller appears; only reader scroll releases a pin (BUG-1043, BUG-1044)
- The anchor listener and follow observer now attach whenever the scroller
  element itself appears (checked after every commit, no-op unless element,
  active or resetKey changed). The home page mounts `.conversation` after its
  loading screen with unchanged active/resetKey, so a directly opened session
  never got a listener, never landed on its newest content, showed the jump
  chip under short replies and did not follow after pressing it.
- After a pin, geometry no longer releases the hold: only a wheel, touch drag,
  scroll key or scrollbar press followed by a scroll within 1s does. The
  rectification pin rests 94px from the bottom, inside the 96px threshold,
  which dragged long replies to their last line.
- Real React lifecycle tests (loading screen -> reveal, 94px rest), DESIGN,
  BUG history, PROGRESS, CHANGELOG, device checklist and CDP screenshots.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 10:42:12 +08:00