Commit Graph
258 Commits
Author SHA1 Message Date
Jesse_ChenandClaude Fable 5.1 88dca263fe docs(research): holdout v5 full-set baseline scorecard and progress close-out (BUG-1090)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:52:43 +08:00
Jesse_ChenandClaude Fable 5.1 e380026641 docs(research): holdout v5 build notes, v4 subset reproduction, progress and BUG-1090
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:52:43 +08:00
Jesse_ChenandClaude Fable 5.1 54e63fe46e research(rectification): holdout v5 — event cutoff rule in protocol and schema test
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:52:43 +08:00
Jesse_ChenandClaude Fable 5.1 8ceb94da02 research(rectification): holdout v5 — drop current-year events (harness cutoff 2026-09-14), index pages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:52:43 +08:00
Jesse_ChenandClaude Fable 5.1 e9dc8023f2 research(rectification): holdout v5 checkpoint — protocol, roster (47 new AA cases), quoted events, build script
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:52:18 +08:00
Jesse_ChenandClaude Fable 5.1 abf2ad6aea research(rectification): scoring-method scaffold — LR weights, absence, precision follow-up pipelines on v4, verdicts pending v5 (BUG-1091)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:06:45 +08:00
jesse-uxandClaude Code 64aa0ec20c feat(chart): polish chart presentation and symbol explanations
Independent Staging Quality Gate / publish (push) Canceled after 0s
Independent Staging Quality Gate / validate (push) Canceled after 1m27s
Hide engine branding, gate supported divisional columns, align house cells, and add accessible symbol help. Reduce desktop chart width and stabilize the waiting slot.

Validated with 78 browser checks and 87 focused frontend tests. Full suite retains 29 identical baseline loader failures; no new failures.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-29 19:03:28 +08:00
Jesse_ChenandClaude Opus 5.5 ea21743b09 fix(rectification): stop spoken collect once the training gate opens (BUG-1084..1087)
Once the discriminator training gate is open, only choice cards are asked
and the range card goes out when they are exhausted; targeted lines, their
re-ask and guided windows no longer hold the card or invite more events.
Delivery body says how many choice questions were used instead of the event
fit percent; narration names an excluded cluster instead of "range
unchanged"; a delivered turn no longer carries a collect question.

Offline replay (v4, 3 radii x 2 directions): truth in range 20/20 in every
cell; guided-window injections give the same width in truth and opposite
directions, so red line 1 was revised by product to truth-in-range only.
Skill 10.0.31 -> 10.0.32 (10.0.31 kept as deprecated for pinned cases).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:57 +08:00
Jesse_ChenandClaude Opus 5.5 bcf0c86feb research(rectification): typed events scored as answered probes — no benefit (BUG-1089)
v4 open holdout, ±10/±30/±60, raw and percent priors: only 4/9/14 of 57
day-precision training events split candidates; widths unchanged, top-1
drops. Narrowed year blocking (R3) mixed. All arms no_benefit; no
implementation brief recommended. Two runs byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:13 +08:00
Jesse_ChenandClaude Opus 5.5 56b51e2170 fix(rectification): name education quality events by kind, no month for year precision (BUG-1088)
Wording only: _quality_user_meaning names start/change/interruption as
升学/学业变动/学业中断; _display_date_label drops the stored month for
year-precision events. Split hash and month field unchanged; same-machine
A/B (PYTHONHASHSEED=0) differs only in user_meaning/display_date_label.
Re-frozen per ERR-110 under new quality_wording_2026_09_29 records;
sealed rerun (20) and reported-offset sweep (900) identical to 09-21.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:13 +08:00
jesse-ux e3bd3930e3 research: compare Jev intent state with the previous turn
Independent Staging Quality Gate / validate (push) Successful in 12m0s
Independent Staging Quality Gate / publish (push) Successful in 3m35s
V0 on the existing 157 real rows matches the 09-19 cache. The staging extract has no case linkage, so V1/V2 are unmeasured and the verdict stays 缺数据.
2026-09-27 11:41:53 +08:00
jesse-ux d8d03b5a69 research: consultation evidence-card inventory and draft
Independent Staging Quality Gate / validate (push) Failing after 6m40s
Independent Staging Quality Gate / publish (push) Skipped
Measure the model-visible consultation payload on three public charts, draft per-domain cards, and record the Narayana/pratyantar projection gap as BUG-1054. No runtime behavior change.
2026-09-27 01:27:57 +08:00
Jesse_ChenandClaude Opus 5.5 e53052a258 fix: restore vendored qizheng CLI and upload-artifact action deleted by 12cbe6f8
Independent Staging Quality Gate / validate (push) Successful in 12m26s
Independent Staging Quality Gate / publish (push) Successful in 6m7s
12cbe6f8 (docs-only intent) swept two pre-existing working-tree deletions in
via commit -a. Restore both byte-for-byte from e801fcf5; log ERR-112.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 20:36:08 +08:00
Jesse_ChenandClaude Opus 5.5 0f5442cea2 research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
  a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
  (two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
  identity at ±30/±60 by construction and leave six-question metrics
  unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
  pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
  gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
  (all-pairs adds no dated probes); the bottleneck is monthly evaluation.
  New finding recorded as BUG-1048 (investigating): _boundary_windows
  year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
  research docs; README board row -> 待验收. No production code, scoring,
  thresholds, gates or Skill changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:47:48 +08:00
Jesse_ChenandClaude Opus 5.5 4e6e8d87ef feat(rectification): 少问几道就出卡,交付卡以范围为主(D1–D4)
- D1 引导窗口题每个校正最多 2 道(前端 GUIDED_WINDOW_CASE_LIMIT;引擎
  GUIDED_COLLECT_LIMIT 未动:event_probes.py 属冻结评分身份,改它需重新冻结)
- D2 七条定向线与跳过线重问问完即出卡,没问到的引导窗口不再挡卡,出卡后也不再挂窗口题
- D3 卡头加副标题「最可能 HH:MM」
- D4 前两列相差 ≥5 个百分点才显示相对可能性,否则一句「这几个时刻目前区分不开……」
- 离线回放 scripts/research/fewer_probes_card_replay.py:真值不降、宽度中位 ±1 分钟、提问 11.4→7.8
- Skill 10.0.29 → 10.0.30;DESIGN / VOICE / CHANGELOG / PROGRESS / 真机清单

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:27:12 +08:00
Jesse_ChenandClaude Opus 5.5 86ff9a40d1 docs(rectification): BUG-1047 records, progress, device checklist, ledger ERR-110/111
- BUG_HISTORY BUG-1047 (investigating), PROGRESS with timing estimates, A/B
  evidence, diagnostics guide, test diffs and env gaps.
- CHANGELOG (Skill version not bumped), docs/testing checklist + headless
  screenshots, tasks README row -> 待验收, BLOCKED entry.
- pre-work ledger: ERR-110 (byte-identical engine perf changes still trip the
  frozen-scoring integrity gate), ERR-111 (Shadbala fingerprint depends on
  PYTHONHASHSEED).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 11:53:27 +08:00
jesse-uxandClaude Code 4d801e53c3 fix: integrate people archive, report reader and western chart corrections
Independent Staging Quality Gate / validate (push) Successful in 16m28s
Independent Staging Quality Gate / publish (push) Successful in 3m33s
Validate on Linux Node 22 and PostgreSQL 17: 3894 frontend tests and 64 database tests pass, with no removed test names or new failures. Preserve static Home, bounded gzip, assertion-change records and manual acceptance gaps.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-25 20:41:16 +08:00
jesse-ux 3c2f7bd559 fix(report): use the upstream reader edition for new longform bodies
New reports request reader_main from upstream origin/main 23b9609e.
KP and transit no longer copy an empty vars() dict, solar returns keep
birth_asc_sign_idx, and ordinary projection deletes internal lines whole.
2026-09-25 02:21:46 +08:00
jesse-uxandClaude Code 1420471ab1 feat: add chart waiting states and report block exports
Independent Staging Quality Gate / validate (push) Successful in 10m34s
Independent Staging Quality Gate / publish (push) Successful in 3m36s
Add bounded chart loading, per-layer retry states, and shared SVG skeletons. Add report block downloads with SVG, localized metadata, and inline deletion.

Record verification and retain CRLF export, full-build, and controlled-device acceptance blockers. User authorized staging delivery with these gaps documented.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-24 12:33:09 +08:00
jesse-uxandClaude Code 8902e48468 fix(chat): honor new-chat intent from secondary pages
Independent Staging Quality Gate / validate (push) Successful in 13m41s
Independent Staging Quality Gate / publish (push) Successful in 3m30s
Create a fresh local consultation for explicit new-chat navigation and keep reserved recovery from taking over its landing. Add regression tests and record validation gaps for remote review.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-24 00:25:28 +08:00
jesse-uxandClaude Code 15036465ee docs: record staging gate result
Record Gitea run 2857 evidence: H1-H3 pass, BUG-1011 remains the sole E2BIG failure, publish skipped, and staging is not deployed. Mark BUG-1012 through BUG-1014 resolved at the code-contract level while retaining the deployment blocker.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-23 17:18:09 +08:00
jesse-ux b85c4a686a fix(rectification): anchor candidate windows to civil dates across midnight
Independent Staging Quality Gate / validate (push) Successful in 13m27s
Independent Staging Quality Gate / publish (push) Failing after 1h0m1s
Carry explicit local date intervals instead of inferring the day from clock
order. Cluster width, delivery, adoption, and reports keep the actual civil
date; adopted date is stored separately from the reported birth_date.

Algorithm identity is scoring-9 / spec-v5. Scoring weights, confirmation
thresholds, and Skill version are unchanged. Isolated Linux final-3 gates
passed; four pre-existing Python failures remain. This is not a production
release.
2026-09-21 02:55:00 +08:00
jesse-uxandClaude Code 8d0359fc62 fix(rectification): enforce trusted result identity and preserve receipt provenance
Independent Staging Quality Gate / validate (push) Successful in 10m4s
Independent Staging Quality Gate / publish (push) Successful in 10m26s
Unify minute and block cache identity, keep unverifiable historical results read-only across server tools and write entrypoints, and aggregate completed receipt sources chronologically through a compatible function migration.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 18:03:38 +08:00
jesse-uxandClaude Code d575e89a83 fix(rectification): stop stamping pending receipts and document aggregate blocker
Implement BUG-984 F2 option A and reproduce mixed completed identities through real tools. Stop at the required SQL authorization boundary; F1/F3/F4 remain pending.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 15:55:09 +08:00
jesse-uxandClaude Code 3f39bafc4a docs: record staging authentication blocker before BUG-984
Independent Staging Quality Gate / validate (push) Successful in 12m48s
Independent Staging Quality Gate / publish (push) Successful in 14m50s
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 15:15:52 +08:00
jesse-uxandClaude Code 25232ce4ce test(rectification): compare same-day scores against legacy path in process
Replace machine-specific float hashes with strict score and matrix byte comparisons. Keep quick bridge coverage and document duplicate collection. Verify Windows/Linux float behavior and in-memory candidate-date reversal; record the separate-machine acceptance gap and BUG-984 end-to-end blocker.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 14:38:29 +08:00
jesse-uxandClaude Code 777bd53d7e merge: sync staging task sheet into cross-midnight gate fix
Preserve both the reviewed implementation and latest staging records. No production scoring changes beyond aa46da10.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 14:19:17 +08:00
jesse-uxandClaude Code aa46da1016 fix(rectification): use candidate dates for cross-midnight dasha scoring
Add date-isolated caches and regression coverage, align scoring identity, and freeze full research reruns while preserving historical artifacts. Record unresolved cache/receipt identity and end-to-end acceptance gaps for branch review only.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 13:56:11 +08:00
jesse-uxandClaude Code 539d4daee4 feat(consult): add free model-classified smalltalk fast path
Independent Staging Quality Gate / validate (push) Successful in 9m35s
Independent Staging Quality Gate / publish (push) Successful in 3m52s
Keep full consultation tool contracts unchanged. Persist short replies and refund the original reservation atomically while recording actual model usage. Verify Linux frontend 3566/3566, database 40/40, Static home and gzip +0.0493%.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 12:58:34 +08:00
jesse-uxandClaude Code 932f2fffba research(rectification): add reported-offset evaluation and frozen rerun integrity
Independent Staging Quality Gate / validate (push) Successful in 12m7s
Independent Staging Quality Gate / publish (push) Successful in 3m46s
Preserve closed confirmation gates and previously-exposed dataset boundaries. Add auditable 900-trial sensitivity results, current scorer freshness checks, and the v5 collection protocol.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 12:03:52 +08:00
jesse-uxandClaude Code 5cbc2d0097 chore(privacy): purge upstream private material and add import guards
Prepare branch for remote review; preserve documented validation gaps and protected historical packages.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 00:55:18 +08:00
jesse-uxandClaude Code 7e0037177d docs: record staging push authentication blocker
Independent Staging Quality Gate / validate (push) Failing after 9m56s
Independent Staging Quality Gate / publish (push) Skipped
Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-19 14:35:57 +08:00
jesse-ux 507ef959f5 research(jev-intent): fix2 把来源 B 现行与高置信错误补进报告
Independent Staging Quality Gate / validate (push) Failing after 9m3s
Independent Staging Quality Gate / publish (push) Skipped
离线从 cache 聚合,不重跑模型。无焦点层 Flash 69.7% 低于 Jev 78.8%。采集层相对门槛标不可判。
2026-09-19 12:28:43 +08:00
jesse-ux c2ecbc74cd research(jev-intent): 修复轮重造语料并全量对照
Independent Staging Quality Gate / publish (push) Canceled after 0s
Independent Staging Quality Gate / validate (push) Canceled after 1m37s
来源 C 改为 DeepSeek Flash 生成+独立复核,撤回模板拼接结论。来源 B 157 条人工标注后跑 Jev x2 与 Flash 全量对照,结论为缺数据。
2026-09-19 11:30:31 +08:00
jesse-ux d6fc4fb8b3 research(rectification): DeepSeek Flash 抽 33% 对照现行分类器提示词
Independent Staging Quality Gate / validate (push) Successful in 10m16s
Independent Staging Quality Gate / publish (push) Successful in 3m57s
来源 B 仍为 0。Flash 套生产提示词,采集题相对门槛未过。结论仍不可接。
2026-09-19 09:45:29 +08:00
jesse-ux fde541c2ca research(rectification): Jev 意图分类离线对照,结论不可接
Independent Staging Quality Gate / validate (push) Successful in 10m0s
Independent Staging Quality Gate / publish (push) Successful in 3m49s
来源 C 900 条 + jev-1.13.0 双跑。高置信错误 7%、点选题 answer_class 65%。不改线上分类器。
2026-09-19 09:27:10 +08:00
Jesse_ChenandClaude Fable 5.1 dc732825d8 fix(rectification): 有题就接着问,题问完才出卡;引导窗口不再硬贴领域
出卡时机只看题源有没有空:撤回「门槛达标就短路采集线」的写法,同时
按 D2 保住「题源全空就按现行规则出卡」——门槛只在还有题可问时挡住
出卡,precision_gate_met 改成只上报(新挂在决策与公开投影上),不再
单独决定时机。引导窗口题在无领域轨道上改问开放题,一个时间窗只问一
次;录入卡提交的是「YYYY 年 M 月,<领域>方面有一件事」,不再是题干
的三选一列表。记忆化 golden 只补一个新键并冻结墙钟。离线回放改成注
入真值方向的边界事件,另跑一组反方向对照。Skill 10.0.28。

BUG-747~752

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JUei7K13cYxLHE3Axe4A45
2026-09-16 12:16:55 +00:00
jesse-ux cfb41daf3d feat(rectification): 出卡加精度门槛,补经历改成系统点名
Independent Staging Quality Gate / validate (push) Failing after 6m28s
Independent Staging Quality Gate / publish (push) Skipped
宽度超过 10 分钟或头名并列时不再出交付卡,改为按大运边界逐条问、
用类型芯片和年/月选择器录入。跳过的线换问法再问一次;答「这类事
都没有过」的不再问。用户说「没有了」仍立刻给目前范围。Skill 10.0.27。

BUG-740~743
2026-09-16 18:35:27 +08:00
Jesse_ChenandClaude Fable 5 ce1939b074 docs(tasks): 星盘页与 VedAstro 外网调用脱钩 + 运行期真相两单
星盘页首屏那一发 /api/chart 没传 skip_vedastro_main_entry_overview,
本机实测冷算 0.40–0.66 秒里约 0.36 秒是 VedAstro 空转(连 endpoint 都
没配的情况下);带标志的同一调用是 5 毫秒。生产 env 开着 network 与
fanout,等于首屏同步等 24 个外部请求加 3 次领域扫描,而前端 mapper 与
contract 根本不读这份证据。BUG-718,复发自 BUG-161。

运行期单记录两条新发现:官方 vedastro==1.23.25 其实是 REST 客户端,
且 import 时请求 pypi 并 pip install --upgrade 自升级(实测 pin 装完
一 import 即变 1.23.26);无 key 时免费层排队是同步 sleep 加进程级
全局锁,24 个请求约 4.8 分钟堵住前台线程。台账补 ERR-107 / ERR-108,
生产 env 核对清单交产品负责人执行。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155nFCgCHtoA7jhSDGmZmMu
2026-09-15 15:47:32 +00:00
jesse-ux 039b0a2608 fix(rectification): say 这两分钟 only when two candidates remain (BUG-692)
Independent Staging Quality Gate / validate (push) Successful in 9m43s
Independent Staging Quality Gate / publish (push) Successful in 8m50s
Count-aware closed-pool invite, catch the five stale contract assertions, and mark M1b V1/V2 as not_measured.
2026-09-15 09:43:51 +08:00
jesse-ux dfb572fc7f docs(research): measure precision-adaptive probe gates; no variant clears the bar
Independent Staging Quality Gate / validate (push) Failing after 13m3s
Independent Staging Quality Gate / publish (push) Skipped
20 public AA cases, raman/mean. Lowering MIN_BOUNDARY_DAYS narrows ±10 from
15 to 11 minutes but drops top-1 0.80→0.75; wider radii get wider ranges.
Varga sensitivity weights (V1/V2) match production; D60 (V3) hurts ±10.
Day-precision events offset ±7 never squeeze the true minute out. Production
45/30 gate and equal varga weights stay unchanged.
2026-09-15 09:04:03 +08:00
Jesse_ChenandClaude Fable 5 62803c5d01 docs: defer tightening the domain gate until the invite ships
上游 R3 要求淘汰与定案需 ≥3 领域,本仓是 MIN_ACCEPTANCE_DOMAINS=2。产品
2026-09-14 决定先不动,等 BUG-689 的补经历邀请上线、有真机数据后再评估:那时
用户更容易补到第三个领域,收紧代价更小;现在收紧会与「永远给结果」冲突。
写进 BUG-689 单的 §4b 与研究定论页的产品决策存档,并明令执行方不得顺手改常量。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155nFCgCHtoA7jhSDGmZmMu
2026-09-14 14:46:48 +00:00
Jesse_ChenandClaude Fable 5 a643452b44 docs(research): closure page for the two rounds on unseparable candidates
把 09-14 两轮离线研究和上游对照合并成一页定论:加权重类改法全部不过门、
簇合并从未触发(假线索)、宽度按线上口径是 15/33/56 且真值 20/20 从未被挤出、
问答链有效(回放六题后头名命中 0.80/0.55/0.35)。剩下的唯一痛点是区间内部并列,
只能靠带年月的新经历或诚实呈现。附 KP 计分与风格题前置两条产品决策存档,
以及给下一个人的三条提醒(先确认指标口径、不要从 docstring 推断行为、
真值覆盖率优先于区间宽度)。ACTIVE_FRONTS 加索引。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155nFCgCHtoA7jhSDGmZmMu
2026-09-14 14:18:13 +00:00
jesse-ux cf972f4040 docs(research): measure why delivered range width equals the search window
Independent Staging Quality Gate / validate (push) Successful in 10m20s
Independent Staging Quality Gate / publish (push) Successful in 2m18s
Offline v4 holdout probe. Last round's width=window result was the
no-elimination metric; production still-valid ranges after six answers
are 15/33/56 minutes. Adjacent merge never fires under step-2 radii.
W1/W2 match baseline; W3 is uncertain after one coverage squeeze.
No production clustering or scoring defaults changed.
2026-09-14 22:11:47 +08:00
jesse-ux 2d2467dca1 docs(research): measure minute-resolution scoring; no variant clears all radii
Independent Staging Quality Gate / validate (push) Failing after 9m27s
Independent Staging Quality Gate / publish (push) Skipped
Build holdout v4 from the public AA set, correct the two v3 dates, and sweep
R1–R5 plus pairs offline. No production scoring defaults change. No
implementation brief: delivered width stays the full window on every radius.
2026-09-14 20:43:27 +08:00
Jesse_ChenandClaude Fable 5 9943a05a01 docs(research): brief for minute-resolution scoring, plus the upstream gap report
上游 b9a0ef8f..92d3a47a 的 43 个提交里生时校正零改动,没有新技法可取。候选
分不开的根因在本仓自己的打分结构:一个日精度事件里窗口内恒定的项上限 11.5
分,随分钟变化的项上限 2.125 分,约 5:1。研究单先修封存基准(v3 每例仅 3 件
事、被标 invalidated)出 v4,再离线量五个改法:分盘除数、去底座、KP 宫头子主
计分(产品 2026-09-14 拍板,推翻 BUG-325 的「不得计分」一条)、年精度事件改
边际似然、聚类签名层与计分层对齐。有收益才另立实现单。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0155nFCgCHtoA7jhSDGmZmMu
2026-09-14 11:17:33 +00:00
Jesse_ChenandClaude Opus 5 007a05fe48 docs(research): add the chance baseline the house-lord measurement was missing
Block-layer unique top-1 0.20 sits at or below the 0.35 random expectation for
2.95 candidate ascendants, so H1-H4 had no signal to amplify. Not significant at
n=20; recorded as "no positive signal", not "worse than chance".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0193vBv6w5MV2cifdTUu9H5P
2026-09-13 17:24:08 +00:00
jesse-ux 53565d0da1 docs(research): measure house-lord gochara rules with no benefit
Independent Staging Quality Gate / validate (push) Successful in 15m27s
Independent Staging Quality Gate / publish (push) Successful in 2m20s
Offline H1-H4 measurement on 20 public AA holdout cases.
Block unique top-1 did not rise; H4 made it worse; minute layer unchanged.
Leave production scoring untouched. Transits must not drive minute conclusions.
2026-09-14 01:07:11 +08:00
Jesse_ChenandCursor b063c66835 docs(research): measure dated probe supply after six answers
Independent Staging Quality Gate / validate (push) Successful in 11m55s
Independent Staging Quality Gate / publish (push) Successful in 8m53s
Offline holdout replay shows refresh-only R3+R4 add discriminative dated probes; R1/R2 do not meet the gate.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 12:59:46 +08:00
Jesse_ChenandCursor bf8ad0d1ff fix(web): keep consultation conclusions across turns and surface cache hits (BUG-555, BUG-556)
Session history was silently clipped to the first 4000 characters of the last 12 messages, so follow-ups could not see timing or audit tables. Keep an append-only tail plus a checkpoint summary, retry overflow in the same request, and expose cache hit rate in admin usage.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-06 15:16:18 +08:00