Commit Graph
28 Commits
Author SHA1 Message Date
Jesse_ChenandClaude Fable 5.1 abf2ad6aea research(rectification): scoring-method scaffold — LR weights, absence, precision follow-up pipelines on v4, verdicts pending v5 (BUG-1091)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 19:06:45 +08:00
Jesse_ChenandClaude Opus 5.5 ea21743b09 fix(rectification): stop spoken collect once the training gate opens (BUG-1084..1087)
Once the discriminator training gate is open, only choice cards are asked
and the range card goes out when they are exhausted; targeted lines, their
re-ask and guided windows no longer hold the card or invite more events.
Delivery body says how many choice questions were used instead of the event
fit percent; narration names an excluded cluster instead of "range
unchanged"; a delivered turn no longer carries a collect question.

Offline replay (v4, 3 radii x 2 directions): truth in range 20/20 in every
cell; guided-window injections give the same width in truth and opposite
directions, so red line 1 was revised by product to truth-in-range only.
Skill 10.0.31 -> 10.0.32 (10.0.31 kept as deprecated for pinned cases).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:57 +08:00
Jesse_ChenandClaude Opus 5.5 bcf0c86feb research(rectification): typed events scored as answered probes — no benefit (BUG-1089)
v4 open holdout, ±10/±30/±60, raw and percent priors: only 4/9/14 of 57
day-precision training events split candidates; widths unchanged, top-1
drops. Narrowed year blocking (R3) mixed. All arms no_benefit; no
implementation brief recommended. Two runs byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:13 +08:00
Jesse_ChenandClaude Opus 5.5 56b51e2170 fix(rectification): name education quality events by kind, no month for year precision (BUG-1088)
Wording only: _quality_user_meaning names start/change/interruption as
升学/学业变动/学业中断; _display_date_label drops the stored month for
year-precision events. Split hash and month field unchanged; same-machine
A/B (PYTHONHASHSEED=0) differs only in user_meaning/display_date_label.
Re-frozen per ERR-110 under new quality_wording_2026_09_29 records;
sealed rerun (20) and reported-offset sweep (900) identical to 09-21.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-29 10:06:13 +08:00
Jesse_ChenandClaude Opus 5.5 f6fa367f2b fix(engine): D9 double-transit targets check the D9 signs they name (BUG-1060)
`cmd_double_transit_pac` placed the `D9_{N}宫` target on the D9 lagna sign
and the `D9_{lord}(宫主)` target on that planet's D1 longitude. Both now sit
on D9 signs (D9 house-N sign; the lord's navamsa sign), built in one
testable helper. Target names, D1 and Chandra Lagna layers are unchanged
(pre-fix golden, byte-identical). Evidence-card golden regenerated with the
capture script (one public chart: double-transit conclusions 9 -> 10).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 17:30:05 +08:00
Jesse_ChenandClaude Opus 5.5 db17074430 feat(consult): evidence card v2 per the astrologer's review
TASK-consult-evidence-card-v2-20260927 T3. Card version evidence-card-v2.

- Base section (every card): the engine's D9 summary (D9 lagna, each
  planet's D9 sign and dignity, Vargottama, D1/D9 reversals); D9 houses and
  aspects stay out.
- Career: + AL (engine pada A1), 10H SAV, SAV of the Jupiter / Saturn
  transit signs.
- Marriage: + 5H / 5L, day / night, Punarphoo (observation_only), Double
  Transit on 7H / 7L (house-7 run) and DK / UL, Vivah Saham, gender unknown.
- Wealth: 8H / 12H named.
- Annual: the annual Tajika chart verbatim (parameter_sensitive), or
  年盘未接入 when the pack is blocked / not attached (never natal data);
  houses 1 + running / next AD lords' houses + the year's Jupiter / Saturn
  transit houses, with their basis; ingress / station dates only.
- Timing: Rahu / Ketu with ingress dates, Jupiter / Saturn SAV and BAV,
  Double Transit conclusions (no degrees); vargas follow the turn's other
  domain, D9 when alone.
- Chara Dasha leaves every card; KP is never on a card. The lookup enum adds
  karakamsha, dispositor_chains, inter_chart_linkage, argala, moon_transit;
  a KP lookup travels with its blocked note.
- System prompt and tool descriptions name the D9 summary, the
  parameter_sensitive / observation_only tags, 年盘未接入 and the lookups.
- Telemetry card version / agentVersion bumped to v2 (three-column notes in
  the touched tests).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 16:44:19 +08:00
Jesse_ChenandClaude Opus 5.5 d5616c347e feat(engine): expose native technique layers in the consultation output
TASK-consult-evidence-card-v2-20260927 T1. New module
scripts/consultation_native_layers.py, thinly registered after the merged
engine fields in _attach_local_consultation_layers (no handler method, no new
forgery site). It writes one new top-level chart key,
chart.consultation_native_layers, kept out of chart.modules so the thematic
report's full_reading_module_count does not move:

- d9_summary: D9 lagna, each planet's D9 sign and dignity
  (jyotish_engine._get_dignity_level), Vargottama (_calc_vargottama),
  D1<->D9 reversals
- punarphoo (punarphoo.detect_punarphoo, observation_only)
- vivah_saham (jyotish_engine._calc_vivah_saham), day_night (gulika)
- slow_transits: Jupiter / Saturn / Rahu / Ketu now with natal house and the
  Jupiter / Saturn SAV / BAV, and the next twelve months' ingress / station
  dates (ephemeris_events; nodes by the same daily-noon sampling)
- double_transit: cmd_double_transit_pac for houses 1-12, conclusions only,
  plus DK / UL targets by the same PAC rule
- karakamsha, argala, dispositor_chains, inter_chart_linkage, moon_transit
  (lookup-only layers)
- annual_tajika on the annual route: build_annual_tajika_pack (annual lagna,
  Varshesha, Muntha, Mudda Dasha with dates, Sun / Moon / year-lord Tajika
  aspects), parameter_sensitive; blocked when the pack is not reliable,
  never natal data

A/B on 3 public AA charts x 4 routes: every existing output is identical,
only the new key is added; about 40 ms per workflow. Golden regenerated with
annual and timing routes (family workflow byte-identical apart from the new
key).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 16:27:24 +08:00
jesse-ux e3bd3930e3 research: compare Jev intent state with the previous turn
Independent Staging Quality Gate / validate (push) Successful in 12m0s
Independent Staging Quality Gate / publish (push) Successful in 3m35s
V0 on the existing 157 real rows matches the 09-19 cache. The staging extract has no case linkage, so V1/V2 are unmeasured and the verdict stays 缺数据.
2026-09-27 11:41:53 +08:00
Jesse_ChenandClaude Opus 5.5 850f18605b feat(consult): the answer model reads the answer contract plus an evidence card
New frontend/src/lib/consultation-evidence-card.ts ports the research
CARD_SPECS: base section (ascendant, house signs, placements with degrees,
functional benefics/malefics with lordship, Vimshottari MD/AD/PD with dates,
Narayana md/ad/pd) plus a section per domain; values copied verbatim from
the projection or engine context, gaps listed, never filled. The tool now
returns toModelEvidenceView: status, evidence_contract (policy, blockers,
layers, limitation), claim_cards = the card, evidence_card meta (D4 line,
supplementable sections), rectification, methodology, domains; the audit
table, spectra, must_use_layers and presentation stay server-side and the
single-domain consultations copy is gone. 「本轮技法」 rows are unchanged.
Golden tests over three public AA engine captures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 75a1844cdf fix(consult): project the running Narayana period and pratyantar dates to the model (BUG-1054)
timingKeys now admits current_dasha.md/ad/pd (sign, lord, years,
start_age, end_age), remaining_years and pratyantar_dasha_timeline. Depth
and item caps unchanged. Golden regression over three public AA engine
captures asserts values, not key presence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 02:54:11 +08:00
Jesse_ChenandClaude Opus 5.5 560d4fc2dd fix(research): reuse existing handler instead of new __new__ forgeries
Independent Staging Quality Gate / validate (push) Successful in 11m55s
Independent Staging Quality Gate / publish (push) Successful in 3m28s
Gate run 2958 failed tests/test_api_server_growth_contract.py: the
evidence-card research script added two JyotishAPIHandler forgery sites.
Reuse capture_report_blocked_repairs_golden._handler(); output unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-27 01:40:47 +08:00
jesse-ux d8d03b5a69 research: consultation evidence-card inventory and draft
Independent Staging Quality Gate / validate (push) Failing after 6m40s
Independent Staging Quality Gate / publish (push) Skipped
Measure the model-visible consultation payload on three public charts, draft per-domain cards, and record the Narayana/pratyantar projection gap as BUG-1054. No runtime behavior change.
2026-09-27 01:27:57 +08:00
Jesse_ChenandClaude Opus 5.5 0f5442cea2 research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by
  a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays
  (two flips = 8 points = SEPARATION_LEAD).
- R2: weights do apply (research scorer == production at V0); V1/V2 are
  identity at ±30/±60 by construction and leave six-question metrics
  unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not
  pass the gate.
- R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day
  gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted
  (all-pairs adds no dated probes); the bottleneck is monthly evaluation.
  New finding recorded as BUG-1048 (investigating): _boundary_windows
  year-straddle exemption and positional zip misalignment bypass the gate.
- Dated errata appended (no deletions) to the 09-14/09-16 briefs and
  research docs; README board row -> 待验收. No production code, scoring,
  thresholds, gates or Skill changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:47:48 +08:00
Jesse_ChenandClaude Opus 5.5 4e6e8d87ef feat(rectification): 少问几道就出卡,交付卡以范围为主(D1–D4)
- D1 引导窗口题每个校正最多 2 道(前端 GUIDED_WINDOW_CASE_LIMIT;引擎
  GUIDED_COLLECT_LIMIT 未动:event_probes.py 属冻结评分身份,改它需重新冻结)
- D2 七条定向线与跳过线重问问完即出卡,没问到的引导窗口不再挡卡,出卡后也不再挂窗口题
- D3 卡头加副标题「最可能 HH:MM」
- D4 前两列相差 ≥5 个百分点才显示相对可能性,否则一句「这几个时刻目前区分不开……」
- 离线回放 scripts/research/fewer_probes_card_replay.py:真值不降、宽度中位 ±1 分钟、提问 11.4→7.8
- Skill 10.0.29 → 10.0.30;DESIGN / VOICE / CHANGELOG / PROGRESS / 真机清单

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
2026-09-26 14:27:12 +08:00
jesse-ux b85c4a686a fix(rectification): anchor candidate windows to civil dates across midnight
Independent Staging Quality Gate / validate (push) Successful in 13m27s
Independent Staging Quality Gate / publish (push) Failing after 1h0m1s
Carry explicit local date intervals instead of inferring the day from clock
order. Cluster width, delivery, adoption, and reports keep the actual civil
date; adopted date is stored separately from the reported birth_date.

Algorithm identity is scoring-9 / spec-v5. Scoring weights, confirmation
thresholds, and Skill version are unchanged. Isolated Linux final-3 gates
passed; four pre-existing Python failures remain. This is not a production
release.
2026-09-21 02:55:00 +08:00
jesse-uxandClaude Code aa46da1016 fix(rectification): use candidate dates for cross-midnight dasha scoring
Add date-isolated caches and regression coverage, align scoring identity, and freeze full research reruns while preserving historical artifacts. Record unresolved cache/receipt identity and end-to-end acceptance gaps for branch review only.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 13:56:11 +08:00
jesse-uxandClaude Code 932f2fffba research(rectification): add reported-offset evaluation and frozen rerun integrity
Independent Staging Quality Gate / validate (push) Successful in 12m7s
Independent Staging Quality Gate / publish (push) Successful in 3m46s
Preserve closed confirmation gates and previously-exposed dataset boundaries. Add auditable 900-trial sensitivity results, current scorer freshness checks, and the v5 collection protocol.

Co-Authored-By: Claude Code <noreply@anthropic.com>
2026-09-20 12:03:52 +08:00
jesse-ux 507ef959f5 research(jev-intent): fix2 把来源 B 现行与高置信错误补进报告
Independent Staging Quality Gate / validate (push) Failing after 9m3s
Independent Staging Quality Gate / publish (push) Skipped
离线从 cache 聚合,不重跑模型。无焦点层 Flash 69.7% 低于 Jev 78.8%。采集层相对门槛标不可判。
2026-09-19 12:28:43 +08:00
jesse-ux c2ecbc74cd research(jev-intent): 修复轮重造语料并全量对照
Independent Staging Quality Gate / publish (push) Canceled after 0s
Independent Staging Quality Gate / validate (push) Canceled after 1m37s
来源 C 改为 DeepSeek Flash 生成+独立复核,撤回模板拼接结论。来源 B 157 条人工标注后跑 Jev x2 与 Flash 全量对照,结论为缺数据。
2026-09-19 11:30:31 +08:00
jesse-ux d6fc4fb8b3 research(rectification): DeepSeek Flash 抽 33% 对照现行分类器提示词
Independent Staging Quality Gate / validate (push) Successful in 10m16s
Independent Staging Quality Gate / publish (push) Successful in 3m57s
来源 B 仍为 0。Flash 套生产提示词,采集题相对门槛未过。结论仍不可接。
2026-09-19 09:45:29 +08:00
jesse-ux fde541c2ca research(rectification): Jev 意图分类离线对照,结论不可接
Independent Staging Quality Gate / validate (push) Successful in 10m0s
Independent Staging Quality Gate / publish (push) Successful in 3m49s
来源 C 900 条 + jev-1.13.0 双跑。高置信错误 7%、点选题 answer_class 65%。不改线上分类器。
2026-09-19 09:27:10 +08:00
Jesse_ChenandClaude Fable 5.1 dc732825d8 fix(rectification): 有题就接着问,题问完才出卡;引导窗口不再硬贴领域
出卡时机只看题源有没有空:撤回「门槛达标就短路采集线」的写法,同时
按 D2 保住「题源全空就按现行规则出卡」——门槛只在还有题可问时挡住
出卡,precision_gate_met 改成只上报(新挂在决策与公开投影上),不再
单独决定时机。引导窗口题在无领域轨道上改问开放题,一个时间窗只问一
次;录入卡提交的是「YYYY 年 M 月,<领域>方面有一件事」,不再是题干
的三选一列表。记忆化 golden 只补一个新键并冻结墙钟。离线回放改成注
入真值方向的边界事件,另跑一组反方向对照。Skill 10.0.28。

BUG-747~752

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JUei7K13cYxLHE3Axe4A45
2026-09-16 12:16:55 +00:00
jesse-ux cfb41daf3d feat(rectification): 出卡加精度门槛,补经历改成系统点名
Independent Staging Quality Gate / validate (push) Failing after 6m28s
Independent Staging Quality Gate / publish (push) Skipped
宽度超过 10 分钟或头名并列时不再出交付卡,改为按大运边界逐条问、
用类型芯片和年/月选择器录入。跳过的线换问法再问一次;答「这类事
都没有过」的不再问。用户说「没有了」仍立刻给目前范围。Skill 10.0.27。

BUG-740~743
2026-09-16 18:35:27 +08:00
jesse-ux dfb572fc7f docs(research): measure precision-adaptive probe gates; no variant clears the bar
Independent Staging Quality Gate / validate (push) Failing after 13m3s
Independent Staging Quality Gate / publish (push) Skipped
20 public AA cases, raman/mean. Lowering MIN_BOUNDARY_DAYS narrows ±10 from
15 to 11 minutes but drops top-1 0.80→0.75; wider radii get wider ranges.
Varga sensitivity weights (V1/V2) match production; D60 (V3) hurts ±10.
Day-precision events offset ±7 never squeeze the true minute out. Production
45/30 gate and equal varga weights stay unchanged.
2026-09-15 09:04:03 +08:00
jesse-ux cf972f4040 docs(research): measure why delivered range width equals the search window
Independent Staging Quality Gate / validate (push) Successful in 10m20s
Independent Staging Quality Gate / publish (push) Successful in 2m18s
Offline v4 holdout probe. Last round's width=window result was the
no-elimination metric; production still-valid ranges after six answers
are 15/33/56 minutes. Adjacent merge never fires under step-2 radii.
W1/W2 match baseline; W3 is uncertain after one coverage squeeze.
No production clustering or scoring defaults changed.
2026-09-14 22:11:47 +08:00
jesse-ux 2d2467dca1 docs(research): measure minute-resolution scoring; no variant clears all radii
Independent Staging Quality Gate / validate (push) Failing after 9m27s
Independent Staging Quality Gate / publish (push) Skipped
Build holdout v4 from the public AA set, correct the two v3 dates, and sweep
R1–R5 plus pairs offline. No production scoring defaults change. No
implementation brief: delivered width stays the full window on every radius.
2026-09-14 20:43:27 +08:00
jesse-ux 53565d0da1 docs(research): measure house-lord gochara rules with no benefit
Independent Staging Quality Gate / validate (push) Successful in 15m27s
Independent Staging Quality Gate / publish (push) Successful in 2m20s
Offline H1-H4 measurement on 20 public AA holdout cases.
Block unique top-1 did not rise; H4 made it worse; minute layer unchanged.
Leave production scoring untouched. Transits must not drive minute conclusions.
2026-09-14 01:07:11 +08:00
Jesse_ChenandCursor b063c66835 docs(research): measure dated probe supply after six answers
Independent Staging Quality Gate / validate (push) Successful in 11m55s
Independent Staging Quality Gate / publish (push) Successful in 8m53s
Offline holdout replay shows refresh-only R3+R4 add discriminative dated probes; R1/R2 do not meet the gate.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 12:59:46 +08:00