research(rectification): scoring-method scaffold — LR weights, absence, precision follow-up pipelines on v4, verdicts pending v5 (BUG-1091)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
Jesse_Chen
2026-09-29 19:06:45 +08:00
co-authored by Claude Fable 5.1
parent 5f84637ea9
commit abf2ad6aea
12 changed files with 8092 additions and 1 deletions
+15
View File
@@ -14688,6 +14688,21 @@
- 复发自:无
- 修复版本:未修
## BUG-1091 | 生时校正打分只奖不罚、权重未校准、高分辨率信号未计分
- 状态:investigating(2026-09-29 研究单 `TASK-rectification-scoring-research-20260929` 脚手架已落地;判定按任务书须在 v5(≥60 例)上做,v4 上只跑通流水线、所有判定字段写 `pending_v5`)
- 首次发现 / 最近更新:2026-09-29 / 2026-09-29
- 影响面:`scripts/active_rectification_event_engine.py::_score_event`(11 条加分规则,权重手拍)、`OBSERVATION_ONLY_LAYERS`(KP 宫头子主每分钟在算但不计分)、`_fine_minute_indices`(Pranapada / Hora / Ghati / Bhava lagna 已算未计分)、计分分盘无 D60。
- 用户现象:开场说 12 件带年份经历、答约 10 张卡后范围整窗不动;引擎原始分头名簇命中在 v4 上只有 0.35 / 0.15 / 0.10(±10 / ±30 / ±60),约为随机水平的 4–5 倍。
- 触发条件:任何一次校正。打分对候选「能解释事件」加分,对「预测了没发生的事」不扣分;一条规则对真分钟和错分钟同样常命中时照样 +2。
- 根因(已确认):(1) 只奖不罚——负证据只来自选择题答「没有」(`core/build-state.ts::applyProbeOutcome`,±2);(2) 权重是人拍的,R1–R5 / V1 / V2 / V1n 都是「再拍一个权重看结果」,没有按数据校准;(3) KP / Pranapada / D60 这些分钟级信号没进计分,09-14 KP@默认权重更差是「权重没校准」的另一个症状。
- 修复:未修。本轮只做离线脚手架:`scripts/research/scoring_research_lib.py`(特征记录 provider 与线上 `score_candidates` 逐分对账、似然比权重、leave-one-case-out、加权 provider、缺席 / 精度追问辅助)、`scripts/research/scoring_research.py`(R-A / R-B / R-C 流水线)、`tests/test_scoring_research.py`。生产代码与冻结文件零改动。
- 验证:`tests/test_scoring_research.py`(含 v4 公开案例对账 0 差);v4 全集流水线跑通、`PYTHONHASHSEED=0` 两次复跑 JSON 逐字节一致(见 `docs/tasks/PROGRESS-rectification-scoring-research-20260929.md`)。
- 防复发:任何计分改动必须先在 v5 上过任务书硬红线(六格真值在区间不降、留一法数字、稳健性三项),再按 ERR-110 重冻结;不得在 v4 上下结论(BUG-1090)。
- 相关记录:BUG-560、BUG-1048、BUG-1089、BUG-1090
- 复发自:无
- 修复版本:未修(研究分支 `codex/rectification-scoring-research-20260929`)
## BUG-1092 | 报告技术页标题出现「字段」,行星主星表头是看不懂的 RL / NL / SL / SS / SB
- 状态:resolved(代码 + 回归测试;未部署,真机待合入 staging 后复核)
+1
View File
@@ -121,3 +121,4 @@ Do not open new product surfaces before at least one of these four lanes is clos
- 定论页(先读这个):`<repo>/docs/research/rectification_minute_resolution_closure_2026_09_14.md`
- 两个已被证伪的假设:加权重能拉开候选;簇合并把候选并没了(合并从未触发)。
- 未解决的只剩「区间内部并列」,目前唯一被证明有效的信息源是用户提供的带年月新经历。
- 2026-09-29 打分方法研究(脚手架,判定待 v5):`<repo>/docs/research/rectification_scoring_research_2026_09_29.md`——似然比校准权重 / 缺席证据 / 精度追问三条流水线已在 v4 上跑通,所有判定 `pending_v5`;任务书 `docs/tasks/TASK-rectification-scoring-research-20260929.md`,前置 `TASK-rectification-holdout-expansion-20260929.md`。
@@ -94,3 +94,8 @@
- 任务书和 BUG-689 邀请语依据里的「1 分钟 ≈ 1.1 天大运边界位移」有误,实测中位 3.8 天(1.3–5.9)。45 天闸门对应 8–34 分钟,不是 40 分钟。
- 新增答错容错:答错 1 题,真值在区间 98–100%;答错 2 题,±30 / ±60 上 7–10% 的回放会挤出真值(两道反答 = 8 分 = 淘汰线)。
- 全文见 `docs/research/rectification_offline_research_2026_09_26.md`。
## 9. 后续(2026-09-29 打分方法研究单,脚手架阶段)
- §1.3「加权重类改法全部不过门」说的是**人拍权重**。2026-09-29 研究单换了路线:每条规则的权重由数据估(似然比,leave-one-case-out),允许负权重(只奖不罚是结构性问题),KP / Pranapada / D60 作为特征进入由数据定权重;另测「缺席证据」与「对已说事件追问月份」。
- 脚手架与 v4 流水线见 `docs/research/rectification_scoring_research_2026_09_29.md`;**判定必须在 v5(≥60 例)上做**,v4 上的数字只是调试参考(BUG-1090 / BUG-1091)。本页 §1–§7 的结论不因此改变。
@@ -0,0 +1,139 @@
# 生时校正打分方法研究:似然比权重 / 缺席证据 / 精度追问(脚手架阶段,2026-09-29)
- 任务书:`docs/tasks/TASK-rectification-scoring-research-20260929.md`;进度:`docs/tasks/PROGRESS-rectification-scoring-research-20260929.md`
- 分支:`codex/rectification-scoring-research-20260929`,基线 `origin/staging` @ `6e393780`(开工时 head;任务书写作时为 `5febb111`,其间只有文档提交)。
- 性质:**离线测量,不改线上代码、打分、阈值、出题闸门、Skill、冻结文件。** 研究计分器是 `precision_gate_lib.score_event_with_policy`(09-14 / 09-26 已证与线上逐分相同)加上不计分的"额外观察",通过 `build_event_contribution_matrix(row_provider=…)` 走线上同一条矩阵 / 出题 / 六题回放链。
- 数据:`references/real_case_calibration/minute_rectification_holdout_v4.json`,20 例公开 Rodden-AA。**这是开放集回放,不是盲测;本页所有数字都是 v4 调试参考,任务书规定判定只能在 v5(≥60 例)上做,所以每一项的判定栏都是 `pending_v5`。**
- 口径:ayanamsa `raman`,node mode `mean`,步长 2 分钟,半径 ±10 / ±30 / ±60;六题回放(`ASK_COUNT = 6`,按真值作答,题目由线上 G0 出题器按**线上矩阵**出一次、各方案共用);交付区间 = 未淘汰且落后头名不足 8 分的簇并集。两种先验都报:`raw`(引擎行原始分,09-14 / 09-26 口径)与 `percent`(线上 `relative_support` 百分比口径;似然比方案用带温度的 softmax 后验,见 R-A)。
- 复跑:`PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir <临时目录>`(约 14 分钟)→ `docs/research/scoring_research_{lr,absence,precision_followup}_2026_09_29.json`。同机两次运行 JSON 逐字节一致(见进度记录)。
- 判门:`precision_gate_lib.gate_verdict`(头名不降**且**宽度下降、覆盖不降);JSON 里以 `debug_verdict_v4` 记录,`verdict` 字段一律 `pending_v5`。
## 结论
| 项 | 一句话结论(v4 调试参考) | 判定 |
| --- | --- | --- |
| R-0 脚手架 | 研究计分器与线上 `score_candidates` 在 20 例 × 3 档 = 60 个 case-radius 上**逐分对账 0 差**(分数与规则 id 都相同);每个 (候选, 事件, 采样日期) 记录 42 个特征(线上规则 25 个 + 额外观察 17 个) | 完成 |
| R-A 似然比权重 | 流水线跑通。v4 上:留一法 `raw` 口径头名 / 覆盖与基线持平或略动(±30 A3 头名 0.55 → 0.65、宽度 33 → 31),引擎自身头名(答题前)±10 0.30 → 0.40–0.45;`percent` 口径的 softmax 后验即使按训练集拟合温度仍过度自信,把真值挤出区间(±10 squeezed 3–5 例)。20 例上 42 个特征里 31–38 个的置信区间跨 0 | `pending_v5` |
| R-B 缺席证据 | 流水线跑通。v4 上按卡片同等强度(2 分)扣分会大量挤出真值(±60 覆盖 1.00 → 0.40)——**v4 的事件表是传记摘录,不是完整时间线,空白年不等于没发生**;真人口述同样不完整。0.5 分档不挤真值但头名下降 | `pending_v5`(v4 信号明确为负,v5 上若不翻转即关闭) |
| R-C 精度追问 | 流水线跑通。把日精度经历降成"只记得年"后,**每例平均只有 0.2 / 0.45 / 0.7 件经历在知道月份后能把候选分到两边**(±10 只有 2/20 例有);对这些例子把那一件恢复到月精度重算先验(`C1_prior`)在 ±30 上 7 例头名 0.71 vs 该 7 例的 C0,其余档位样本太少 | `pending_v5` |
## R-0 · 脚手架与对账
- `scripts/research/scoring_research_lib.py`
- `make_recording_provider`:对每个 legacy 单事件请求,用 `score_event_with_policy`(V0 基线)+ 线上的过运 / Ashtakavarga / Shadbala 辅助(取 context 里线上算好的结果)产出**与线上完全相同**的 evidence,同时把该候选在该事件下的额外观察写进 `FeatureStore`:
- 线上规则(记为特征):`vim_{md,ad,pd}_domain_{house,lord,varga}`、`{label}_functional_{benefic,malefic}_auxiliary`、`narayana_{md,ad}_domain_house`、`{label}_arudha_auxiliary`、`controlled_transit_{jupiter,saturn}_domain_house`、`ashtakavarga_target_house_{support,pressure}_auxiliary`、`shadbala_sthana_drik_naisargika_{support,pressure}_auxiliary`。`event_kind:*`、`no_domain_activation`、`*_auxiliary_not_primary` 不作特征。
- 额外观察(线上算了没计分):`kp_target_cusp_{sublord,starlord}_is_vim_{md,ad,pd}`、`kp_asc_sublord_is_vim_{md,ad,pd}`(`feature.kp_cusps.houses`,Placidus + Krishnamurti)、`vim_{md,ad,pd}_domain_varga_d60`(D60 现算)、`{pranapada,hora_lagna,ghati_lagna,bhava_lagna}_in_target_house`(`_fine_minute_indices`)。
- `reconcile_rows`:逐分比较分数与规则 id。**60/60 case-radius 差异 0。**
- `event_feature_matrix`:特征值 = 该规则在该事件采样日期上命中的比例(年精度 12 个采样日、月 3 个、日 1 个),与线上按采样日平均分数同口径。
- 真值标签:候选属于真值所在的**签名簇**(`raw_signature_clusters`)。
- 测试:`tests/test_scoring_research.py`(10 条:LR 估计、留一法不泄漏、特征矩阵、规则过滤、加权 provider、缩放答案、缺席年、精度升降、百分比、以及 v4 公开案例 ±10 对账 0 差)。
## R-A · 似然比校准权重
**方法。** 对每个特征 f,`log LR_present = log P(f | 真值簇) − log P(f | 非真值簇)`,拉普拉斯平滑 α = 1(`(Σx + 1) / (N + 2)`);`log LR_absent` 为朴素贝叶斯互补项。每档半径单独估。
| 方案 | 特征 | 权重规则 |
| --- | --- | --- |
| A1 | 只用线上 25 条规则 | 只奖不罚:`max(log LR_present, 0)`,不用 absent 项 |
| A2 | 加 17 个额外观察 | 同上 |
| A3 | 同 A2 | 朴素贝叶斯:present + absent 两项都用,允许负权重 |
候选分 = Σ 事件 Σ 特征(x·w_present + (1−x)·w_absent),经线上矩阵的事件类型系数、精度权重、采样日平均与大运边界邻近项(后者按线上原样附加,不重加权)。两个只从**训练折**估的标定量:`scale`(研究分的例内极差中位 → 线上极差中位,让 `raw` 回放的 ±2 / 8 分线意义不变)、`temperature`(Platt 单温度,最大化训练例真值簇的对数后验;网格 0.25–256),`percent` 先验 = softmax(分 / T) × 100。**留一法按案例留**(`loo_folds`),全集拟合另报只作参照。
### 六题回放(留一法;括号内为全集拟合)
`raw` 口径:
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 引擎头名(答题前) | 挤出真值 | debug 判定 |
| --- | --- | ---: | ---: | ---: | ---: | ---: | --- |
| ±10 | 基线 | 0.80 | 1.00 | 15 | 0.30 | 0 | — |
| ±10 | A1 | 0.70 (0.75) | 1.00 | 14 | 0.45 | 0 | no_benefit |
| ±10 | A2 | 0.75 (0.80) | 1.00 | 15 | 0.40 (0.50) | 0 | no_benefit |
| ±10 | A3 | 0.80 (0.85) | 1.00 | 15 | 0.40 (0.50) | 0 | no_benefit |
| ±30 | 基线 | 0.55 | 1.00 | 33 | 0.20 | 0 | — |
| ±30 | A1 | 0.55 (0.70) | 1.00 | 32 | 0.25 | 0 | benefit(宽度 −1) |
| ±30 | A2 | 0.60 (0.65) | 1.00 | 34 | 0.10 (0.30) | 0 | no_benefit |
| ±30 | A3 | 0.65 (0.75) | 1.00 | 31 | 0.20 (0.45) | 0 | benefit(宽度 −2) |
| ±60 | 基线 | 0.40 | 1.00 | 56 | 0.05 | 0 | — |
| ±60 | A1 | 0.45 (0.50) | 1.00 | 61 | 0.15 (0.20) | 0 | no_benefit |
| ±60 | A2 | 0.55 (0.60) | 1.00 | 61 | 0.00 (0.10) | 0 | no_benefit |
| ±60 | A3 | 0.50 (0.70) | **0.95** | 61 | 0.05 (0.30) | 1 | no_benefit |
`percent` 口径(基线 = 线上比例百分比;A* = 拟合温度的 softmax 后验):
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 挤出真值 | 温度(各折) |
| --- | --- | ---: | ---: | ---: | ---: | --- |
| ±10 | 基线 | 0.90 | 1.00 | 11 | 0 | — |
| ±10 | A1 / A2 / A3 | 0.60 / 0.50 / 0.55 | 0.80 / 0.85 / 0.75 | 7 / 7 / 5 | 4 / 3 / 5 | 1–2 / 1 / 0.5–1 |
| ±30 | 基线 | 0.85 | 1.00 | 33 | 0 | — |
| ±30 | A1 / A2 / A3 | 0.65 / 0.55 / 0.55 | 0.95 / 0.80 / 0.90 | 31 / 10 / 13 | 1 / 4 / 2 | 2–4 / 1–2 / 1–2 |
| ±60 | 基线 | 0.80 | 1.00 | 53 | 0 | — |
| ±60 | A1 / A2 / A3 | 0.65 / 0.50 / 0.40 | 1.00 / 0.80 / 0.80 | 57 / 32 / 38 | 0 / 4 / 4 | 4–8 / 2–4 / 2 |
**读法(v4,仅供 v5 前设计参考,不是结论)。**
1. 对账通过后,`raw` 口径的留一法数字与基线基本持平:似然比权重没有让真值掉出区间(±60 A3 一例除外),也没有明显收窄;引擎答题前的头名在 ±10 从 0.30 到 0.40–0.45、±60 A1 从 0.05 到 0.15,全集拟合再高一些——**全集与留一法的差距就是 20 例过拟合的量**。
2. `percent` 口径下 softmax 后验**过度自信**:即使温度按训练折拟合,仍把 15–25% 例子的真值挤出区间。校准曲线(A3 留一法)在 ±10 上 [0.20,0.30) 桶预测 0.24 / 实际 0.40、[0.50,0.70) 预测 0.58 / 实际 0.40、[0.70,1.00) 只有 1 个点且错;±30 / ±60 高桶几乎无样本。要上线必须先解决后验的标定,而不是先验本身。
3. **42 个特征里 31 / 36 / 38 个(±10 / ±30 / ±60)的 5–95% 置信区间跨 0**(案例级 bootstrap 200 次)。±10 上区间不跨 0 的:`ashtakavarga_target_house_pressure_auxiliary`(+0.41)、`vim_ad_domain_varga`(+0.30)、`vim_md_domain_varga`(+0.27)、`ghati_lagna_in_target_house` / `hora_lagna_in_target_house`(+0.31)、`pranapada_in_target_house`(**−0.36**)、`vim_ad_arudha_auxiliary`(−0.21)、`narayana_ad_domain_house`(−0.21)、`kp_target_cusp_sublord_is_vim_ad`(−0.19)。方向与线上手拍权重相反的几条(Arudha、Narayana AD、KP AD 子主为负)正是 20 例上说不清的地方——这就是任务书要先扩集的原因。
4. 稳健性(留一法 A3,`raw`):±7 天抖动与 ±1–3 月平移三档真值在区间 0.95–1.00;答错 1 题 0.96–1.00、答错 2 题 0.91–0.99(与 09-26 基线同量级)。`percent` 口径下三项都更差(0.70–0.90),同样是后验过自信的表现。
## R-B · 缺席证据
**方法。** 只用训练事件(留一件 holdout 不用)。每个领域取第一件与最后一件带年月事件之间的年份,没有该领域事件的年份 = 缺席年。对每个缺席 (领域, 年) 用出题器同一判定函数 `_evaluate_contexts(month=None)` 在代表候选上求"这一年该领域是否激活"的分边;分得开就当作一道答了"没有"的题,按 0.5 / 1.0 / 2.0 分扣(卡片是 2.0;低于卡片强度的扣分不计强冲突、不会淘汰)。之后照常六题。漏说稳健性:随机删掉 1–2 件训练事件再算缺席年(每例 5 次,种子固定)。
`raw` 口径六题回放(`percent` 口径同向,见 JSON):
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 挤出真值 | 漏说 1 件(100 次回放)真值在区间 | 漏说 2 件 |
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
| ±10 | B0 | 0.80 | 1.00 | 15 | 0 | — | — |
| ±10 | 缺席 @0.5 分 | 0.70 | 1.00 | 15 | 0 | 1.00 | 1.00 |
| ±10 | 缺席 @1.0 分 | 0.70 | 1.00 | 14 | 0 | 1.00 | 1.00 |
| ±10 | 缺席 @2.0 分(=卡片) | 0.55 | **0.85** | 13 | 3 | 0.94 | 0.94 |
| ±30 | B0 | 0.55 | 1.00 | 33 | 0 | — | — |
| ±30 | 缺席 @0.5 分 | 0.50 | 1.00 | 34 | 0 | 1.00 | 1.00 |
| ±30 | 缺席 @1.0 分 | 0.45 | 1.00 | 33 | 0 | 1.00 | 1.00 |
| ±30 | 缺席 @2.0 分 | 0.25 | **0.70** | 19 | 6 | 0.80 | 0.87 |
| ±60 | B0 | 0.40 | 1.00 | 56 | 0 | — | — |
| ±60 | 缺席 @0.5 分 | 0.35 | 1.00 | 68 | 0 | 1.00 | 1.00 |
| ±60 | 缺席 @1.0 分 | 0.25 | 1.00 | 61 | 0 | 1.00 | 1.00 |
| ±60 | 缺席 @2.0 分 | 0.10 | **0.40** | 9 | 12 | 0.54 | 0.73 |
每例平均缺席年 17.95;能分边的缺席题 2.05 / 4.30 / 6.15 道(有题的例子 12 / 17 / 17)。三档 debug 判定全部 no_benefit。
读法:
1. **按卡片强度扣,真值被大量挤出**(±60 覆盖 1.00 → 0.40)。原因不是方法错,而是 v4 的事件表是传记摘录:那些"空白年"里公众人物多半确有该领域的事,只是没进数据集;扣分等于用不完整的时间线否定真值。真人口述同样不完整(用户不会把每年的事都说全),所以这条在 v5 上大概率也是负的——**若 v5 上不翻转,R-B 关闭**。
2. 0.5 / 1.0 分档不挤真值(覆盖 1.00)但头名一律下降、宽度不降,等于只加噪声。漏说 1–2 件在低强度下不伤覆盖,在卡片强度下把覆盖压到 0.54–0.94。
## R-C · 精度追问可行性
**方法。** 把每例所有日精度经历降成"只记得年"(v4 没有月精度)作为口述形态 `C0_spoken`,先验按线上重算。对每件训练中的日精度经历,用 `_evaluate_contexts(month=真实月份)` 判断"知道月份后能否把代表候选分到两边";能分的按信息增益排序取前 1 / 2 件:`C{k}_prior`(那件恢复到月精度重算先验)、`C{k}_probe`(当作答"是"的卡)、`C{k}_both`。
| 半径 | 每例可追问件数均值 | ≥1 件的例子 | ≥2 件的例子 |
| --- | ---: | ---: | ---: |
| ±10 | 0.20 | 2 / 20 | 1 / 20 |
| ±30 | 0.45 | 7 / 20 | 1 / 20 |
| ±60 | 0.70 | 10 / 20 | 3 / 20 |
| 半径 | 方案 | n | 头名命中(raw) | 真值在区间 | 宽度中位 | 同 n 例的 C0(raw 头名 / 宽度) |
| --- | --- | ---: | ---: | ---: | ---: | --- |
| ±10 | C1_prior / C1_probe | 2 | 1.00 / 0.50 | 1.00 | 16 | 见 JSON per-arm(样本太少) |
| ±30 | C1_prior | 7 | 0.71 | 1.00 | 31 | debug 判定 benefit(percent 口径头名 1.00) |
| ±30 | C1_probe | 7 | 0.57 | 1.00 | 33 | no_benefit |
| ±60 | C1_prior / C1_probe | 10 | 0.40 / 0.40 | 1.00 | 58 / 63 | no_benefit |
**读法。** 可追问的经历很少(与 09-29 打字经历研究一致:绝大多数经历落在所有候选都一样的大运段),但在 ±30 上"恢复到月精度重算先验"(走线上已有的打字经历通道,不加卡片)对那 7 例有正向信号;"当作答是的卡"没有。样本太少,只能说值得在 v5 上按同口径重测;产品层实现(只对落在候选分歧点的那一件追问月份,算在定向追问 2 条额度内)要等 v5 判定。
## 给产品的话
- 三条流水线都跑通、与线上对账 0 差,**但没有任何一条在 v4 上能下结论**,20 例上 42 个特征有四分之三分不清正负。这一轮交付的是"能在 v5 上一键复跑的测量工具",不是新打分。
- v4 上唯一方向明确的信号是负的:把口述时间线的空白当"没发生"会伤真值(R-B)。
- v5 交付后的复跑命令:`PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --holdout references/real_case_calibration/minute_rectification_holdout_v5.json --dataset-label v5 --out-suffix v5_<日期>`。
## 附:复跑命令
```bash
PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir /tmp/scoring_ctx # 约 14 分钟,写三份 JSON
PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --limit 3 --radii 10 --fast --no-write # 开发用
python3 -m pytest tests/test_scoring_research.py -q
```
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,66 @@
# PROGRESS · 生时校正打分方法研究:似然比权重 / 缺席证据 / 精度追问(脚手架阶段,2026-09-29)
- 任务书:`docs/tasks/TASK-rectification-scoring-research-20260929.md`
- 执行:Claude(直接执行模式,fork 子代理),分支 `codex/rectification-scoring-research-20260929`,基线 `origin/staging` @ `6e393780`(任务书写作时 `5febb111`,其间只有文档提交)。**未推送。**
- 研究报告:`docs/research/rectification_scoring_research_2026_09_29.md` + `scoring_research_{lr,absence,precision_followup}_2026_09_29.json`
- 本轮范围(按调度方指令):在 v4 上完成 R-0 脚手架并把 R-A / R-B / R-C 流水线跑通;**所有判定字段写 `pending_v5`**,v4 数字只作调试参考。v5 交付后在 v5 上跑判定(复跑命令见研究报告末尾)。
## 开工
- `git fetch origin --prune` → worktree `.worktrees/rectification-scoring-research-20260929` 基于 `origin/staging`;任何 git 操作前 `git status -sb` 核对分支。
- 先读:定论页(含 §7 勘误)、`rectification_offline_research_2026_09_26.md`、`rectification_typed_event_scoring_2026_09_29.md`、ERR-110 / ERR-111。
- 环境:本机没有 `.venv`(所有 worktree 的 `.venv` 软链都指向不存在的 `/workspace/Jyotisha/.venv`),全部用系统 `python3` 3.13.5(swisseph 9.1.1、pytest、timezonefinder 齐)。快速门的前端步需要 Node 22:`frontend/node_modules` 软链到主检出,`npm` 用 `/exec-daemon/node` + `/exec-daemon/lib/node_modules/npm` 的 shim(系统 `node` 是 20.19,`/usr/bin/npm` 跑的 `tsx` 不在 PATH)。`PYTHONHASHSEED=0`(ERR-111)。
- 另一子代理并行做 v5 扩集(`.worktrees/rectification-holdout-expansion-20260929`),未触碰。
## 交付物
| 文件 | 内容 |
| --- | --- |
| `scripts/research/scoring_research_lib.py` | R-0:特征记录 provider(线上同源 evidence + 额外观察)、逐分对账、特征矩阵、似然比估计(α = 1)、案例级 bootstrap、加权 provider、leave-one-case-out、缩放答案、Platt 温度、缺席年 / 精度升降 / 月份平移辅助 |
| `scripts/research/scoring_research.py` | R-A / R-B / R-C 流水线;`--limit / --radii / --parts / --fast / --cache-dir / --holdout / --dataset-label / --out-suffix` |
| `tests/test_scoring_research.py` | 10 条:纯函数 9 条 + v4 公开案例 ±10 对账 0 差 1 条 |
| `docs/research/rectification_scoring_research_2026_09_29.md` | 报告(结构按任务书 R-D;结论栏全部 `pending_v5`;复跑命令) |
| `docs/research/scoring_research_*_2026_09_29.json` | 三份结果(`dataset: v4, verdict: pending_v5`) |
| `docs/BUG_HISTORY.md` | BUG-1091 `investigating`(开工核对最大号 1089;1090 由扩集单占用,未冲突) |
| `docs/research/ACTIVE_FRONTS.md`、定论页 §9、`docs/tasks/README.md` | 索引与状态板 |
生产代码、`sealed_holdout_rerun.py::PRODUCTION_FILES` 的 10 个冻结文件、任何常数:**零改动**(`git status` 只有上表文件)。
## R-0 · 对账(任务书硬红线 4)
- 研究计分器 vs 线上 `score_candidates`(经 `build_event_contribution_matrix` 同一条矩阵链):**20 例 × 3 档 = 60 个 case-radius,分数差 0、规则 id 差 0**(`reconciliation.unreconciled = 0`,三份 JSON 均记录)。
- 做法:provider 用 `precision_gate_lib.score_event_with_policy`(V0)+ 线上 `_controlled_transit_rules`(带缓存)/ `_ashtakavarga_auxiliary` / `_shadbala_verified_components_auxiliary`(取 context 里线上算好的结果),返回与 `_candidate_row` 相同的 evidence;额外观察只进 `FeatureStore`,不进 rule_ids / points,所以矩阵、`event_kind` 系数、`technique_layers`、出题都与线上一致。
- 特征 42 个:线上规则 25 + 额外观察 17(KP 宫头子主 / 星主 × MD/AD/PD、KP 上升子主、D60 × MD/AD/PD、Pranapada / Hora / Ghati / Bhava lagna 落领域宫)。特征值 = 该规则在该事件采样日期上命中的比例。真值标签 = 候选属于真值所在签名簇。
- 留一法按**案例**留(`loo_folds`),`scale` 与 `temperature` 也只从训练折估。
## R-A / R-B / R-C · 流水线状态
| 项 | 状态 | v4 调试参考(`raw` 口径,留一法) | 判定 |
| --- | --- | --- | --- |
| R-A A1 / A2 / A3 | 跑通(留一法 + 全集、六题回放、引擎头名、校准曲线、稳健性四项、LR 表 + 案例级 bootstrap 区间) | 头名 ±10 0.70 / 0.75 / 0.80(基线 0.80);±30 0.55 / 0.60 / 0.65(0.55);±60 0.45 / 0.55 / 0.50(0.40)。真值在区间除 ±60 A3 0.95 外全 1.00。`percent` 口径 softmax 后验过自信,挤出 3–5 例 | `pending_v5` |
| R-B 缺席 @0.5 / 1.0 / 2.0 分 | 跑通(含漏说 1–2 件 × 5 次) | 2.0 分(=卡片)把覆盖压到 0.85 / 0.70 / 0.40;0.5 / 1.0 分覆盖 1.00 但头名一律降 | `pending_v5`(v4 信号为负) |
| R-C C1 / C2 × prior / probe / both | 跑通 | 可追问件数每例均值 0.20 / 0.45 / 0.70;±30 上 C1_prior 7 例头名 0.71 | `pending_v5`(样本太少) |
各特征在 v4 上的 LR 与「LR≈1」名单在报告 R-A 第 3 条与 JSON `radii.*.lr_table` / `noise_features`:**42 个特征里 31 / 36 / 38 个(±10 / ±30 / ±60)的 5–95% 区间跨 0**。
## 复跑与确定性(任务书硬红线 5)
- `PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir <scratch>` 连跑两次(758 s / 748 s,静态 context 走磁盘缓存),三份 JSON `cmp` **逐字节一致**;`--out-suffix rerun` 的副本比对后删除,不入库。
- JSON 不写耗时字段,`sort_keys=True`;随机(漏说抽样、月份平移、答错抽样、bootstrap)全部按 `f"{SEED}:{case_id}:{radius}:…"` 种子。
## 快速门
- 开工基线(`origin/staging` @ `6e393780`,未改任何文件)与改后各跑一次 `python3 scripts/run_quality_gate.py --profile quick`(PATH 带 Node 22 shim):
- Python:**1001 passed, 1 skipped** 两次相同(改后含新增的 10 条 `test_scoring_research.py`——门禁的 pytest 集合按清单收集,新文件不在清单里,计数没变;定向 `python3 -m pytest tests/test_scoring_research.py -q` 10 passed)。
- 前端 `npm test`:基线 4357 / pass 4301 / fail 25;改后 4357 / pass **4302** / fail **24**、cancelled 0。失败清单逐条比对:改后 24 条是基线 25 条的子集(少的一条 `D3 · the first stage line shows on send…` 是基线那次偶发红、改后绿,前端零改动);24 条全是需要 Docker / PostgreSQL / 部署环境的套件(Better Auth / billing / migration / staging sync 等),与 09-29 上一单验收记录的 24 条环境失败一致。
- 门禁整体 exit 1 = 该 24 条环境失败所致,基线同样 exit 1;**无新增失败**。
- 隐私:`tests/test_repo_privacy_markers.py` 在快速门 Python 步内通过(1001 passed 含之)。
## 缺口与说明
- **判定未做**:按任务书,R-A / R-B / R-C 的有无收益只能在 v5 上判;本轮 JSON 的 `debug_verdict_v4` 只是 `gate_verdict` 的机械输出,不作结论。
- R-A 的 `percent` 口径需要先解决后验标定(单温度 Platt 不够,校准曲线高桶过自信);v5 上若仍过自信,考虑按训练折做等分位标定或直接用 `raw` 口径(缩放到线上极差)。
- 六道题由线上矩阵出一次、各方案共用(隔离先验的影响);若某方案在 v5 上过门,实现前还应按方案矩阵重新出题再测一次。
- 大运边界邻近项(`merge_transition_proximity`)按线上原样附加、未重加权。
- R-B 的缺席年在 v4 上不可靠(传记摘录不是完整时间线);这条限制在真人口述上同样存在,报告已写明。
- 未做:R-C 的产品层实现、任何实现单;`CHANGELOG.md` 无用户可见变化,未改。
+1 -1
View File
@@ -256,7 +256,7 @@
| `TASK-mobile-chart-and-confirmed-edit-20260929.md` | `PROGRESS-mobile-chart-confirmed-edit-20260929.md` | **手机星盘 + 确认时间可改 + 天空定格**:iPhone 星盘被宽表撑出屏幕(BUG-1083);confirmed 状态改声明字段也重置排盘时间(推翻 BUG-264 约定);「那一刻的天空」改为重放汇聚→定格→右上角分享;去掉报告星图封面 | 已验收(Claude 2026-09-29),已推 staging | BUG-1083 |
| `TASK-rectification-futile-collect-stop-20260929.md` | `PROGRESS-rectification-futile-collect-stop-20260929.md` | **生时校正停掉无效补经历循环(止血)**:打字经历不收窄(BUG-560 后果),流程却一路索要,真机 22 件整窗不动;采集只为开闸、门开后只问点选卡问完即出卡;交付正文去吻合率;多段时旁白误报「范围没变」;交付轮带采集题 | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1084~1087 |
| `TASK-rectification-holdout-expansion-20260929.md` | — | **开放评价集扩到 ≥60 例(v5)**:所有打分判定都在 20 例上做,1 例 = 5pp,KP / 精度闸 / V1n 的 no_benefit 都只差 1–3 例。只用 Rodden AA 公开名人,事件人工核对出处,分层(年代 / 纬度 / 南半球 / 跨午夜 / UTC+8),v4 不动;基线成绩单 v4 子集须与已发布数字逐格一致。研究单的判定以 v5 为准 | 待领取 | — |
| `TASK-rectification-scoring-research-20260929.md` | — | **打分方法研究(离线)**:R-A 似然比校准权重(leave-one-case-out,KP / Pranapada / D60 作为特征由数据定权重、允许负权重)、R-B 缺席证据(口述时间线覆盖时段内的空白年扣分,必测漏说稳健性)、R-C 精度追问可行性(对已说事件追问月份,算在定向追问 2 条额度内)。判定必须在 v5 上做;有收益才立实现单并按 ERR-110 重冻结 | 待领取(脚手架可先在 v4 上开发) | — |
| `TASK-rectification-scoring-research-20260929.md` | `PROGRESS-rectification-scoring-research-20260929.md` | **打分方法研究(离线)**:R-A 似然比校准权重(leave-one-case-out,KP / Pranapada / D60 作为特征由数据定权重、允许负权重)、R-B 缺席证据(口述时间线覆盖时段内的空白年扣分,必测漏说稳健性)、R-C 精度追问可行性(对已说事件追问月份,算在定向追问 2 条额度内)。判定必须在 v5 上做;有收益才立实现单并按 ERR-110 重冻结 | 已实现待验收(Claude fork 2026-09-29:R-0 脚手架 + R-A / R-B / R-C 流水线在 v4 上跑通,对账 0 差,判定全部 `pending_v5`;生产代码零改动) | 分支 `codex/rectification-scoring-research-20260929`(BUG-1091 investigating) |
| `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1088、1089 |
| `TASK-report-reader-polish-20260929.md` | `PROGRESS-report-reader-polish-20260929.md` | **报告页打磨**:生成入口挪进页面主体(删标题栏按钮)、详情页到底部按钮(懒渲染一次到底)、导出按钮带文字、目录一级/二级分层、去掉「字段」与 RL/NL 缩写表头、状态列同义重复去重、报告表格淡底色、分块导出显示真实文件大小 | Claude 直接执行并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +7 B、快速门 Python 1000 passed),已推 staging,真机欠 | BUG-1092~1094 |
| `TASK-report-english-edition-20260929.md` | — | **报告中英两版**:同一次引擎计算渲染 zh/en 两遍(不用模型翻译),瑜伽库 406 条补英文、模板与前端表头英文化;中文逐字节不变、英文零汉字、两版数字序列一致;阅读页 `?lang=en` 切换,导出当前语言 + 「问 AI 建议导出英文版」提示;旧报告不补英文 | 待领取(依赖 reader-polish 合入) | — |
+529
View File
@@ -0,0 +1,529 @@
#!/usr/bin/env python3
"""Offline research runner for TASK-rectification-scoring-research-20260929.
R-0 research scorer reconciled per minute against the production scorer, then
per-(candidate, event) rule features recorded (eleven production rules,
auxiliary rules, and unscored extras: KP cusp sub/star lords, Pranapada /
Hora / Ghati / Bhava lagna, D60).
R-A likelihood-ratio weights (Laplace α = 1) per radius, leave-one-CASE-out
and full-fit; variants A1 (production features, reward only), A2 (+extras,
reward only), A3 (+extras, naive-Bayes present/absent, negative allowed);
six-probe replay, engine top1, calibration curve, robustness (±7-day and
±1–3-month date shifts, 1–2 flipped answers).
R-B absence evidence: gap years inside a domain's spoken span answered "no"
at 0.5 / 1.0 / 2.0 points (a card is 2.0), plus "user omitted 1–2 events"
robustness.
R-C precision follow-up: with all day events degraded to year ("spoken"
form), how many events would split candidates if the month were known,
and the gain from asking 1 / 2 of them (as a prior recompute and as an
answered probe).
Every verdict field is written as ``pending_v5``: the task brief allows a
verdict only on the v5 open set. v4 numbers are debugging references.
Run:
PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --limit 2 --radii 10 --no-write
PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py
"""
from __future__ import annotations
import argparse
import json
import pickle
import sys
import time as clock_module
import traceback
from collections import defaultdict
from pathlib import Path
from random import Random
from typing import Any, Sequence
ROOT = Path(__file__).resolve().parents[2]
if str(ROOT) not in sys.path:
sys.path.insert(0, str(ROOT))
from scripts.active_rectification_event_engine import AYANAMSA, NODE_MODE, compute_candidate_static_contexts # noqa: E402
from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix, scoreable_request # noqa: E402
from scripts.research.minute_resolution_sweep import MINUTE_STEP, scoring_request_for # noqa: E402
from scripts.research.offline_research_20260926_lib import exhaustive_flip_sets, flipped_answers, sample_flip_sets # noqa: E402
from scripts.research.precision_gate_lib import gate_verdict, jitter_day_events # noqa: E402
from scripts.research.probe_supply_after_six import ASK_COUNT, optimal_answer # noqa: E402
import scripts.research.scoring_research_lib as lib # noqa: E402
HOLDOUT_V4 = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v4.json"
OUT_DIR = ROOT / "docs" / "research"
RADII = (10, 30, 60)
PRIORS = ("raw", "percent")
ABSENCE_POINTS = (0.5, 1.0, 2.0) # score points per "no"; a card is 2.0
CARD_POINTS = 2.0
OMIT_REPEATS = 5
SHIFT_REPEATS = 3
FLIP_REPEATS = 10
SEED = 20260929
ABSENCE_SOURCE = "absence_research"
FOLLOWUP_SOURCE = "precision_followup_research"
# ---------------------------------------------------------------------------
# per-case data
# ---------------------------------------------------------------------------
class CaseData:
def __init__(self, case: dict[str, Any], radius: int, cache_dir: Path | None) -> None:
self.case = case
self.case_id = str(case["case_id"])
self.radius = radius
self.true_time = str(case["birth"]["time"])[:5]
self.request = scoring_request_for({**case, "candidate_radius_minutes": radius}, radius)
self.contexts = self._contexts(cache_dir)
self.times = [ctx["candidate_at"].strftime("%H:%M") for ctx in self.contexts]
self.truth_times = lib.truth_cluster_times(self.contexts, self.true_time)
self.d60 = lib.d60_charts(self.contexts)
# production
prod_built = build_event_contribution_matrix(self.request, static_contexts=self.contexts)
self.prod_rows = score_from_matrix(self.request, prod_built)
self.prod_built = prod_built
self.probes = lib.production_probes(self.request, prod_built, self.times, self.true_time)
self.prod_public = lib.public_for(self.prod_rows, self.contexts)
# research recording pass (must reconcile to production)
self.store = lib.FeatureStore()
built = build_event_contribution_matrix(
self.request, row_provider=lib.make_recording_provider(self.contexts, self.store, self.d60),
static_contexts=self.contexts,
)
rows = score_from_matrix(self.request, built)
self.reconciliation = lib.reconcile_rows(self.prod_rows, rows)
self.training_ids = lib.training_event_ids(self.request)
self.representatives = lib.representatives_for(self.contexts)
self.events_by_id = {str(e["id"]): e for e in self.request["events"]}
def _contexts(self, cache_dir: Path | None) -> list[dict[str, Any]]:
if cache_dir is not None:
path = cache_dir / f"{self.case_id}_{self.radius}.pkl"
if path.exists():
with path.open("rb") as handle:
return pickle.load(handle)
contexts = compute_candidate_static_contexts(self.request)
if cache_dir is not None:
cache_dir.mkdir(parents=True, exist_ok=True)
with (cache_dir / f"{self.case_id}_{self.radius}.pkl").open("wb") as handle:
pickle.dump(contexts, handle)
return contexts
# features -------------------------------------------------------------
def examples(self, *, include_extras: bool) -> list[tuple[dict[str, float], bool]]:
out: list[tuple[dict[str, float], bool]] = []
truth = set(self.truth_times)
for event_id in self.training_ids:
matrix = lib.event_feature_matrix(self.store, event_id, self.times, include_extras=include_extras)
for time in self.times:
out.append((matrix.get(time, {}), time in truth))
return out
# scoring with a weight table ------------------------------------------
def weighted_rows(
self, table: lib.LRTable, variant: str, scale: float,
store: lib.FeatureStore | None = None, request: dict[str, Any] | None = None,
) -> tuple[list[dict[str, Any]], dict[str, Any]]:
request = request or self.request
built = build_event_contribution_matrix(
request,
row_provider=lib.make_weighted_provider(store or self.store, table, variant=variant, scale=scale),
static_contexts=self.contexts,
)
return score_from_matrix(request, built), built
def rescored_store(self, events: Sequence[dict[str, Any]]) -> tuple[lib.FeatureStore, dict[str, Any], list[dict[str, Any]]]:
"""Re-record features for an altered event list (date shifts): (store, request, production rows)."""
request = scoring_request_for({**self.case, "events": list(events), "candidate_radius_minutes": self.radius}, self.radius)
store = lib.FeatureStore()
built = build_event_contribution_matrix(
request, row_provider=lib.make_recording_provider(self.contexts, store, self.d60), static_contexts=self.contexts,
)
return store, request, score_from_matrix(request, built)
def production_rows_for(self, events: Sequence[dict[str, Any]]) -> list[dict[str, Any]]:
request = scoring_request_for({**self.case, "events": list(events), "candidate_radius_minutes": self.radius}, self.radius)
built = build_event_contribution_matrix(request, static_contexts=self.contexts)
return score_from_matrix(request, built)
def evaluate(data: CaseData, rows: Sequence[dict[str, Any]], *, softmax_prior: bool, pre_answers=(), answers=None, temperature: float = 1.0) -> dict[str, dict[str, Any]]:
public = lib.public_for(rows, data.contexts)
out: dict[str, dict[str, Any]] = {}
for mode in PRIORS:
prior = lib.prior_for(public, mode=mode, softmax=softmax_prior, temperature=temperature)
result = lib.replay(
public=public, prior=prior, probes=data.probes, true_time=data.true_time,
window_times=data.times, pre_answers=pre_answers, answers=answers,
)
result["engine_top1"] = lib.engine_top1(public, data.true_time)
out[mode] = result
return out
# ---------------------------------------------------------------------------
# R-A
# ---------------------------------------------------------------------------
def fit_prior(training: Sequence[CaseData], table: lib.LRTable, variant: str) -> tuple[float, float]:
"""(scale, temperature) from training cases only.
scale: research within-case score range matched to the production median
range (so the raw-prior replay's ±2 / 8-lead constants mean the same);
temperature: single Platt temperature for the percent (softmax) prior.
"""
prod = [lib.within_case_range(d.prod_rows) for d in training]
rows_by_case = {d.case_id: d.weighted_rows(table, variant, 1.0)[0] for d in training}
research = [lib.within_case_range(rows) for rows in rows_by_case.values()]
p_med, r_med = lib.median(prod), lib.median(research)
scale = round(p_med / r_med, 6) if p_med and r_med else 1.0
publics = [
(lib.public_for([{**r, "score": float(r["score"]) * scale} for r in rows_by_case[d.case_id]], d.contexts), d.truth_times)
for d in training
]
return scale, lib.fit_temperature(publics)
def run_ra(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]:
out: dict[str, Any] = {"variants": list(lib.VARIANTS), "alpha": lib.LAPLACE_ALPHA, "radii": {}}
for radius, cases in sorted(by_radius.items()):
block: dict[str, Any] = {"n": len(cases), "baseline": {}, "full_fit": {}, "loo": {}, "robustness": {}, "calibration": {}, "lr_table": {}, "noise_features": {}}
baseline_rows = []
for data in cases:
row = evaluate(data, data.prod_rows, softmax_prior=False)
baseline_rows.append({"case_id": data.case_id, **{f"{m}": row[m] for m in PRIORS}})
block["baseline"] = {m: lib.summarize([r[m] for r in baseline_rows]) for m in PRIORS}
per_case_examples = {
include: {d.case_id: d.examples(include_extras=include) for d in cases} for include in (False, True)
}
altered_stores: dict[tuple[str, str, int], tuple[lib.FeatureStore, dict[str, Any]]] = {}
for variant in lib.VARIANTS:
include_extras, _neg = lib.VARIANT_SPEC[variant]
examples = per_case_examples[include_extras]
# full fit -------------------------------------------------------
full_table = lib.estimate_lr([ex for rows in examples.values() for ex in rows])
full_scale, full_temperature = fit_prior(cases, full_table, variant)
full_rows = []
for data in cases:
rows, _ = data.weighted_rows(full_table, variant, full_scale)
full_rows.append(evaluate(data, rows, softmax_prior=True, temperature=full_temperature))
block["full_fit"][variant] = {
"scale": full_scale,
"temperature": full_temperature,
**{m: lib.summarize([r[m] for r in full_rows]) for m in PRIORS},
}
if variant == "A3":
block["lr_table"] = {
name: {
"log_lr_present": full_table.present[name],
"log_lr_absent": full_table.absent[name],
**full_table.support[name],
"extra": lib.is_extra_feature(name),
}
for name in sorted(full_table.present)
}
ci = lib.bootstrap_lr(examples, resamples=50 if fast else 200)
for name, interval in ci.items():
block["lr_table"].setdefault(name, {}).update(interval)
block["noise_features"] = sorted(
name for name, row in block["lr_table"].items()
if row.get("p05") is not None and row["p05"] <= 0.0 <= row["p95"]
)
# leave-one-case-out -------------------------------------------
loo_rows = []
robustness = defaultdict(list)
calibration_points: list[tuple[float, bool]] = []
temperatures: list[float] = []
for held in cases:
training = [d for d in cases if d.case_id != held.case_id]
table = lib.estimate_lr([ex for d in training for ex in examples[d.case_id]])
scale, temperature = fit_prior(training, table, variant)
temperatures.append(temperature)
rows, _ = held.weighted_rows(table, variant, scale)
result = evaluate(held, rows, softmax_prior=True, temperature=temperature)
loo_rows.append(result)
if variant == "A3":
public = lib.public_for(rows, held.contexts)
posterior = lib.percent_softmax({str(r["time"])[:5]: float(r["score"]) for r in public}, temperature)
truth = set(held.truth_times)
for rep in public:
stamp = str(rep["time"])[:5]
members = {str(t)[:5] for t in (rep.get("cluster_times") or [stamp])}
calibration_points.append((posterior.get(stamp, 0.0) / 100.0, bool(members & truth)))
# robustness: date shifts (features re-recorded with fold weights)
if not fast or variant == "A3":
events = list(held.case.get("events") or [])
shifted_sets = {
"jitter_7d": [jitter_day_events(events, 7, case_id=held.case_id)],
"shift_1_3_months": [
lib.shift_day_events_by_months(events, seed=f"{SEED}:{held.case_id}:{radius}:{k}")
for k in range(1 if fast else SHIFT_REPEATS)
],
}
for label, variants_of_events in shifted_sets.items():
for index, altered in enumerate(variants_of_events):
cache_key = (held.case_id, label, index)
if cache_key not in altered_stores:
altered_stores[cache_key] = held.rescored_store(altered)[:2]
store, request_alt = altered_stores[cache_key]
rows_alt, _ = held.weighted_rows(table, variant, scale, store=store, request=request_alt)
robustness[label].append(evaluate(held, rows_alt, softmax_prior=True, temperature=temperature))
# answer flips on the LOO rows
asked = list(held.probes)[:ASK_COUNT]
optimal = [optimal_answer(p, held.true_time) for p in asked]
for k in (1, 2):
sets = exhaustive_flip_sets(optimal, k) if k == 1 else sample_flip_sets(
optimal, k, 3 if fast else FLIP_REPEATS, case_id=held.case_id, radius=radius,
)
for picked in sets:
robustness[f"flip_{k}"].append(evaluate(
held, rows, softmax_prior=True, answers=flipped_answers(optimal, picked), temperature=temperature,
))
block["loo"][variant] = {m: lib.summarize([r[m] for r in loo_rows]) for m in PRIORS}
block["loo"][variant]["temperatures"] = sorted(set(temperatures))
block["loo"][variant]["debug_verdict_v4"] = {
m: gate_verdict(block["baseline"][m], block["loo"][variant][m]) for m in PRIORS
}
block["loo"][variant]["verdict"] = "pending_v5"
if robustness:
block["robustness"][variant] = {
label: {m: lib.summarize([r[m] for r in rows_]) for m in PRIORS}
for label, rows_ in robustness.items()
}
if variant == "A3":
block["calibration"] = lib.calibration_bins(calibration_points)
out["radii"][str(radius)] = block
return out
# ---------------------------------------------------------------------------
# R-B
# ---------------------------------------------------------------------------
def absence_probes(data: CaseData, events: Sequence[dict[str, Any]]) -> tuple[list[dict[str, Any]], dict[str, list[int]]]:
gaps = lib.absent_years(events)
probes: list[dict[str, Any]] = []
for domain, years in gaps.items():
for year in years:
probe = lib.split_probe(
data.representatives, birth_date=str(data.request["birth_date"]), domain=domain,
year=year, month=None, source=ABSENCE_SOURCE,
)
if probe is not None:
probes.append(probe)
return probes, gaps
def run_rb(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]:
out: dict[str, Any] = {"points": list(ABSENCE_POINTS), "card_points": CARD_POINTS, "omit_repeats": OMIT_REPEATS, "radii": {}}
for radius, cases in sorted(by_radius.items()):
rows_by_arm: dict[str, list[dict[str, Any]]] = defaultdict(list)
gap_stats = []
omit_rows: dict[str, list[dict[str, Any]]] = defaultdict(list)
for data in cases:
training = [data.events_by_id[i] for i in data.training_ids]
spoken = [{**e, "date": e.get("date_start"), "domain": e["domain"]} for e in training]
probes, gaps = absence_probes(data, spoken)
gap_stats.append({
"case_id": data.case_id, "absent_years": sum(len(v) for v in gaps.values()),
"domains_with_gaps": len(gaps), "absence_probes": len(probes),
})
base = evaluate(data, data.prod_rows, softmax_prior=False)
rows_by_arm["B0"].append(base)
for points in ABSENCE_POINTS:
pre = [(probe, "no", points / CARD_POINTS) for probe in probes]
rows_by_arm[f"absence@{points}"].append(evaluate(data, data.prod_rows, softmax_prior=False, pre_answers=pre))
# omission robustness: user did not mention 1–2 of the training events
for k in (1, 2):
for rep in range(1 if fast else OMIT_REPEATS):
kept = lib.drop_events(spoken, k, seed=f"{SEED}:{data.case_id}:{radius}:{k}:{rep}")
probes_k, _ = absence_probes(data, kept)
for points in ABSENCE_POINTS:
pre = [(probe, "no", points / CARD_POINTS) for probe in probes_k]
omit_rows[f"omit_{k}@{points}"].append(evaluate(data, data.prod_rows, softmax_prior=False, pre_answers=pre))
block: dict[str, Any] = {
"n": len(cases),
"gap_stats": {
"mean_absent_years": round(sum(g["absent_years"] for g in gap_stats) / len(gap_stats), 2),
"mean_absence_probes": round(sum(g["absence_probes"] for g in gap_stats) / len(gap_stats), 2),
"cases_with_probes": sum(1 for g in gap_stats if g["absence_probes"]),
},
"arms": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in rows_by_arm.items()},
"omission": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in omit_rows.items()},
}
for arm in block["arms"]:
if arm == "B0":
continue
block["arms"][arm]["debug_verdict_v4"] = {m: gate_verdict(block["arms"]["B0"][m], block["arms"][arm][m]) for m in PRIORS}
block["arms"][arm]["verdict"] = "pending_v5"
out["radii"][str(radius)] = block
return out
# ---------------------------------------------------------------------------
# R-C
# ---------------------------------------------------------------------------
def run_rc(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]:
out: dict[str, Any] = {"radii": {}}
for radius, cases in sorted(by_radius.items()):
rows_by_arm: dict[str, list[dict[str, Any]]] = defaultdict(list)
askable_counts = []
for data in cases:
originals = {str(e["id"]): e for e in data.case.get("events") or []}
raw_events = list(data.case.get("events") or [])
spoken = lib.degrade_to_year(raw_events)
spoken_rows = data.production_rows_for(spoken)
rows_by_arm["B0_true_precision"].append(evaluate(data, data.prod_rows, softmax_prior=False))
rows_by_arm["C0_spoken"].append(evaluate(data, spoken_rows, softmax_prior=False))
# which spoken (year) events would split candidates if the month were known?
training = set(data.training_ids)
id_map = {str(e["id"]): str(rid) for rid, e in zip(
[str(x["id"]) for x in data.request["events"]], data.case.get("events") or [],
)}
candidates = []
for event in raw_events:
if str(event.get("precision")) != "day":
continue
request_id = id_map.get(str(event["id"]))
if request_id not in training:
continue
year, month = int(str(event["date"])[:4]), int(str(event["date"])[5:7])
probe = lib.split_probe(
data.representatives, birth_date=str(data.request["birth_date"]), domain=str(event["domain"]),
year=year, month=month, source=FOLLOWUP_SOURCE,
)
if probe is None:
continue
candidates.append((float(probe.get("information_gain") or 0), str(event["id"]), probe))
candidates.sort(key=lambda item: (-item[0], item[1]))
askable_counts.append(len(candidates))
for k in (1, 2):
picked = candidates[:k]
if len(picked) < k:
continue
ids = {item[1] for item in picked}
upgraded = lib.upgrade_to_month(spoken, ids, originals)
rows_by_arm[f"C{k}_prior"].append(evaluate(data, data.production_rows_for(upgraded), softmax_prior=False))
pre = [(item[2], "yes", 1.0) for item in picked]
rows_by_arm[f"C{k}_probe"].append(evaluate(data, spoken_rows, softmax_prior=False, pre_answers=pre))
rows_by_arm[f"C{k}_both"].append(evaluate(data, data.production_rows_for(upgraded), softmax_prior=False, pre_answers=pre))
block: dict[str, Any] = {
"n": len(cases),
"askable_per_case": {
"mean": round(sum(askable_counts) / len(askable_counts), 2) if askable_counts else None,
"distribution": {str(k): askable_counts.count(k) for k in sorted(set(askable_counts))},
"cases_with_at_least_one": sum(1 for c in askable_counts if c >= 1),
"cases_with_at_least_two": sum(1 for c in askable_counts if c >= 2),
},
"arms": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in rows_by_arm.items()},
}
for arm in block["arms"]:
if arm.startswith("C0") or arm.startswith("B0"):
continue
block["arms"][arm]["debug_verdict_v4_vs_C0"] = {m: gate_verdict(block["arms"]["C0_spoken"][m], block["arms"][arm][m]) for m in PRIORS}
block["arms"][arm]["verdict"] = "pending_v5"
out["radii"][str(radius)] = block
return out
# ---------------------------------------------------------------------------
# main
# ---------------------------------------------------------------------------
def load_cases(path: Path) -> tuple[dict[str, Any], list[dict[str, Any]]]:
payload = json.loads(path.read_text(encoding="utf-8"))
return payload, list(payload["cases"])
def write_json(path: Path, payload: dict[str, Any]) -> None:
path.write_text(json.dumps(payload, ensure_ascii=False, indent=1, sort_keys=True) + "\n", encoding="utf-8")
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--holdout", default=str(HOLDOUT_V4))
parser.add_argument("--dataset-label", default="v4")
parser.add_argument("--limit", type=int, default=0)
parser.add_argument("--radii", nargs="+", type=int, default=list(RADII))
parser.add_argument("--parts", nargs="+", default=["ra", "rb", "rc"])
parser.add_argument("--cache-dir", default="")
parser.add_argument("--fast", action="store_true", help="fewer repeats (development)")
parser.add_argument("--no-write", action="store_true")
parser.add_argument("--out-suffix", default="2026_09_29")
args = parser.parse_args()
started = clock_module.time()
holdout_path = Path(args.holdout)
meta, cases = load_cases(holdout_path)
if args.limit:
cases = cases[: args.limit]
cache_dir = Path(args.cache_dir) if args.cache_dir else None
by_radius: dict[int, list[CaseData]] = defaultdict(list)
errors: list[dict[str, Any]] = []
reconciliation: list[dict[str, Any]] = []
for case in cases:
for radius in args.radii:
try:
data = CaseData(case, radius, cache_dir)
except Exception as exc: # noqa: BLE001
errors.append({"case_id": case.get("case_id"), "radius": radius, "error": f"{type(exc).__name__}: {exc}", "trace": traceback.format_exc()})
continue
reconciliation.append({"case_id": data.case_id, "radius": radius, **data.reconciliation})
by_radius[radius].append(data)
print(f"loaded {case.get('case_id')} ({clock_module.time() - started:.0f}s)", flush=True)
unreconciled = [row for row in reconciliation if row["changed"]]
header = {
"generated_for": "TASK-rectification-scoring-research-20260929",
"dataset": args.dataset_label,
"holdout": str(holdout_path.relative_to(ROOT)) if holdout_path.is_relative_to(ROOT) else str(holdout_path),
"case_count": len(cases),
"radii": list(args.radii),
"minute_step": MINUTE_STEP,
"ayanamsa": AYANAMSA,
"node_mode": NODE_MODE,
"ask_count": ASK_COUNT,
"probe_today": lib.TODAY.isoformat(),
"probes": "production G0 probes generated once per case/radius from the production matrix; reused for every arm",
"open_set_not_blind": True,
"verdict": "pending_v5",
"reconciliation": {"cases": len(reconciliation), "unreconciled": len(unreconciled), "rows": reconciliation},
"errors": errors,
"seed": SEED,
"fast": bool(args.fast),
}
if unreconciled:
print(json.dumps({"UNRECONCILED": unreconciled}, ensure_ascii=False, indent=1))
return 2
results: dict[str, dict[str, Any]] = {}
if "ra" in args.parts:
results["lr"] = {**header, "part": "R-A", **run_ra(by_radius, fast=args.fast)}
print("R-A done", f"{clock_module.time() - started:.0f}s", flush=True)
if "rb" in args.parts:
results["absence"] = {**header, "part": "R-B", **run_rb(by_radius, fast=args.fast)}
print("R-B done", f"{clock_module.time() - started:.0f}s", flush=True)
if "rc" in args.parts:
results["precision_followup"] = {**header, "part": "R-C", **run_rc(by_radius, fast=args.fast)}
print("R-C done", f"{clock_module.time() - started:.0f}s", flush=True)
for key, payload in results.items():
payload["elapsed_seconds"] = round(clock_module.time() - started, 1)
if not args.no_write:
write_json(OUT_DIR / f"scoring_research_{key}_{args.out_suffix}.json", {k: v for k, v in payload.items() if k != "elapsed_seconds"})
brief: dict[str, Any] = {"reconciliation_unreconciled": len(unreconciled), "errors": len(errors)}
for key, payload in results.items():
brief[key] = {}
for radius, block in payload["radii"].items():
if key == "lr":
brief[key][radius] = {
"baseline": {m: (block["baseline"][m]["top1"], block["baseline"][m]["coverage"], block["baseline"][m]["width_median"], block["baseline"][m]["engine_top1"]) for m in PRIORS},
**{v: {m: (block["loo"][v][m]["top1"], block["loo"][v][m]["coverage"], block["loo"][v][m]["width_median"], block["loo"][v][m]["engine_top1"]) for m in PRIORS} for v in lib.VARIANTS},
}
else:
brief[key][radius] = {arm: {m: (s[m]["top1"], s[m]["coverage"], s[m]["width_median"]) for m in PRIORS} for arm, s in block["arms"].items()}
print(json.dumps(brief, ensure_ascii=False, indent=1))
return 0 if not errors else 1
if __name__ == "__main__":
sys.exit(main())
+820
View File
@@ -0,0 +1,820 @@
"""Scaffold (R-0) for TASK-rectification-scoring-research-20260929.
Offline only. Nothing here changes a production default or a frozen scoring
file. The module records, for every (case, candidate minute, event, sample
date), which scoring rules fire — the eleven production rules from
`active_rectification_event_engine._score_event` plus the auxiliary
transit / Ashtakavarga / Shadbala rules — and a set of *extra* observations
that production computes but does not score (KP cusp sub-lords, Pranapada /
Hora / Ghati / Bhava lagna, D60). Likelihood-ratio weights are estimated per
feature with leave-one-**case**-out folds, and a weighted row provider feeds
the same contribution-matrix / probe / six-question replay machinery that the
09-14 / 09-26 / 09-29 studies used.
Recording provider vs production: the provider returns exactly the production
evidence (same rule ids, same points), so `build_event_contribution_matrix`
yields the same matrix as the production path. `reconcile_rows` asserts that
per case before any experiment (task hard line 4).
"""
from __future__ import annotations
import math
import statistics
import sys
from collections import defaultdict
from dataclasses import dataclass, field
from datetime import date, datetime
from pathlib import Path
from random import Random
from typing import Any, Callable, Sequence
ROOT = Path(__file__).resolve().parents[2]
if str(ROOT) not in sys.path:
sys.path.insert(0, str(ROOT))
from scripts.active_rectification_event_engine import ( # noqa: E402
DOMAIN_CONFIG,
OBSERVATION_ONLY_LAYERS,
_active_narayana,
_active_vimshottari,
_ashtakavarga_auxiliary,
_controlled_transit_rules,
_event_datetime,
_relative_house,
_shadbala_verified_components_auxiliary,
_varga_chart,
_varga_house,
)
import scripts.rectification.event_probes as event_probes # noqa: E402
from scripts.rectification.candidate_contrast import cluster_contexts_by_signature # noqa: E402
from scripts.rectification.case_holdout import holdout_event_ids # noqa: E402
from scripts.rectification.contracts import is_scoreable_event # noqa: E402
from scripts.rectification.refinement_packet import window_scan # noqa: E402
from scripts.rectification.scoring_service import ( # noqa: E402
build_event_contribution_matrix,
score_from_matrix,
)
from scripts.research.cluster_width_lib import ( # noqa: E402
SEPARATION_LEAD,
delivery_from_public,
merge_adjacent_traced,
metrics_bundle,
public_from_clusters,
raw_signature_clusters,
shannon_entropy,
still_valid_public,
)
from scripts.research.precision_gate_lib import ( # noqa: E402
VargaPolicy,
score_event_with_policy,
)
from scripts.research.probe_supply_after_six import ( # noqa: E402
ASK_COUNT,
SCORE_DELTA,
apply_answer,
inherit_direction,
optimal_answer,
outcome_groups,
)
import varga # noqa: E402
TODAY = date(2026, 9, 16) # same probe "today" as the 09-16 / 09-29 replays
LAPLACE_ALPHA = 1.0
EPS = 1e-9
EXTRA_PREFIXES = ("kp_", "pranapada_", "hora_lagna_", "ghati_lagna_", "bhava_lagna_")
D60_SUFFIX = "_domain_varga_d60"
IGNORED_RULE_PREFIXES = ("event_kind:", "event_kind_profile:", "no_domain_activation")
CONSTANT_RULES = frozenset({"occupation_auxiliary_not_primary", "appearance_auxiliary_not_primary"})
VARIANTS = ("A1", "A2", "A3")
VARIANT_SPEC = {
# (uses extra features, allows negative / absence weights)
"A1": (False, False),
"A2": (True, False),
"A3": (True, True),
}
# ---------------------------------------------------------------------------
# small helpers
# ---------------------------------------------------------------------------
def hhmm(value: object) -> str | None:
text = str(value or "")[:5]
return text if len(text) == 5 and text[2] == ":" else None
def clock(value: str) -> int:
return int(value[:2]) * 60 + int(value[3:5])
def is_extra_feature(name: str) -> bool:
return name.startswith(EXTRA_PREFIXES) or name.endswith(D60_SUFFIX)
def is_scoring_rule(name: str) -> bool:
if name.startswith(IGNORED_RULE_PREFIXES) or name in CONSTANT_RULES:
return False
return not is_extra_feature(name)
def median(values: Sequence[float]) -> float | None:
items = [float(item) for item in values if item is not None]
return statistics.median(items) if items else None
# ---------------------------------------------------------------------------
# R-0 · feature recording provider (production-identical evidence + extras)
# ---------------------------------------------------------------------------
@dataclass
class FeatureStore:
"""(event_id, sample_date) -> HH:MM -> {rules, extras}."""
rows: dict[tuple[str, str], dict[str, dict[str, Any]]] = field(default_factory=dict)
def put(self, event_id: str, sample: str, time: str, rules: Sequence[str], extras: Sequence[str]) -> None:
self.rows.setdefault((event_id, sample), {})[time] = {
"rules": sorted(set(rules)),
"extras": sorted(set(extras)),
}
def samples_for(self, event_id: str) -> list[str]:
return sorted(sample for (item, sample) in self.rows if item == event_id)
def d60_charts(contexts: Sequence[dict[str, Any]]) -> dict[str, dict[str, Any] | None]:
out: dict[str, dict[str, Any] | None] = {}
for context in contexts:
stamp = context["candidate_at"].strftime("%H:%M")
try:
computed = varga.calc_all_vargas(
context["planet_longitudes"], float(context["ascendant_longitude"]), divisions=[60],
)
out[stamp] = _varga_chart(computed, "D60")
except (KeyError, TypeError, ValueError):
out[stamp] = None
return out
def extra_features(
context: dict[str, Any],
target_houses: tuple[int, ...],
vimshottari: tuple[str, str, str],
d60: dict[str, Any] | None,
) -> list[str]:
feature = context.get("feature") if isinstance(context.get("feature"), dict) else {}
ascendant_index = int(context["ascendant_index"])
kp_houses = ((feature.get("kp_cusps") or {}).get("houses") or {})
target_sub = {str((kp_houses.get(str(h)) or {}).get("sub_lord") or "") for h in target_houses}
target_star = {str((kp_houses.get(str(h)) or {}).get("nakshatra_lord") or "") for h in target_houses}
asc_sub = str((kp_houses.get("1") or {}).get("sub_lord") or "")
out: list[str] = []
for lord, label in zip(vimshottari, ("md", "ad", "pd")):
if lord in target_sub:
out.append(f"kp_target_cusp_sublord_is_vim_{label}")
if lord in target_star:
out.append(f"kp_target_cusp_starlord_is_vim_{label}")
if asc_sub and lord == asc_sub:
out.append(f"kp_asc_sublord_is_vim_{label}")
if isinstance(d60, dict) and _varga_house(d60, lord) in target_houses:
out.append(f"vim_{label}{D60_SUFFIX}")
for key, name in (
("pranapada_sign_index", "pranapada"),
("hora_sign_index", "hora_lagna"),
("ghati_sign_index", "ghati_lagna"),
("bhava_sign_index", "bhava_lagna"),
):
index = feature.get(key)
if isinstance(index, int) and _relative_house(index, ascendant_index) in target_houses:
out.append(f"{name}_in_target_house")
return out
def recorded_row(
request: dict[str, Any],
context: dict[str, Any],
*,
store: FeatureStore,
transit_cache: dict[tuple[Any, ...], dict[str, Any]],
d60_by_time: dict[str, dict[str, Any] | None],
) -> dict[str, Any]:
"""Production `_candidate_row` for one legacy (single event, one date) request, plus extras."""
candidate_at = context["candidate_at"]
chart = context["chart"]
planet_longitudes = context["planet_longitudes"]
ascendant_index = context["ascendant_index"]
arudha_padas = context["arudha_padas"]
varga_charts = context["varga_charts"]
moon_longitude = planet_longitudes["Moon"]
stamp = candidate_at.strftime("%H:%M")
policy = VargaPolicy(name="V0", window_minutes=0.0)
evidence: list[dict[str, Any]] = []
missing: list[str] = []
for event in request["events"]:
event_at = _event_datetime(event)
prefixes, target_houses = DOMAIN_CONFIG[event["domain"]]
selected = {prefix: varga_charts.get(prefix) for prefix in prefixes}
if any(item is None for item in selected.values()):
missing.extend(prefixes)
continue
try:
vimshottari = _active_vimshottari(
candidate_at.date().isoformat(), moon_longitude, event_at, context.get("vimshottari_timeline"),
)
except (KeyError, TypeError, ValueError):
missing.append("Vimshottari_MD_AD_PD")
continue
try:
narayana = _active_narayana(
ascendant_index, planet_longitudes, candidate_at, event_at, context.get("narayana_periods"),
)
except (KeyError, TypeError, ValueError):
missing.append("Narayana_MD_AD")
continue
if narayana[0] is None or narayana[1] is None:
missing.append("Narayana_MD_AD")
continue
row = score_event_with_policy(
candidate_time=stamp, event=event, natal_chart=chart,
varga_by_prefix={k: v for k, v in selected.items() if isinstance(v, dict)},
vimshottari=vimshottari, narayana=narayana, arudha_padas=arudha_padas, policy=policy,
)
weight = 1.0 # legacy requests carry precision "day"; the matrix applies the real weight
transit_rules = _controlled_transit_rules(
request, event, ascendant_index, target_houses, transit_cache,
)
if transit_rules:
row["rule_ids"].extend(transit_rules)
row["points"] = round(row["points"] + 0.25 * len(transit_rules) * weight, 4)
av_rules, av_points = _ashtakavarga_auxiliary(
chart, ascendant_index, target_houses, context.get("ashtakavarga_result"),
)
if av_rules:
row["rule_ids"].extend(av_rules)
row["points"] = round(row["points"] + av_points * weight, 4)
sb_rules, sb_points = _shadbala_verified_components_auxiliary(
chart, candidate_at.hour + candidate_at.minute / 60, vimshottari, context.get("shadbala_result"),
)
if sb_rules:
row["rule_ids"].extend(sb_rules)
row["points"] = round(row["points"] + sb_points * weight, 4)
evidence.append(row)
store.put(
str(event["id"]), str(event["date"]), stamp, row["rule_ids"],
extra_features(context, target_houses, vimshottari, d60_by_time.get(stamp)),
)
return {
"time": stamp,
"score": round(sum(item["points"] for item in evidence), 4),
"evidence": evidence,
"missing_layers": sorted(set(
missing + [layer for layer in context["feature"]["blocked_layers"] if layer not in OBSERVATION_ONLY_LAYERS]
)),
}
def make_recording_provider(
contexts: Sequence[dict[str, Any]],
store: FeatureStore,
d60_by_time: dict[str, dict[str, Any] | None],
) -> Callable[[dict[str, Any]], list[dict[str, Any]]]:
transit_cache: dict[tuple[Any, ...], dict[str, Any]] = {}
def provider(request: dict[str, Any]) -> list[dict[str, Any]]:
return [
recorded_row(request, context, store=store, transit_cache=transit_cache, d60_by_time=d60_by_time)
for context in contexts
]
return provider
def score_map(rows: Sequence[dict[str, Any]]) -> dict[str, float]:
return {str(row["time"])[:5]: float(row.get("score") or 0) for row in rows}
def reconcile_rows(production: Sequence[dict[str, Any]], research: Sequence[dict[str, Any]]) -> dict[str, Any]:
"""Per-minute score and rule-id comparison. `changed == 0` is the R-0 gate."""
prod = {str(r["time"])[:5]: r for r in production}
rese = {str(r["time"])[:5]: r for r in research}
changed: list[dict[str, Any]] = []
for time, row in prod.items():
other = rese.get(time)
if other is None:
changed.append({"time": time, "reason": "missing"})
continue
if abs(float(row.get("score") or 0) - float(other.get("score") or 0)) > EPS:
changed.append({"time": time, "reason": "score", "production": row.get("score"), "research": other.get("score")})
continue
prod_rules = {
(item["event_id"], tuple(sorted(item["rule_ids"]))) for item in row.get("evidence") or []
}
rese_rules = {
(item["event_id"], tuple(sorted(item["rule_ids"]))) for item in other.get("evidence") or []
}
if prod_rules != rese_rules:
changed.append({"time": time, "reason": "rule_ids"})
return {"candidates": len(prod), "changed": len(changed), "details": changed[:5]}
# ---------------------------------------------------------------------------
# feature matrices and likelihood-ratio weights
# ---------------------------------------------------------------------------
def event_feature_matrix(
store: FeatureStore,
event_id: str,
times: Sequence[str],
*,
include_extras: bool,
) -> dict[str, dict[str, float]]:
"""HH:MM -> feature -> fraction of sample dates on which it fired."""
samples = store.samples_for(event_id)
out: dict[str, dict[str, float]] = {}
if not samples:
return {time: {} for time in times}
for time in times:
counts: dict[str, int] = defaultdict(int)
for sample in samples:
cell = store.rows.get((event_id, sample), {}).get(time)
if not cell:
continue
names = [name for name in cell["rules"] if is_scoring_rule(name)]
if include_extras:
names += list(cell["extras"])
for name in names:
counts[name] += 1
out[time] = {name: round(count / len(samples), 6) for name, count in counts.items()}
return out
def training_event_ids(request: dict[str, Any]) -> list[str]:
holdout = holdout_event_ids(request["events"])
return [str(e["id"]) for e in request["events"] if is_scoreable_event(e) and str(e["id"]) not in holdout]
@dataclass
class LRTable:
alpha: float
positives: float
negatives: float
present: dict[str, float] # log LR of the feature firing
absent: dict[str, float] # log LR of the feature not firing (naive Bayes complement)
support: dict[str, dict[str, float]]
def weight(self, name: str, *, allow_negative: bool) -> tuple[float, float]:
"""(present weight, absent weight) under the variant policy."""
present = self.present.get(name, 0.0)
absent = self.absent.get(name, 0.0)
if allow_negative:
return present, absent
return max(present, 0.0), 0.0
def estimate_lr(examples: Sequence[tuple[dict[str, float], bool]], *, alpha: float = LAPLACE_ALPHA) -> LRTable:
"""Per-feature log likelihood ratios from (feature fractions, is-truth-cluster) examples."""
positives = float(sum(1 for _values, truth in examples if truth))
negatives = float(sum(1 for _values, truth in examples if not truth))
names: set[str] = set()
pos_sum: dict[str, float] = defaultdict(float)
neg_sum: dict[str, float] = defaultdict(float)
for values, truth in examples:
for name, value in values.items():
names.add(name)
if truth:
pos_sum[name] += float(value)
else:
neg_sum[name] += float(value)
present: dict[str, float] = {}
absent: dict[str, float] = {}
support: dict[str, dict[str, float]] = {}
for name in sorted(names):
p_pos = (pos_sum[name] + alpha) / (positives + 2 * alpha)
p_neg = (neg_sum[name] + alpha) / (negatives + 2 * alpha)
present[name] = round(math.log(p_pos / p_neg), 6)
absent[name] = round(math.log((1 - p_pos) / (1 - p_neg)), 6)
support[name] = {
"fires_truth": round(pos_sum[name], 4),
"fires_other": round(neg_sum[name], 4),
"p_truth": round(p_pos, 6),
"p_other": round(p_neg, 6),
}
return LRTable(alpha=alpha, positives=positives, negatives=negatives, present=present, absent=absent, support=support)
def bootstrap_lr(
per_case_examples: dict[str, list[tuple[dict[str, float], bool]]],
*,
resamples: int = 200,
seed: int = 20260929,
alpha: float = LAPLACE_ALPHA,
) -> dict[str, dict[str, float]]:
"""Case-level bootstrap 5–95 % interval of the present-log-LR per feature."""
ids = sorted(per_case_examples)
rng = Random(f"{seed}:lr-bootstrap")
draws: dict[str, list[float]] = defaultdict(list)
for _ in range(resamples):
picked = [rng.choice(ids) for _ in ids]
table = estimate_lr([ex for cid in picked for ex in per_case_examples[cid]], alpha=alpha)
for name, value in table.present.items():
draws[name].append(value)
out: dict[str, dict[str, float]] = {}
for name, values in draws.items():
ordered = sorted(values)
lo = ordered[int(0.05 * (len(ordered) - 1))]
hi = ordered[int(0.95 * (len(ordered) - 1))]
out[name] = {"p05": round(lo, 4), "p95": round(hi, 4), "draws": len(ordered)}
return out
def weighted_points(values: dict[str, float], table: LRTable, *, allow_negative: bool, universe: Sequence[str]) -> float:
total = 0.0
for name in universe:
present, absent = table.weight(name, allow_negative=allow_negative)
x = float(values.get(name, 0.0))
total += x * present + (1.0 - x) * absent
return total
def make_weighted_provider(
store: FeatureStore,
table: LRTable,
*,
variant: str,
scale: float = 1.0,
baseline_rows_by_sample: dict[tuple[str, str], dict[str, dict[str, Any]]] | None = None,
) -> Callable[[dict[str, Any]], list[dict[str, Any]]]:
"""Row provider that scores each (event, sample date) as Σ feature weights.
Production rule ids are kept on the evidence so the matrix's kind factor
and technique layers behave as in production; only the points change.
"""
include_extras, allow_negative = VARIANT_SPEC[variant]
universe = [
name for name in sorted(table.present)
if (include_extras or not is_extra_feature(name))
]
def provider(request: dict[str, Any]) -> list[dict[str, Any]]:
rows: list[dict[str, Any]] = []
for event in request["events"]:
key = (str(event["id"]), str(event["date"]))
cells = store.rows.get(key, {})
for time in sorted(cells, key=clock):
cell = cells[time]
names = [name for name in cell["rules"] if is_scoring_rule(name)]
if include_extras:
names += list(cell["extras"])
values = {name: 1.0 for name in names}
points = round(scale * weighted_points(values, table, allow_negative=allow_negative, universe=universe), 4)
rows.append({
"time": time,
"score": points,
"evidence": [{
"event_id": event["id"], "domain": event["domain"], "candidate_time": time,
"rule_ids": list(cell["rules"]), "points": points,
}],
"missing_layers": [],
})
return rows
return provider
def loo_folds(case_ids: Sequence[str]) -> list[tuple[str, list[str]]]:
ids = list(case_ids)
return [(held, [other for other in ids if other != held]) for held in ids]
# ---------------------------------------------------------------------------
# truth labels, priors, replay
# ---------------------------------------------------------------------------
def truth_cluster_times(contexts: Sequence[dict[str, Any]], true_time: str) -> list[str]:
for cluster in raw_signature_clusters(contexts):
times = [str(item)[:5] for item in cluster.get("times") or []]
if true_time[:5] in times:
return sorted(times, key=clock)
return [true_time[:5]]
def public_for(rows: Sequence[dict[str, Any]], contexts: Sequence[dict[str, Any]]) -> list[dict[str, Any]]:
raw = raw_signature_clusters(contexts)
by_time = {stamp: row for row in rows if (stamp := hhmm(row.get("time")))}
merged, _trace = merge_adjacent_traced(raw, by_time)
public = public_from_clusters(merged, rows)
for row in public:
row["score"] = float(row.get("score") or 0)
return public
def percent_proportional(scores: dict[str, float]) -> dict[str, float]:
total = sum(max(value, 0.0) for value in scores.values())
if total <= 0:
return {key: 0.0 for key in scores}
return {key: float(round(max(value, 0.0) / total * 100)) for key, value in scores.items()}
def softmax_probabilities(scores: dict[str, float], temperature: float = 1.0) -> dict[str, float]:
if not scores:
return {}
temp = float(temperature) if temperature and temperature > 0 else 1.0
peak = max(scores.values())
weights = {key: math.exp((value - peak) / temp) for key, value in scores.items()}
total = sum(weights.values())
return {key: weight / total for key, weight in weights.items()}
def percent_softmax(scores: dict[str, float], temperature: float = 1.0) -> dict[str, float]:
return {key: float(round(p * 100)) for key, p in softmax_probabilities(scores, temperature).items()}
TEMPERATURE_GRID = (0.25, 0.5, 1.0, 2.0, 4.0, 8.0, 16.0, 32.0, 64.0, 128.0, 256.0)
def truth_mass(public: Sequence[dict[str, Any]], probabilities: dict[str, float], truth_times: Sequence[str]) -> float:
truth = {str(t)[:5] for t in truth_times}
mass = 0.0
for row in public:
stamp = str(row["time"])[:5]
members = {str(t)[:5] for t in (row.get("cluster_times") or [stamp])}
if members & truth:
mass += probabilities.get(stamp, 0.0)
return mass
def fit_temperature(
training: Sequence[tuple[Sequence[dict[str, Any]], Sequence[str]]],
grid: Sequence[float] = TEMPERATURE_GRID,
) -> float:
"""Platt-style single temperature: maximise Σ log P(truth cluster) on training cases only."""
best_t, best_ll = 1.0, None
for temp in grid:
ll = 0.0
for public, truth_times in training:
base = {str(r["time"])[:5]: float(r.get("score") or 0) for r in public}
mass = truth_mass(public, softmax_probabilities(base, temp), truth_times)
ll += math.log(max(mass, 1e-9))
if best_ll is None or ll > best_ll + 1e-12:
best_t, best_ll = float(temp), ll
return best_t
def prior_for(public: Sequence[dict[str, Any]], *, mode: str, softmax: bool, temperature: float = 1.0) -> dict[str, float]:
base = {str(row["time"])[:5]: float(row.get("score") or 0) for row in public}
if mode == "raw":
return base
return percent_softmax(base, temperature) if softmax else percent_proportional(base)
def apply_scaled_answer(
scores: dict[str, float],
conflicts: dict[str, int],
eliminated: set[str],
probe: dict[str, Any],
answer: str,
candidate_times: Sequence[str],
factor: float,
) -> tuple[dict[str, float], dict[str, int], set[str]]:
"""`apply_answer` at card strength (factor == 1); weaker evidence scales the
delta and never counts a strong conflict (so it can never eliminate)."""
if factor >= 1.0 - EPS:
return apply_answer(scores, conflicts, eliminated, probe, answer, candidate_times)
yes, no = outcome_groups(probe)
next_scores = dict(scores)
for time in candidate_times:
if time in eliminated:
continue
if time in yes:
raw = "support"
elif time in no:
raw = "conflict"
else:
raw = inherit_direction(time, yes, no)
if answer == "no":
raw = {"support": "conflict", "conflict": "support", "neutral": "neutral"}[raw]
next_scores[time] = next_scores.get(time, 0.0) + SCORE_DELTA[raw] * factor
return next_scores, dict(conflicts), set(eliminated)
def replay(
*,
public: Sequence[dict[str, Any]],
prior: dict[str, float],
probes: Sequence[dict[str, Any]],
true_time: str,
window_times: Sequence[str],
pre_answers: Sequence[tuple[dict[str, Any], str, float]] = (),
answers: Sequence[str | None] | None = None,
) -> dict[str, Any]:
"""Prior → optional pre-answers (probe, answer, factor) → six probes answered from truth."""
reps = [str(row["time"])[:5] for row in public]
scores = {time: float(prior.get(time, 0.0)) for time in reps}
conflicts = {time: 0 for time in reps}
eliminated: set[str] = set()
for probe, answer, factor in pre_answers:
scores, conflicts, eliminated = apply_scaled_answer(scores, conflicts, eliminated, probe, answer, reps, factor)
asked = list(probes)[:ASK_COUNT]
given = list(answers) if answers is not None else [optimal_answer(p, true_time) for p in asked]
answered = 0
for probe, answer in zip(asked, given):
if answer is None:
continue
answered += 1
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, answer, reps)
posterior = [{**row, "score": scores.get(str(row["time"])[:5], row.get("score") or 0)} for row in public]
valid = still_valid_public(posterior, scores, eliminated, lead=SEPARATION_LEAD)
delivery = delivery_from_public(valid)
alive = [row for row in posterior if str(row["time"])[:5] not in eliminated]
metrics = metrics_bundle(
public=alive, true_time=true_time, window_times=window_times,
delivery_times=delivery["times"], delivery_width=delivery["width"], independent=True,
entropy_scores=[scores.get(str(row["time"])[:5], 0.0) for row in alive],
)
truth_eliminated = any(
true_time in (row.get("cluster_times") or [str(row.get("time"))[:5]]) and str(row["time"])[:5] in eliminated
for row in public
)
return {
"top1": bool(metrics["top1_hit"]),
"coverage": bool(metrics["coverage"]),
"width": delivery["width"],
"tie": bool(metrics["tie"]),
"squeezed": bool(metrics["truth_squeezed"]),
"eliminated": len(eliminated),
"truth_eliminated": bool(truth_eliminated),
"questions": answered,
"entropy": round(shannon_entropy(max(s, 0.0) for t, s in scores.items() if t not in eliminated), 4),
}
def engine_top1(public: Sequence[dict[str, Any]], true_time: str) -> bool:
from scripts.research.cluster_width_lib import top1_from_public
return bool(top1_from_public(public, true_time))
def within_case_range(rows: Sequence[dict[str, Any]]) -> float:
scores = [float(row.get("score") or 0) for row in rows]
return (max(scores) - min(scores)) if scores else 0.0
def summarize(rows: Sequence[dict[str, Any]]) -> dict[str, Any]:
"""Same keys as precision_gate_sweep.summarize so `gate_verdict` can read it."""
if not rows:
return {"n": 0, "top1": None, "coverage": None, "width_median": None, "tie": None, "squeezed": 0,
"engine_top1": None, "refresh_mean": 0.0, "entropy": None, "truth_eliminated": 0}
n = len(rows)
return {
"n": n,
"top1": round(sum(1 for r in rows if r.get("top1")) / n, 4),
"coverage": round(sum(1 for r in rows if r.get("coverage")) / n, 4),
"width_median": median([r.get("width") for r in rows]),
"tie": round(sum(1 for r in rows if r.get("tie")) / n, 4),
"squeezed": sum(1 for r in rows if r.get("squeezed")),
"engine_top1": round(sum(1 for r in rows if r.get("engine_top1")) / n, 4),
"refresh_mean": 0.0,
"entropy": round(sum(float(r.get("entropy") or 0) for r in rows) / n, 4),
"truth_eliminated": sum(1 for r in rows if r.get("truth_eliminated")),
}
# ---------------------------------------------------------------------------
# probes on representatives (typed-event convention), absence, precision follow-up
# ---------------------------------------------------------------------------
def production_probes(request: dict[str, Any], built: dict[str, Any], times: Sequence[str], true_time: str) -> list[dict[str, Any]]:
payload = {**request, "refresh_probes": False, "asked_probe_keys": []}
return event_probes.discriminating_event_probes(
payload, built, scan=window_scan(built), candidate_times=list(times),
representative_time=true_time, today=TODAY,
)
@dataclass
class Representatives:
reps: list[dict[str, Any]]
clusters: list[dict[str, Any]]
set_version: str
def representatives_for(contexts: Sequence[dict[str, Any]]) -> Representatives:
full = [item for item in contexts if event_probes._context_time(item)]
full.sort(key=lambda item: event_probes._clock(str(event_probes._context_time(item))))
clusters = cluster_contexts_by_signature(full)
reps = [c["representative"] for c in clusters if event_probes._scoreable(c["representative"])]
if len(reps) < 2:
reps = [item for item in full if event_probes._scoreable(item)]
version = event_probes.candidate_set_version([c["times"] for c in clusters])
return Representatives(reps=reps, clusters=clusters, set_version=version)
def split_probe(
representatives: Representatives,
*,
birth_date: str,
domain: str,
year: int,
month: int | None,
source: str,
) -> dict[str, Any] | None:
canonical = event_probes.canonical_domain(domain)
if canonical not in event_probes.DOMAIN_CATALOG:
return None
return event_probes._evaluate_contexts(
representatives.reps, birth_date=birth_date, domain=canonical, year=year, month=month,
source=source, clusters=representatives.clusters, set_version=representatives.set_version,
)
def event_year(event: dict[str, Any]) -> int | None:
raw = str(event.get("date_start") or event.get("date") or "")
return int(raw[:4]) if raw[:4].isdigit() else None
def event_month(event: dict[str, Any]) -> int | None:
raw = str(event.get("date_start") or event.get("date") or "")
return int(raw[5:7]) if len(raw) >= 7 and raw[4] == "-" else None
def absent_years(events: Sequence[dict[str, Any]]) -> dict[str, list[int]]:
"""Per domain: years strictly inside the span of that domain's dated events with no event."""
years: dict[str, set[int]] = defaultdict(set)
for event in events:
year = event_year(event)
if year is None:
continue
years[str(event.get("domain") or "")].add(year)
out: dict[str, list[int]] = {}
for domain, known in years.items():
if len(known) < 2:
continue
lo, hi = min(known), max(known)
gaps = [year for year in range(lo + 1, hi) if year not in known]
if gaps:
out[domain] = gaps
return dict(sorted(out.items()))
def drop_events(events: Sequence[dict[str, Any]], count: int, *, seed: str) -> list[dict[str, Any]]:
rng = Random(seed)
if count <= 0 or count >= len(events):
return list(events)
dropped = set(rng.sample(range(len(events)), count))
return [event for index, event in enumerate(events) if index not in dropped]
def degrade_to_year(events: Sequence[dict[str, Any]]) -> list[dict[str, Any]]:
"""The 'spoken' form of a case: every day/month event becomes year precision."""
out = []
for event in events:
item = dict(event)
if str(item.get("precision")) in {"day", "month"}:
item["precision"] = "year"
item["date"] = str(item.get("date") or "")[:4]
out.append(item)
return out
def upgrade_to_month(events: Sequence[dict[str, Any]], event_ids: set[str], originals: dict[str, dict[str, Any]]) -> list[dict[str, Any]]:
out = []
for event in events:
item = dict(event)
original = originals.get(str(item.get("id")))
if str(item.get("id")) in event_ids and original is not None and len(str(original.get("date") or "")) >= 7:
item["precision"] = "month"
item["date"] = str(original["date"])[:7]
out.append(item)
return out
def shift_day_events_by_months(events: Sequence[dict[str, Any]], *, seed: str, choices: Sequence[int] = (-3, -2, -1, 1, 2, 3)) -> list[dict[str, Any]]:
rng = Random(seed)
out = []
for event in events:
item = dict(event)
raw = str(item.get("date") or "")
if str(item.get("precision")) == "day" and len(raw) >= 10:
year, month, day = int(raw[:4]), int(raw[5:7]), int(raw[8:10])
index = year * 12 + (month - 1) + rng.choice(list(choices))
new_year, new_month = divmod(index, 12)
item["date"] = f"{new_year:04d}-{new_month + 1:02d}-{min(day, 28):02d}"
item["shift_months"] = index - (year * 12 + month - 1)
out.append(item)
return out
def calibration_bins(points: Sequence[tuple[float, bool]], edges: Sequence[float] = (0.0, 0.05, 0.1, 0.2, 0.3, 0.5, 0.7, 1.0001)) -> list[dict[str, Any]]:
out = []
for lo, hi in zip(edges, edges[1:]):
inside = [(p, t) for p, t in points if lo <= p < hi]
out.append({
"bin": f"[{lo:.2f},{min(hi, 1.0):.2f})",
"n": len(inside),
"predicted_mean": round(sum(p for p, _ in inside) / len(inside), 4) if inside else None,
"observed_rate": round(sum(1 for _, t in inside if t) / len(inside), 4) if inside else None,
})
return out
__all__ = [name for name in dir() if not name.startswith("__")]
+159
View File
@@ -0,0 +1,159 @@
"""Regression tests for the 2026-09-29 scoring-method research scaffold.
Pure helpers use fictional inputs. The one integration test reconciles the
research recording scorer against the production scorer on a public v4 case
(task hard line 4: zero difference before any experiment).
"""
from __future__ import annotations
import json
from pathlib import Path
import pytest
import scripts.research.scoring_research_lib as lib
ROOT = Path(__file__).resolve().parents[1]
HOLDOUT_V4 = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v4.json"
def test_estimate_lr_rewards_discriminating_feature_and_leaves_noise_near_zero() -> None:
examples = []
for _ in range(20):
examples.append(({"good": 1.0, "noise": 1.0}, True))
examples.append(({"good": 1.0, "noise": 0.0}, True))
for _ in range(80):
examples.append(({"noise": 1.0}, False))
examples.append(({"noise": 0.0}, False))
table = lib.estimate_lr(examples, alpha=1.0)
assert table.positives == 40 and table.negatives == 160
assert table.present["good"] > 1.5
assert table.absent["good"] < -1.5
assert abs(table.present["noise"]) < 0.05
# reward-only policy clips the negative absent term; naive-Bayes keeps it
assert table.weight("good", allow_negative=False) == (table.present["good"], 0.0)
assert table.weight("good", allow_negative=True) == (table.present["good"], table.absent["good"])
def test_loo_folds_never_contain_the_held_out_case() -> None:
folds = lib.loo_folds(["a", "b", "c"])
assert [held for held, _ in folds] == ["a", "b", "c"]
for held, training in folds:
assert held not in training and len(training) == 2
def test_event_feature_matrix_is_the_fraction_of_sample_dates() -> None:
store = lib.FeatureStore()
store.put("e1", "1990-01-15", "10:00", ["vim_md_domain_house", "event_kind:career"], ["pranapada_in_target_house"])
store.put("e1", "1990-02-15", "10:00", ["vim_ad_domain_lord"], [])
store.put("e1", "1990-01-15", "10:02", ["no_domain_activation"], [])
store.put("e1", "1990-02-15", "10:02", ["no_domain_activation"], [])
matrix = lib.event_feature_matrix(store, "e1", ["10:00", "10:02"], include_extras=True)
assert matrix["10:00"] == {"vim_md_domain_house": 0.5, "vim_ad_domain_lord": 0.5, "pranapada_in_target_house": 0.5}
assert matrix["10:02"] == {}
without = lib.event_feature_matrix(store, "e1", ["10:00"], include_extras=False)
assert "pranapada_in_target_house" not in without["10:00"]
def test_scoring_rule_filter_drops_event_kind_and_constant_rules() -> None:
assert lib.is_scoring_rule("vim_md_domain_varga")
assert lib.is_scoring_rule("controlled_transit_jupiter_domain_house")
assert not lib.is_scoring_rule("event_kind:career")
assert not lib.is_scoring_rule("event_kind_profile:career:change")
assert not lib.is_scoring_rule("no_domain_activation")
assert not lib.is_scoring_rule("occupation_auxiliary_not_primary")
assert not lib.is_scoring_rule("kp_asc_sublord_is_vim_md") # extras are a separate namespace
assert lib.is_extra_feature("vim_md_domain_varga_d60")
def test_weighted_provider_scores_each_sample_from_the_store() -> None:
store = lib.FeatureStore()
store.put("e1", "1990-01-15", "10:00", ["vim_md_domain_house"], ["kp_asc_sublord_is_vim_md"])
store.put("e1", "1990-01-15", "10:02", [], [])
table = lib.LRTable(
alpha=1.0, positives=1, negatives=1,
present={"vim_md_domain_house": 1.0, "kp_asc_sublord_is_vim_md": 0.5},
absent={"vim_md_domain_house": -0.25, "kp_asc_sublord_is_vim_md": -0.1},
support={},
)
request = {"events": [{"id": "e1", "date": "1990-01-15", "domain": "career"}]}
a1 = lib.make_weighted_provider(store, table, variant="A1")(request)
a2 = lib.make_weighted_provider(store, table, variant="A2")(request)
a3 = lib.make_weighted_provider(store, table, variant="A3", scale=2.0)(request)
assert [(r["time"], r["score"]) for r in a1] == [("10:00", 1.0), ("10:02", 0.0)]
assert [(r["time"], r["score"]) for r in a2] == [("10:00", 1.5), ("10:02", 0.0)]
assert [(r["time"], r["score"]) for r in a3] == [("10:00", 3.0), ("10:02", -0.7)]
assert a1[0]["evidence"][0]["rule_ids"] == ["vim_md_domain_house"]
def _probe() -> dict:
return {"expected_outcomes": [
{"answer_class": "yes", "supports": ["10:00"], "conflicts": ["10:04"]},
{"answer_class": "no", "supports": ["10:04"], "conflicts": ["10:00"]},
]}
def test_scaled_answer_halves_the_delta_and_never_counts_a_conflict() -> None:
scores = {"10:00": 10.0, "10:02": 10.0, "10:04": 10.0}
conflicts = {time: 0 for time in scores}
weak, weak_conflicts, weak_elim = lib.apply_scaled_answer(scores, conflicts, set(), _probe(), "no", list(scores), 0.5)
assert weak == {"10:00": 9.0, "10:02": 10.0, "10:04": 11.0}
assert weak_conflicts == conflicts and weak_elim == set()
full, full_conflicts, _ = lib.apply_scaled_answer(scores, conflicts, set(), _probe(), "no", list(scores), 1.0)
assert full == {"10:00": 8.0, "10:02": 10.0, "10:04": 12.0}
assert full_conflicts["10:00"] == 1
def test_absent_years_are_gaps_strictly_inside_each_domain_span() -> None:
events = [
{"domain": "career", "date": "2017-05-01", "precision": "day"},
{"domain": "career", "date": "2019", "precision": "year"},
{"domain": "career", "date": "2025-02", "precision": "month"},
{"domain": "relocation", "date": "2000", "precision": "year"},
]
assert lib.absent_years(events) == {"career": [2018, 2020, 2021, 2022, 2023, 2024]}
def test_degrade_and_upgrade_precision_helpers() -> None:
events = [{"id": "a", "domain": "career", "date": "2017-05-20", "precision": "day"},
{"id": "b", "domain": "career", "date": "2019", "precision": "year"}]
spoken = lib.degrade_to_year(events)
assert spoken[0]["precision"] == "year" and spoken[0]["date"] == "2017"
assert spoken[1] == events[1]
upgraded = lib.upgrade_to_month(spoken, {"a"}, {e["id"]: e for e in events})
assert upgraded[0]["precision"] == "month" and upgraded[0]["date"] == "2017-05"
shifted = lib.shift_day_events_by_months(events, seed="fictional")
assert shifted[0]["shift_months"] in {-3, -2, -1, 1, 2, 3}
assert shifted[1] == events[1]
assert lib.drop_events(events, 1, seed="x") != events and len(lib.drop_events(events, 1, seed="x")) == 1
def test_softmax_and_proportional_percent_sum_to_about_one_hundred() -> None:
scores = {"10:00": 12.0, "10:02": 11.0, "10:04": 9.0}
assert abs(sum(lib.percent_proportional(scores).values()) - 100) <= 2
soft = lib.percent_softmax(scores)
assert abs(sum(soft.values()) - 100) <= 2 and soft["10:00"] > soft["10:02"] > soft["10:04"]
assert lib.calibration_bins([(0.9, True), (0.02, False)])[-1]["observed_rate"] == 1.0
@pytest.mark.skipif(not HOLDOUT_V4.exists(), reason="v4 open set not present")
def test_recording_scorer_reconciles_with_production_on_a_public_case() -> None:
from scripts.active_rectification_event_engine import compute_candidate_static_contexts
from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix
from scripts.research.minute_resolution_sweep import scoring_request_for
case = json.loads(HOLDOUT_V4.read_text(encoding="utf-8"))["cases"][0]
request = scoring_request_for({**case, "candidate_radius_minutes": 10}, 10)
contexts = compute_candidate_static_contexts(request)
production = score_from_matrix(request, build_event_contribution_matrix(request, static_contexts=contexts))
store = lib.FeatureStore()
built = build_event_contribution_matrix(
request, row_provider=lib.make_recording_provider(contexts, store, lib.d60_charts(contexts)), static_contexts=contexts,
)
research = score_from_matrix(request, built)
outcome = lib.reconcile_rows(production, research)
assert outcome["changed"] == 0, outcome
assert outcome["candidates"] == len(contexts) == 11
# extras were recorded alongside the production rules
recorded = {name for cells in store.rows.values() for cell in cells.values() for name in cell["extras"]}
assert any(name.startswith("kp_") for name in recorded) or any(name.endswith("_in_target_house") for name in recorded)