diff --git a/docs/BUG_HISTORY.md b/docs/BUG_HISTORY.md index bc4bd22a..0d25ff73 100644 --- a/docs/BUG_HISTORY.md +++ b/docs/BUG_HISTORY.md @@ -14688,6 +14688,21 @@ - 复发自:无 - 修复版本:未修 +## BUG-1091 | 生时校正打分只奖不罚、权重未校准、高分辨率信号未计分 + +- 状态:investigating(2026-09-29 研究单 `TASK-rectification-scoring-research-20260929` 脚手架已落地;判定按任务书须在 v5(≥60 例)上做,v4 上只跑通流水线、所有判定字段写 `pending_v5`) +- 首次发现 / 最近更新:2026-09-29 / 2026-09-29 +- 影响面:`scripts/active_rectification_event_engine.py::_score_event`(11 条加分规则,权重手拍)、`OBSERVATION_ONLY_LAYERS`(KP 宫头子主每分钟在算但不计分)、`_fine_minute_indices`(Pranapada / Hora / Ghati / Bhava lagna 已算未计分)、计分分盘无 D60。 +- 用户现象:开场说 12 件带年份经历、答约 10 张卡后范围整窗不动;引擎原始分头名簇命中在 v4 上只有 0.35 / 0.15 / 0.10(±10 / ±30 / ±60),约为随机水平的 4–5 倍。 +- 触发条件:任何一次校正。打分对候选「能解释事件」加分,对「预测了没发生的事」不扣分;一条规则对真分钟和错分钟同样常命中时照样 +2。 +- 根因(已确认):(1) 只奖不罚——负证据只来自选择题答「没有」(`core/build-state.ts::applyProbeOutcome`,±2);(2) 权重是人拍的,R1–R5 / V1 / V2 / V1n 都是「再拍一个权重看结果」,没有按数据校准;(3) KP / Pranapada / D60 这些分钟级信号没进计分,09-14 KP@默认权重更差是「权重没校准」的另一个症状。 +- 修复:未修。本轮只做离线脚手架:`scripts/research/scoring_research_lib.py`(特征记录 provider 与线上 `score_candidates` 逐分对账、似然比权重、leave-one-case-out、加权 provider、缺席 / 精度追问辅助)、`scripts/research/scoring_research.py`(R-A / R-B / R-C 流水线)、`tests/test_scoring_research.py`。生产代码与冻结文件零改动。 +- 验证:`tests/test_scoring_research.py`(含 v4 公开案例对账 0 差);v4 全集流水线跑通、`PYTHONHASHSEED=0` 两次复跑 JSON 逐字节一致(见 `docs/tasks/PROGRESS-rectification-scoring-research-20260929.md`)。 +- 防复发:任何计分改动必须先在 v5 上过任务书硬红线(六格真值在区间不降、留一法数字、稳健性三项),再按 ERR-110 重冻结;不得在 v4 上下结论(BUG-1090)。 +- 相关记录:BUG-560、BUG-1048、BUG-1089、BUG-1090 +- 复发自:无 +- 修复版本:未修(研究分支 `codex/rectification-scoring-research-20260929`) + ## BUG-1092 | 报告技术页标题出现「字段」,行星主星表头是看不懂的 RL / NL / SL / SS / SB - 状态:resolved(代码 + 回归测试;未部署,真机待合入 staging 后复核) diff --git a/docs/research/ACTIVE_FRONTS.md b/docs/research/ACTIVE_FRONTS.md index 764a8160..0c8d0aa6 100644 --- a/docs/research/ACTIVE_FRONTS.md +++ b/docs/research/ACTIVE_FRONTS.md @@ -121,3 +121,4 @@ Do not open new product surfaces before at least one of these four lanes is clos - 定论页(先读这个):`/docs/research/rectification_minute_resolution_closure_2026_09_14.md` - 两个已被证伪的假设:加权重能拉开候选;簇合并把候选并没了(合并从未触发)。 - 未解决的只剩「区间内部并列」,目前唯一被证明有效的信息源是用户提供的带年月新经历。 +- 2026-09-29 打分方法研究(脚手架,判定待 v5):`/docs/research/rectification_scoring_research_2026_09_29.md`——似然比校准权重 / 缺席证据 / 精度追问三条流水线已在 v4 上跑通,所有判定 `pending_v5`;任务书 `docs/tasks/TASK-rectification-scoring-research-20260929.md`,前置 `TASK-rectification-holdout-expansion-20260929.md`。 diff --git a/docs/research/rectification_minute_resolution_closure_2026_09_14.md b/docs/research/rectification_minute_resolution_closure_2026_09_14.md index a1293351..231f6a1b 100644 --- a/docs/research/rectification_minute_resolution_closure_2026_09_14.md +++ b/docs/research/rectification_minute_resolution_closure_2026_09_14.md @@ -94,3 +94,8 @@ - 任务书和 BUG-689 邀请语依据里的「1 分钟 ≈ 1.1 天大运边界位移」有误,实测中位 3.8 天(1.3–5.9)。45 天闸门对应 8–34 分钟,不是 40 分钟。 - 新增答错容错:答错 1 题,真值在区间 98–100%;答错 2 题,±30 / ±60 上 7–10% 的回放会挤出真值(两道反答 = 8 分 = 淘汰线)。 - 全文见 `docs/research/rectification_offline_research_2026_09_26.md`。 + +## 9. 后续(2026-09-29 打分方法研究单,脚手架阶段) + +- §1.3「加权重类改法全部不过门」说的是**人拍权重**。2026-09-29 研究单换了路线:每条规则的权重由数据估(似然比,leave-one-case-out),允许负权重(只奖不罚是结构性问题),KP / Pranapada / D60 作为特征进入由数据定权重;另测「缺席证据」与「对已说事件追问月份」。 +- 脚手架与 v4 流水线见 `docs/research/rectification_scoring_research_2026_09_29.md`;**判定必须在 v5(≥60 例)上做**,v4 上的数字只是调试参考(BUG-1090 / BUG-1091)。本页 §1–§7 的结论不因此改变。 diff --git a/docs/research/rectification_scoring_research_2026_09_29.md b/docs/research/rectification_scoring_research_2026_09_29.md new file mode 100644 index 00000000..e3648d75 --- /dev/null +++ b/docs/research/rectification_scoring_research_2026_09_29.md @@ -0,0 +1,139 @@ +# 生时校正打分方法研究:似然比权重 / 缺席证据 / 精度追问(脚手架阶段,2026-09-29) + +- 任务书:`docs/tasks/TASK-rectification-scoring-research-20260929.md`;进度:`docs/tasks/PROGRESS-rectification-scoring-research-20260929.md` +- 分支:`codex/rectification-scoring-research-20260929`,基线 `origin/staging` @ `6e393780`(开工时 head;任务书写作时为 `5febb111`,其间只有文档提交)。 +- 性质:**离线测量,不改线上代码、打分、阈值、出题闸门、Skill、冻结文件。** 研究计分器是 `precision_gate_lib.score_event_with_policy`(09-14 / 09-26 已证与线上逐分相同)加上不计分的"额外观察",通过 `build_event_contribution_matrix(row_provider=…)` 走线上同一条矩阵 / 出题 / 六题回放链。 +- 数据:`references/real_case_calibration/minute_rectification_holdout_v4.json`,20 例公开 Rodden-AA。**这是开放集回放,不是盲测;本页所有数字都是 v4 调试参考,任务书规定判定只能在 v5(≥60 例)上做,所以每一项的判定栏都是 `pending_v5`。** +- 口径:ayanamsa `raman`,node mode `mean`,步长 2 分钟,半径 ±10 / ±30 / ±60;六题回放(`ASK_COUNT = 6`,按真值作答,题目由线上 G0 出题器按**线上矩阵**出一次、各方案共用);交付区间 = 未淘汰且落后头名不足 8 分的簇并集。两种先验都报:`raw`(引擎行原始分,09-14 / 09-26 口径)与 `percent`(线上 `relative_support` 百分比口径;似然比方案用带温度的 softmax 后验,见 R-A)。 +- 复跑:`PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir <临时目录>`(约 14 分钟)→ `docs/research/scoring_research_{lr,absence,precision_followup}_2026_09_29.json`。同机两次运行 JSON 逐字节一致(见进度记录)。 +- 判门:`precision_gate_lib.gate_verdict`(头名不降**且**宽度下降、覆盖不降);JSON 里以 `debug_verdict_v4` 记录,`verdict` 字段一律 `pending_v5`。 + +## 结论 + +| 项 | 一句话结论(v4 调试参考) | 判定 | +| --- | --- | --- | +| R-0 脚手架 | 研究计分器与线上 `score_candidates` 在 20 例 × 3 档 = 60 个 case-radius 上**逐分对账 0 差**(分数与规则 id 都相同);每个 (候选, 事件, 采样日期) 记录 42 个特征(线上规则 25 个 + 额外观察 17 个) | 完成 | +| R-A 似然比权重 | 流水线跑通。v4 上:留一法 `raw` 口径头名 / 覆盖与基线持平或略动(±30 A3 头名 0.55 → 0.65、宽度 33 → 31),引擎自身头名(答题前)±10 0.30 → 0.40–0.45;`percent` 口径的 softmax 后验即使按训练集拟合温度仍过度自信,把真值挤出区间(±10 squeezed 3–5 例)。20 例上 42 个特征里 31–38 个的置信区间跨 0 | `pending_v5` | +| R-B 缺席证据 | 流水线跑通。v4 上按卡片同等强度(2 分)扣分会大量挤出真值(±60 覆盖 1.00 → 0.40)——**v4 的事件表是传记摘录,不是完整时间线,空白年不等于没发生**;真人口述同样不完整。0.5 分档不挤真值但头名下降 | `pending_v5`(v4 信号明确为负,v5 上若不翻转即关闭) | +| R-C 精度追问 | 流水线跑通。把日精度经历降成"只记得年"后,**每例平均只有 0.2 / 0.45 / 0.7 件经历在知道月份后能把候选分到两边**(±10 只有 2/20 例有);对这些例子把那一件恢复到月精度重算先验(`C1_prior`)在 ±30 上 7 例头名 0.71 vs 该 7 例的 C0,其余档位样本太少 | `pending_v5` | + +## R-0 · 脚手架与对账 + +- `scripts/research/scoring_research_lib.py` + - `make_recording_provider`:对每个 legacy 单事件请求,用 `score_event_with_policy`(V0 基线)+ 线上的过运 / Ashtakavarga / Shadbala 辅助(取 context 里线上算好的结果)产出**与线上完全相同**的 evidence,同时把该候选在该事件下的额外观察写进 `FeatureStore`: + - 线上规则(记为特征):`vim_{md,ad,pd}_domain_{house,lord,varga}`、`{label}_functional_{benefic,malefic}_auxiliary`、`narayana_{md,ad}_domain_house`、`{label}_arudha_auxiliary`、`controlled_transit_{jupiter,saturn}_domain_house`、`ashtakavarga_target_house_{support,pressure}_auxiliary`、`shadbala_sthana_drik_naisargika_{support,pressure}_auxiliary`。`event_kind:*`、`no_domain_activation`、`*_auxiliary_not_primary` 不作特征。 + - 额外观察(线上算了没计分):`kp_target_cusp_{sublord,starlord}_is_vim_{md,ad,pd}`、`kp_asc_sublord_is_vim_{md,ad,pd}`(`feature.kp_cusps.houses`,Placidus + Krishnamurti)、`vim_{md,ad,pd}_domain_varga_d60`(D60 现算)、`{pranapada,hora_lagna,ghati_lagna,bhava_lagna}_in_target_house`(`_fine_minute_indices`)。 + - `reconcile_rows`:逐分比较分数与规则 id。**60/60 case-radius 差异 0。** + - `event_feature_matrix`:特征值 = 该规则在该事件采样日期上命中的比例(年精度 12 个采样日、月 3 个、日 1 个),与线上按采样日平均分数同口径。 + - 真值标签:候选属于真值所在的**签名簇**(`raw_signature_clusters`)。 +- 测试:`tests/test_scoring_research.py`(10 条:LR 估计、留一法不泄漏、特征矩阵、规则过滤、加权 provider、缩放答案、缺席年、精度升降、百分比、以及 v4 公开案例 ±10 对账 0 差)。 + +## R-A · 似然比校准权重 + +**方法。** 对每个特征 f,`log LR_present = log P(f | 真值簇) − log P(f | 非真值簇)`,拉普拉斯平滑 α = 1(`(Σx + 1) / (N + 2)`);`log LR_absent` 为朴素贝叶斯互补项。每档半径单独估。 + +| 方案 | 特征 | 权重规则 | +| --- | --- | --- | +| A1 | 只用线上 25 条规则 | 只奖不罚:`max(log LR_present, 0)`,不用 absent 项 | +| A2 | 加 17 个额外观察 | 同上 | +| A3 | 同 A2 | 朴素贝叶斯:present + absent 两项都用,允许负权重 | + +候选分 = Σ 事件 Σ 特征(x·w_present + (1−x)·w_absent),经线上矩阵的事件类型系数、精度权重、采样日平均与大运边界邻近项(后者按线上原样附加,不重加权)。两个只从**训练折**估的标定量:`scale`(研究分的例内极差中位 → 线上极差中位,让 `raw` 回放的 ±2 / 8 分线意义不变)、`temperature`(Platt 单温度,最大化训练例真值簇的对数后验;网格 0.25–256),`percent` 先验 = softmax(分 / T) × 100。**留一法按案例留**(`loo_folds`),全集拟合另报只作参照。 + +### 六题回放(留一法;括号内为全集拟合) + +`raw` 口径: + +| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 引擎头名(答题前) | 挤出真值 | debug 判定 | +| --- | --- | ---: | ---: | ---: | ---: | ---: | --- | +| ±10 | 基线 | 0.80 | 1.00 | 15 | 0.30 | 0 | — | +| ±10 | A1 | 0.70 (0.75) | 1.00 | 14 | 0.45 | 0 | no_benefit | +| ±10 | A2 | 0.75 (0.80) | 1.00 | 15 | 0.40 (0.50) | 0 | no_benefit | +| ±10 | A3 | 0.80 (0.85) | 1.00 | 15 | 0.40 (0.50) | 0 | no_benefit | +| ±30 | 基线 | 0.55 | 1.00 | 33 | 0.20 | 0 | — | +| ±30 | A1 | 0.55 (0.70) | 1.00 | 32 | 0.25 | 0 | benefit(宽度 −1) | +| ±30 | A2 | 0.60 (0.65) | 1.00 | 34 | 0.10 (0.30) | 0 | no_benefit | +| ±30 | A3 | 0.65 (0.75) | 1.00 | 31 | 0.20 (0.45) | 0 | benefit(宽度 −2) | +| ±60 | 基线 | 0.40 | 1.00 | 56 | 0.05 | 0 | — | +| ±60 | A1 | 0.45 (0.50) | 1.00 | 61 | 0.15 (0.20) | 0 | no_benefit | +| ±60 | A2 | 0.55 (0.60) | 1.00 | 61 | 0.00 (0.10) | 0 | no_benefit | +| ±60 | A3 | 0.50 (0.70) | **0.95** | 61 | 0.05 (0.30) | 1 | no_benefit | + +`percent` 口径(基线 = 线上比例百分比;A* = 拟合温度的 softmax 后验): + +| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 挤出真值 | 温度(各折) | +| --- | --- | ---: | ---: | ---: | ---: | --- | +| ±10 | 基线 | 0.90 | 1.00 | 11 | 0 | — | +| ±10 | A1 / A2 / A3 | 0.60 / 0.50 / 0.55 | 0.80 / 0.85 / 0.75 | 7 / 7 / 5 | 4 / 3 / 5 | 1–2 / 1 / 0.5–1 | +| ±30 | 基线 | 0.85 | 1.00 | 33 | 0 | — | +| ±30 | A1 / A2 / A3 | 0.65 / 0.55 / 0.55 | 0.95 / 0.80 / 0.90 | 31 / 10 / 13 | 1 / 4 / 2 | 2–4 / 1–2 / 1–2 | +| ±60 | 基线 | 0.80 | 1.00 | 53 | 0 | — | +| ±60 | A1 / A2 / A3 | 0.65 / 0.50 / 0.40 | 1.00 / 0.80 / 0.80 | 57 / 32 / 38 | 0 / 4 / 4 | 4–8 / 2–4 / 2 | + +**读法(v4,仅供 v5 前设计参考,不是结论)。** + +1. 对账通过后,`raw` 口径的留一法数字与基线基本持平:似然比权重没有让真值掉出区间(±60 A3 一例除外),也没有明显收窄;引擎答题前的头名在 ±10 从 0.30 到 0.40–0.45、±60 A1 从 0.05 到 0.15,全集拟合再高一些——**全集与留一法的差距就是 20 例过拟合的量**。 +2. `percent` 口径下 softmax 后验**过度自信**:即使温度按训练折拟合,仍把 15–25% 例子的真值挤出区间。校准曲线(A3 留一法)在 ±10 上 [0.20,0.30) 桶预测 0.24 / 实际 0.40、[0.50,0.70) 预测 0.58 / 实际 0.40、[0.70,1.00) 只有 1 个点且错;±30 / ±60 高桶几乎无样本。要上线必须先解决后验的标定,而不是先验本身。 +3. **42 个特征里 31 / 36 / 38 个(±10 / ±30 / ±60)的 5–95% 置信区间跨 0**(案例级 bootstrap 200 次)。±10 上区间不跨 0 的:`ashtakavarga_target_house_pressure_auxiliary`(+0.41)、`vim_ad_domain_varga`(+0.30)、`vim_md_domain_varga`(+0.27)、`ghati_lagna_in_target_house` / `hora_lagna_in_target_house`(+0.31)、`pranapada_in_target_house`(**−0.36**)、`vim_ad_arudha_auxiliary`(−0.21)、`narayana_ad_domain_house`(−0.21)、`kp_target_cusp_sublord_is_vim_ad`(−0.19)。方向与线上手拍权重相反的几条(Arudha、Narayana AD、KP AD 子主为负)正是 20 例上说不清的地方——这就是任务书要先扩集的原因。 +4. 稳健性(留一法 A3,`raw`):±7 天抖动与 ±1–3 月平移三档真值在区间 0.95–1.00;答错 1 题 0.96–1.00、答错 2 题 0.91–0.99(与 09-26 基线同量级)。`percent` 口径下三项都更差(0.70–0.90),同样是后验过自信的表现。 + +## R-B · 缺席证据 + +**方法。** 只用训练事件(留一件 holdout 不用)。每个领域取第一件与最后一件带年月事件之间的年份,没有该领域事件的年份 = 缺席年。对每个缺席 (领域, 年) 用出题器同一判定函数 `_evaluate_contexts(month=None)` 在代表候选上求"这一年该领域是否激活"的分边;分得开就当作一道答了"没有"的题,按 0.5 / 1.0 / 2.0 分扣(卡片是 2.0;低于卡片强度的扣分不计强冲突、不会淘汰)。之后照常六题。漏说稳健性:随机删掉 1–2 件训练事件再算缺席年(每例 5 次,种子固定)。 + +`raw` 口径六题回放(`percent` 口径同向,见 JSON): + +| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 挤出真值 | 漏说 1 件(100 次回放)真值在区间 | 漏说 2 件 | +| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | +| ±10 | B0 | 0.80 | 1.00 | 15 | 0 | — | — | +| ±10 | 缺席 @0.5 分 | 0.70 | 1.00 | 15 | 0 | 1.00 | 1.00 | +| ±10 | 缺席 @1.0 分 | 0.70 | 1.00 | 14 | 0 | 1.00 | 1.00 | +| ±10 | 缺席 @2.0 分(=卡片) | 0.55 | **0.85** | 13 | 3 | 0.94 | 0.94 | +| ±30 | B0 | 0.55 | 1.00 | 33 | 0 | — | — | +| ±30 | 缺席 @0.5 分 | 0.50 | 1.00 | 34 | 0 | 1.00 | 1.00 | +| ±30 | 缺席 @1.0 分 | 0.45 | 1.00 | 33 | 0 | 1.00 | 1.00 | +| ±30 | 缺席 @2.0 分 | 0.25 | **0.70** | 19 | 6 | 0.80 | 0.87 | +| ±60 | B0 | 0.40 | 1.00 | 56 | 0 | — | — | +| ±60 | 缺席 @0.5 分 | 0.35 | 1.00 | 68 | 0 | 1.00 | 1.00 | +| ±60 | 缺席 @1.0 分 | 0.25 | 1.00 | 61 | 0 | 1.00 | 1.00 | +| ±60 | 缺席 @2.0 分 | 0.10 | **0.40** | 9 | 12 | 0.54 | 0.73 | + +每例平均缺席年 17.95;能分边的缺席题 2.05 / 4.30 / 6.15 道(有题的例子 12 / 17 / 17)。三档 debug 判定全部 no_benefit。 + +读法: + +1. **按卡片强度扣,真值被大量挤出**(±60 覆盖 1.00 → 0.40)。原因不是方法错,而是 v4 的事件表是传记摘录:那些"空白年"里公众人物多半确有该领域的事,只是没进数据集;扣分等于用不完整的时间线否定真值。真人口述同样不完整(用户不会把每年的事都说全),所以这条在 v5 上大概率也是负的——**若 v5 上不翻转,R-B 关闭**。 +2. 0.5 / 1.0 分档不挤真值(覆盖 1.00)但头名一律下降、宽度不降,等于只加噪声。漏说 1–2 件在低强度下不伤覆盖,在卡片强度下把覆盖压到 0.54–0.94。 + +## R-C · 精度追问可行性 + +**方法。** 把每例所有日精度经历降成"只记得年"(v4 没有月精度)作为口述形态 `C0_spoken`,先验按线上重算。对每件训练中的日精度经历,用 `_evaluate_contexts(month=真实月份)` 判断"知道月份后能否把代表候选分到两边";能分的按信息增益排序取前 1 / 2 件:`C{k}_prior`(那件恢复到月精度重算先验)、`C{k}_probe`(当作答"是"的卡)、`C{k}_both`。 + +| 半径 | 每例可追问件数均值 | ≥1 件的例子 | ≥2 件的例子 | +| --- | ---: | ---: | ---: | +| ±10 | 0.20 | 2 / 20 | 1 / 20 | +| ±30 | 0.45 | 7 / 20 | 1 / 20 | +| ±60 | 0.70 | 10 / 20 | 3 / 20 | + +| 半径 | 方案 | n | 头名命中(raw) | 真值在区间 | 宽度中位 | 同 n 例的 C0(raw 头名 / 宽度) | +| --- | --- | ---: | ---: | ---: | ---: | --- | +| ±10 | C1_prior / C1_probe | 2 | 1.00 / 0.50 | 1.00 | 16 | 见 JSON per-arm(样本太少) | +| ±30 | C1_prior | 7 | 0.71 | 1.00 | 31 | debug 判定 benefit(percent 口径头名 1.00) | +| ±30 | C1_probe | 7 | 0.57 | 1.00 | 33 | no_benefit | +| ±60 | C1_prior / C1_probe | 10 | 0.40 / 0.40 | 1.00 | 58 / 63 | no_benefit | + +**读法。** 可追问的经历很少(与 09-29 打字经历研究一致:绝大多数经历落在所有候选都一样的大运段),但在 ±30 上"恢复到月精度重算先验"(走线上已有的打字经历通道,不加卡片)对那 7 例有正向信号;"当作答是的卡"没有。样本太少,只能说值得在 v5 上按同口径重测;产品层实现(只对落在候选分歧点的那一件追问月份,算在定向追问 2 条额度内)要等 v5 判定。 + +## 给产品的话 + +- 三条流水线都跑通、与线上对账 0 差,**但没有任何一条在 v4 上能下结论**,20 例上 42 个特征有四分之三分不清正负。这一轮交付的是"能在 v5 上一键复跑的测量工具",不是新打分。 +- v4 上唯一方向明确的信号是负的:把口述时间线的空白当"没发生"会伤真值(R-B)。 +- v5 交付后的复跑命令:`PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --holdout references/real_case_calibration/minute_rectification_holdout_v5.json --dataset-label v5 --out-suffix v5_<日期>`。 + +## 附:复跑命令 + +```bash +PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir /tmp/scoring_ctx # 约 14 分钟,写三份 JSON +PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --limit 3 --radii 10 --fast --no-write # 开发用 +python3 -m pytest tests/test_scoring_research.py -q +``` diff --git a/docs/research/scoring_research_absence_2026_09_29.json b/docs/research/scoring_research_absence_2026_09_29.json new file mode 100644 index 00000000..96f0cb2c --- /dev/null +++ b/docs/research/scoring_research_absence_2026_09_29.json @@ -0,0 +1,1314 @@ +{ + "ask_count": 6, + "ayanamsa": "raman", + "card_points": 2.0, + "case_count": 20, + "dataset": "v4", + "errors": [], + "fast": false, + "generated_for": "TASK-rectification-scoring-research-20260929", + "holdout": "references/real_case_calibration/minute_rectification_holdout_v4.json", + "minute_step": 2, + "node_mode": "mean", + "omit_repeats": 5, + "open_set_not_blind": true, + "part": "R-B", + "points": [ + 0.5, + 1.0, + 2.0 + ], + "probe_today": "2026-09-16", + "probes": "production G0 probes generated once per case/radius from the production matrix; reused for every arm", + "radii": { + "10": { + "arms": { + "B0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0376, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0346, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "absence@0.5": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0382, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 14.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0342, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 15.0 + }, + "verdict": "pending_v5" + }, + "absence@1.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0325, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 12.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0266, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 14.0 + }, + "verdict": "pending_v5" + }, + "absence@2.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.85, + "engine_top1": 0.3, + "entropy": 1.4545, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.05, + "top1": 0.6, + "truth_eliminated": 4, + "width_median": 11.0 + }, + "raw": { + "coverage": 0.85, + "engine_top1": 0.3, + "entropy": 1.4533, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 4, + "width_median": 13.0 + }, + "verdict": "pending_v5" + } + }, + "gap_stats": { + "cases_with_probes": 12, + "mean_absence_probes": 2.05, + "mean_absent_years": 17.95 + }, + "n": 20, + "omission": { + "omit_1@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0382, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.04, + "top1": 0.83, + "truth_eliminated": 0, + "width_median": 13.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0343, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.71, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "omit_1@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0344, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.79, + "truth_eliminated": 0, + "width_median": 12.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0287, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.71, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "omit_1@2.0": { + "percent": { + "coverage": 0.94, + "engine_top1": 0.3, + "entropy": 1.6687, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.05, + "top1": 0.7, + "truth_eliminated": 10, + "width_median": 11.0 + }, + "raw": { + "coverage": 0.94, + "engine_top1": 0.3, + "entropy": 1.6673, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.0, + "top1": 0.62, + "truth_eliminated": 10, + "width_median": 15.0 + } + }, + "omit_2@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0376, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.04, + "top1": 0.85, + "truth_eliminated": 0, + "width_median": 13.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0338, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.74, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "omit_2@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0349, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.81, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0297, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.74, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "omit_2@2.0": { + "percent": { + "coverage": 0.94, + "engine_top1": 0.3, + "entropy": 1.7352, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.05, + "top1": 0.74, + "truth_eliminated": 8, + "width_median": 11.0 + }, + "raw": { + "coverage": 0.94, + "engine_top1": 0.3, + "entropy": 1.7338, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.0, + "top1": 0.67, + "truth_eliminated": 8, + "width_median": 15.0 + } + } + } + }, + "30": { + "arms": { + "B0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.151, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.85, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1646, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "absence@0.5": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1456, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1641, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 34.0 + }, + "verdict": "pending_v5" + }, + "absence@1.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1172, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1567, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.45, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "verdict": "pending_v5" + }, + "absence@2.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.7, + "engine_top1": 0.2, + "entropy": 1.8761, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.05, + "top1": 0.4, + "truth_eliminated": 6, + "width_median": 19.0 + }, + "raw": { + "coverage": 0.7, + "engine_top1": 0.2, + "entropy": 1.8822, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.0, + "top1": 0.25, + "truth_eliminated": 6, + "width_median": 19.0 + }, + "verdict": "pending_v5" + } + }, + "gap_stats": { + "cases_with_probes": 17, + "mean_absence_probes": 4.3, + "mean_absent_years": 17.95 + }, + "n": 20, + "omission": { + "omit_1@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.147, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.73, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1642, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "omit_1@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1213, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.68, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.159, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.44, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "omit_1@2.0": { + "percent": { + "coverage": 0.8, + "engine_top1": 0.2, + "entropy": 2.1629, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 20, + "tie": 0.04, + "top1": 0.47, + "truth_eliminated": 20, + "width_median": 27.0 + }, + "raw": { + "coverage": 0.8, + "engine_top1": 0.2, + "entropy": 2.1702, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 20, + "tie": 0.0, + "top1": 0.29, + "truth_eliminated": 20, + "width_median": 27.0 + } + }, + "omit_2@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1489, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.81, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1642, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.52, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "omit_2@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1359, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.78, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1616, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.49, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "omit_2@2.0": { + "percent": { + "coverage": 0.87, + "engine_top1": 0.2, + "entropy": 2.4897, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 13, + "tie": 0.04, + "top1": 0.62, + "truth_eliminated": 13, + "width_median": 31.0 + }, + "raw": { + "coverage": 0.87, + "engine_top1": 0.2, + "entropy": 2.4992, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 13, + "tie": 0.0, + "top1": 0.37, + "truth_eliminated": 13, + "width_median": 31.0 + } + } + } + }, + "60": { + "arms": { + "B0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9174, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 53.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9537, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 56.0 + } + }, + "absence@0.5": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.8196, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9473, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.35, + "truth_eliminated": 0, + "width_median": 68.0 + }, + "verdict": "pending_v5" + }, + "absence@1.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.6772, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.921, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.25, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "verdict": "pending_v5" + }, + "absence@2.0": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.4, + "engine_top1": 0.05, + "entropy": 1.7866, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 12, + "tie": 0.1, + "top1": 0.25, + "truth_eliminated": 13, + "width_median": 9.0 + }, + "raw": { + "coverage": 0.4, + "engine_top1": 0.05, + "entropy": 1.8021, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 12, + "tie": 0.0, + "top1": 0.1, + "truth_eliminated": 13, + "width_median": 9.0 + }, + "verdict": "pending_v5" + } + }, + "gap_stats": { + "cases_with_probes": 17, + "mean_absence_probes": 6.15, + "mean_absent_years": 17.95 + }, + "n": 20, + "omission": { + "omit_1@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.8369, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.61, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9481, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.37, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "omit_1@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.7299, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9258, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.28, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "omit_1@2.0": { + "percent": { + "coverage": 0.54, + "engine_top1": 0.05, + "entropy": 2.101, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 46, + "tie": 0.16, + "top1": 0.41, + "truth_eliminated": 51, + "width_median": 25.0 + }, + "raw": { + "coverage": 0.54, + "engine_top1": 0.05, + "entropy": 2.1191, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 46, + "tie": 0.0, + "top1": 0.18, + "truth_eliminated": 51, + "width_median": 25.0 + } + }, + "omit_2@0.5": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.8802, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.67, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9523, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.41, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "omit_2@1.0": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.8027, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.66, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9444, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.33, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "omit_2@2.0": { + "percent": { + "coverage": 0.73, + "engine_top1": 0.05, + "entropy": 2.7068, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 27, + "tie": 0.15, + "top1": 0.53, + "truth_eliminated": 28, + "width_median": 33.0 + }, + "raw": { + "coverage": 0.73, + "engine_top1": 0.05, + "entropy": 2.725, + "n": 100, + "refresh_mean": 0.0, + "squeezed": 27, + "tie": 0.0, + "top1": 0.28, + "truth_eliminated": 28, + "width_median": 39.0 + } + } + } + } + }, + "reconciliation": { + "cases": 60, + "rows": [ + { + "candidates": 11, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + } + ], + "unreconciled": 0 + }, + "seed": 20260929, + "verdict": "pending_v5" +} diff --git a/docs/research/scoring_research_lr_2026_09_29.json b/docs/research/scoring_research_lr_2026_09_29.json new file mode 100644 index 00000000..9a6c7bd9 --- /dev/null +++ b/docs/research/scoring_research_lr_2026_09_29.json @@ -0,0 +1,3837 @@ +{ + "alpha": 1.0, + "ask_count": 6, + "ayanamsa": "raman", + "case_count": 20, + "dataset": "v4", + "errors": [], + "fast": false, + "generated_for": "TASK-rectification-scoring-research-20260929", + "holdout": "references/real_case_calibration/minute_rectification_holdout_v4.json", + "minute_step": 2, + "node_mode": "mean", + "open_set_not_blind": true, + "part": "R-A", + "probe_today": "2026-09-16", + "probes": "production G0 probes generated once per case/radius from the production matrix; reused for every arm", + "radii": { + "10": { + "baseline": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0376, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0346, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "calibration": [ + { + "bin": "[0.00,0.05)", + "n": 19, + "observed_rate": 0.1053, + "predicted_mean": 0.0226 + }, + { + "bin": "[0.05,0.10)", + "n": 32, + "observed_rate": 0.0938, + "predicted_mean": 0.0675 + }, + { + "bin": "[0.10,0.20)", + "n": 37, + "observed_rate": 0.0811, + "predicted_mean": 0.1514 + }, + { + "bin": "[0.20,0.30)", + "n": 15, + "observed_rate": 0.4, + "predicted_mean": 0.24 + }, + { + "bin": "[0.30,0.50)", + "n": 12, + "observed_rate": 0.3333, + "predicted_mean": 0.3692 + }, + { + "bin": "[0.50,0.70)", + "n": 5, + "observed_rate": 0.4, + "predicted_mean": 0.578 + }, + { + "bin": "[0.70,1.00)", + "n": 1, + "observed_rate": 0.0, + "predicted_mean": 0.89 + } + ], + "full_fit": { + "A1": { + "percent": { + "coverage": 0.8, + "engine_top1": 0.45, + "entropy": 1.9993, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 6.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.45, + "entropy": 1.9733, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 14.0 + }, + "scale": 1.972191, + "temperature": 2.0 + }, + "A2": { + "percent": { + "coverage": 0.9, + "engine_top1": 0.5, + "entropy": 1.8004, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 6.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.9875, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 15.0 + }, + "scale": 1.379186, + "temperature": 1.0 + }, + "A3": { + "percent": { + "coverage": 0.8, + "engine_top1": 0.5, + "entropy": 1.791, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.5786, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.85, + "truth_eliminated": 0, + "width_median": 15.0 + }, + "scale": 1.118947, + "temperature": 1.0 + } + }, + "loo": { + "A1": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.8, + "engine_top1": 0.45, + "entropy": 1.9847, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 7.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.45, + "entropy": 1.9797, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 14.0 + }, + "temperatures": [ + 1.0, + 2.0 + ], + "verdict": "pending_v5" + }, + "A2": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.85, + "engine_top1": 0.4, + "entropy": 1.8745, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 7.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4, + "entropy": 1.9935, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 15.0 + }, + "temperatures": [ + 1.0 + ], + "verdict": "pending_v5" + }, + "A3": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.75, + "engine_top1": 0.4, + "entropy": 1.8429, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 5, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4, + "entropy": 1.5789, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 15.0 + }, + "temperatures": [ + 0.5, + 1.0 + ], + "verdict": "pending_v5" + } + }, + "lr_table": { + "ashtakavarga_target_house_pressure_auxiliary": { + "draws": 198, + "extra": false, + "fires_other": 27.0, + "fires_truth": 12.0, + "log_lr_absent": -0.01392, + "log_lr_present": 0.406865, + "p05": 0.1883, + "p95": 0.9269, + "p_other": 0.026794, + "p_truth": 0.040248 + }, + "ashtakavarga_target_house_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 642.0, + "fires_truth": 207.0, + "log_lr_absent": -0.0774, + "log_lr_present": 0.045513, + "p05": 0.0077, + "p95": 0.0835, + "p_other": 0.615311, + "p_truth": 0.643963 + }, + "bhava_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 197.0, + "fires_truth": 59.0, + "log_lr_absent": 0.004573, + "log_lr_present": -0.019803, + "p05": -0.2314, + "p95": 0.1766, + "p_other": 0.189474, + "p_truth": 0.185759 + }, + "controlled_transit_jupiter_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 193.6667, + "fires_truth": 59.3333, + "log_lr_absent": -0.000623, + "log_lr_present": 0.002716, + "p05": -0.1072, + "p95": 0.1176, + "p_other": 0.186284, + "p_truth": 0.18679 + }, + "controlled_transit_saturn_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 153.5833, + "fires_truth": 51.3333, + "log_lr_absent": -0.016682, + "log_lr_present": 0.09102, + "p05": -0.0718, + "p95": 0.246, + "p_other": 0.147927, + "p_truth": 0.162023 + }, + "ghati_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 142.0, + "fires_truth": 59.0, + "log_lr_absent": -0.058341, + "log_lr_present": 0.30562, + "p05": 0.0233, + "p95": 0.5434, + "p_other": 0.136842, + "p_truth": 0.185759 + }, + "hora_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 142.0, + "fires_truth": 59.0, + "log_lr_absent": -0.058341, + "log_lr_present": 0.30562, + "p05": 0.0233, + "p95": 0.5434, + "p_other": 0.136842, + "p_truth": 0.185759 + }, + "kp_asc_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 115.3333, + "fires_truth": 38.1667, + "log_lr_absent": -0.011243, + "log_lr_present": 0.085486, + "p05": -0.2308, + "p95": 0.3943, + "p_other": 0.111324, + "p_truth": 0.121259 + }, + "kp_asc_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 120.1667, + "fires_truth": 60.3333, + "log_lr_absent": -0.08734, + "log_lr_present": 0.493276, + "p05": -0.0258, + "p95": 0.8263, + "p_other": 0.115949, + "p_truth": 0.189886 + }, + "kp_asc_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 149.75, + "fires_truth": 47.4167, + "log_lr_absent": -0.006611, + "log_lr_present": 0.038341, + "p05": -0.2252, + "p95": 0.2252, + "p_other": 0.144258, + "p_truth": 0.149897 + }, + "kp_target_cusp_starlord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 200.5833, + "fires_truth": 60.6667, + "log_lr_absent": 0.002455, + "log_lr_present": -0.010339, + "p05": -0.1118, + "p95": 0.0662, + "p_other": 0.192903, + "p_truth": 0.190918 + }, + "kp_target_cusp_starlord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 233.0, + "fires_truth": 68.0, + "log_lr_absent": 0.013186, + "log_lr_present": -0.047095, + "p05": -0.21, + "p95": 0.1178, + "p_other": 0.223923, + "p_truth": 0.213622 + }, + "kp_target_cusp_starlord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 134.4167, + "fires_truth": 43.1667, + "log_lr_absent": -0.008253, + "log_lr_present": 0.053734, + "p05": -0.0906, + "p95": 0.1935, + "p_other": 0.129585, + "p_truth": 0.136739 + }, + "kp_target_cusp_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 231.3333, + "fires_truth": 58.1667, + "log_lr_absent": 0.049116, + "log_lr_present": -0.193695, + "p05": -0.3737, + "p95": -0.0597, + "p_other": 0.222329, + "p_truth": 0.183179 + }, + "kp_target_cusp_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 217.75, + "fires_truth": 77.6667, + "log_lr_absent": -0.044244, + "log_lr_present": 0.15141, + "p05": -0.1676, + "p95": 0.3785, + "p_other": 0.20933, + "p_truth": 0.24355 + }, + "kp_target_cusp_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 201.25, + "fires_truth": 68.0833, + "log_lr_absent": -0.025544, + "log_lr_present": 0.099929, + "p05": -0.1017, + "p95": 0.2437, + "p_other": 0.193541, + "p_truth": 0.21388 + }, + "narayana_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 140.5, + "fires_truth": 34.5, + "log_lr_absent": 0.029067, + "log_lr_present": -0.208647, + "p05": -0.4929, + "p95": -0.0116, + "p_other": 0.135407, + "p_truth": 0.109907 + }, + "narayana_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 171.75, + "fires_truth": 50.0, + "log_lr_absent": 0.008846, + "log_lr_present": -0.0459, + "p05": -0.1656, + "p95": 0.0157, + "p_other": 0.165311, + "p_truth": 0.157895 + }, + "pranapada_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 157.0, + "fires_truth": 33.0, + "log_lr_absent": 0.052702, + "log_lr_present": -0.362115, + "p05": -0.8036, + "p95": -0.0695, + "p_other": 0.151196, + "p_truth": 0.105263 + }, + "shadbala_sthana_drik_naisargika_pressure_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 639.6667, + "fires_truth": 198.5, + "log_lr_absent": -0.011879, + "log_lr_present": 0.007425, + "p05": -0.0508, + "p95": 0.0587, + "p_other": 0.613078, + "p_truth": 0.617647 + }, + "shadbala_sthana_drik_naisargika_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 397.1667, + "fires_truth": 121.0833, + "log_lr_absent": 0.004921, + "log_lr_present": -0.008047, + "p05": -0.0934, + "p95": 0.0795, + "p_other": 0.381021, + "p_truth": 0.377967 + }, + "vim_ad_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 109.75, + "fires_truth": 26.6667, + "log_lr_absent": 0.02248, + "log_lr_present": -0.212927, + "p05": -0.351, + "p95": -0.075, + "p_other": 0.105981, + "p_truth": 0.085655 + }, + "vim_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 159.5833, + "fires_truth": 52.1667, + "log_lr_absent": -0.013004, + "log_lr_present": 0.068738, + "p05": -0.0373, + "p95": 0.1944, + "p_other": 0.153668, + "p_truth": 0.164603 + }, + "vim_ad_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 229.9167, + "fires_truth": 69.8333, + "log_lr_absent": 0.002147, + "log_lr_present": -0.007607, + "p05": -0.0983, + "p95": 0.0669, + "p_other": 0.220973, + "p_truth": 0.219298 + }, + "vim_ad_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 207.0, + "fires_truth": 86.1667, + "log_lr_absent": -0.092579, + "log_lr_present": 0.304404, + "p05": 0.1603, + "p95": 0.4577, + "p_other": 0.199043, + "p_truth": 0.269866 + }, + "vim_ad_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 153.6667, + "fires_truth": 45.0, + "log_lr_absent": 0.006541, + "log_lr_present": -0.038511, + "p05": -0.3581, + "p95": 0.2665, + "p_other": 0.148006, + "p_truth": 0.142415 + }, + "vim_ad_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 404.1667, + "fires_truth": 116.1667, + "log_lr_absent": 0.039979, + "log_lr_present": -0.066581, + "p05": -0.127, + "p95": 0.0065, + "p_other": 0.387719, + "p_truth": 0.362745 + }, + "vim_ad_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 371.1667, + "fires_truth": 124.4167, + "log_lr_absent": -0.051217, + "log_lr_present": 0.08642, + "p05": -0.0198, + "p95": 0.1959, + "p_other": 0.35614, + "p_truth": 0.388287 + }, + "vim_md_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 80.0, + "fires_truth": 25.0, + "log_lr_absent": -0.003239, + "log_lr_present": 0.037767, + "p05": -0.1267, + "p95": 0.3235, + "p_other": 0.077512, + "p_truth": 0.080495 + }, + "vim_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 140.0, + "fires_truth": 45.6667, + "log_lr_absent": -0.011102, + "log_lr_present": 0.06839, + "p05": -0.055, + "p95": 0.232, + "p_other": 0.134928, + "p_truth": 0.144479 + }, + "vim_md_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 276.5, + "fires_truth": 87.0, + "log_lr_absent": -0.009433, + "log_lr_present": 0.025636, + "p05": -0.0708, + "p95": 0.1151, + "p_other": 0.26555, + "p_truth": 0.272446 + }, + "vim_md_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 216.0833, + "fires_truth": 86.6667, + "log_lr_absent": -0.08379, + "log_lr_present": 0.26738, + "p05": 0.091, + "p95": 0.4654, + "p_other": 0.207735, + "p_truth": 0.271414 + }, + "vim_md_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 159.6667, + "fires_truth": 54.0, + "log_lr_absent": -0.019727, + "log_lr_present": 0.102121, + "p05": -0.2417, + "p95": 0.4246, + "p_other": 0.153748, + "p_truth": 0.170279 + }, + "vim_md_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 214.5, + "fires_truth": 64.3333, + "log_lr_absent": 0.004963, + "log_lr_present": -0.019339, + "p05": -0.4451, + "p95": 0.2187, + "p_other": 0.20622, + "p_truth": 0.20227 + }, + "vim_md_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 564.3333, + "fires_truth": 187.1667, + "log_lr_absent": -0.094932, + "log_lr_present": 0.074032, + "p05": -0.0263, + "p95": 0.1856, + "p_other": 0.540989, + "p_truth": 0.582559 + }, + "vim_pd_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 82.0, + "fires_truth": 25.3333, + "log_lr_absent": -0.002285, + "log_lr_present": 0.026115, + "p05": -0.2546, + "p95": 0.264, + "p_other": 0.079426, + "p_truth": 0.081527 + }, + "vim_pd_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 153.5833, + "fires_truth": 50.75, + "log_lr_absent": -0.014529, + "log_lr_present": 0.079811, + "p05": -0.0906, + "p95": 0.2171, + "p_other": 0.147927, + "p_truth": 0.160217 + }, + "vim_pd_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 249.0833, + "fires_truth": 82.9167, + "log_lr_absent": -0.027305, + "log_lr_present": 0.08215, + "p05": -0.0075, + "p95": 0.1658, + "p_other": 0.239314, + "p_truth": 0.259804 + }, + "vim_pd_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 217.9167, + "fires_truth": 71.1667, + "log_lr_absent": -0.017787, + "log_lr_present": 0.064407, + "p05": -0.1377, + "p95": 0.2605, + "p_other": 0.20949, + "p_truth": 0.223426 + }, + "vim_pd_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 151.25, + "fires_truth": 50.9167, + "log_lr_absent": -0.01776, + "log_lr_present": 0.098236, + "p05": -0.1431, + "p95": 0.3184, + "p_other": 0.145694, + "p_truth": 0.160733 + }, + "vim_pd_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 357.5, + "fires_truth": 101.25, + "log_lr_absent": 0.039544, + "log_lr_present": -0.080388, + "p05": -0.1672, + "p95": -0.0122, + "p_other": 0.343062, + "p_truth": 0.316563 + }, + "vim_pd_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 332.8333, + "fires_truth": 106.5833, + "log_lr_absent": -0.020213, + "log_lr_present": 0.041744, + "p05": -0.0329, + "p95": 0.1148, + "p_other": 0.319458, + "p_truth": 0.333075 + } + }, + "n": 20, + "noise_features": [ + "bhava_lagna_in_target_house", + "controlled_transit_jupiter_domain_house", + "controlled_transit_saturn_domain_house", + "kp_asc_sublord_is_vim_ad", + "kp_asc_sublord_is_vim_md", + "kp_asc_sublord_is_vim_pd", + "kp_target_cusp_starlord_is_vim_ad", + "kp_target_cusp_starlord_is_vim_md", + "kp_target_cusp_starlord_is_vim_pd", + "kp_target_cusp_sublord_is_vim_md", + "kp_target_cusp_sublord_is_vim_pd", + "narayana_md_domain_house", + "shadbala_sthana_drik_naisargika_pressure_auxiliary", + "shadbala_sthana_drik_naisargika_support_auxiliary", + "vim_ad_domain_house", + "vim_ad_domain_lord", + "vim_ad_domain_varga_d60", + "vim_ad_functional_benefic_auxiliary", + "vim_ad_functional_malefic_auxiliary", + "vim_md_arudha_auxiliary", + "vim_md_domain_house", + "vim_md_domain_lord", + "vim_md_domain_varga_d60", + "vim_md_functional_benefic_auxiliary", + "vim_md_functional_malefic_auxiliary", + "vim_pd_arudha_auxiliary", + "vim_pd_domain_house", + "vim_pd_domain_lord", + "vim_pd_domain_varga", + "vim_pd_domain_varga_d60", + "vim_pd_functional_malefic_auxiliary" + ], + "robustness": { + "A1": { + "flip_1": { + "percent": { + "coverage": 0.8588, + "engine_top1": 0.4824, + "entropy": 1.7366, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 12, + "tie": 0.0, + "top1": 0.5529, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4824, + "entropy": 1.7439, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5529, + "truth_eliminated": 0, + "width_median": 13.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.8467, + "engine_top1": 0.4667, + "entropy": 1.4509, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 23, + "tie": 0.0333, + "top1": 0.4667, + "truth_eliminated": 0, + "width_median": 7.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4667, + "entropy": 1.4626, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3067, + "truth_eliminated": 0, + "width_median": 11.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.4, + "entropy": 1.9845, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 7.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4, + "entropy": 1.9769, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 12.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.8833, + "engine_top1": 0.3667, + "entropy": 1.9998, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 7, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 10.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3667, + "entropy": 1.984, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 13.0 + } + } + }, + "A2": { + "flip_1": { + "percent": { + "coverage": 0.7294, + "engine_top1": 0.4824, + "entropy": 1.6357, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 23, + "tie": 0.0, + "top1": 0.6118, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4824, + "entropy": 1.7522, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6706, + "truth_eliminated": 0, + "width_median": 13.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.6533, + "engine_top1": 0.4667, + "entropy": 1.3626, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 52, + "tie": 0.0, + "top1": 0.5467, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 0.9867, + "engine_top1": 0.4667, + "entropy": 1.4771, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.38, + "truth_eliminated": 0, + "width_median": 11.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.75, + "engine_top1": 0.3, + "entropy": 1.8711, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 5, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 6.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 1.9924, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.7, + "engine_top1": 0.25, + "entropy": 1.8559, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 18, + "tie": 0.0, + "top1": 0.5333, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 1.9943, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7333, + "truth_eliminated": 0, + "width_median": 15.0 + } + } + }, + "A3": { + "flip_1": { + "percent": { + "coverage": 0.7647, + "engine_top1": 0.4118, + "entropy": 1.5963, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 20, + "tie": 0.0, + "top1": 0.6353, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.4118, + "entropy": 1.6416, + "n": 85, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6706, + "truth_eliminated": 0, + "width_median": 13.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.6667, + "engine_top1": 0.4, + "entropy": 1.326, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 50, + "tie": 0.0, + "top1": 0.5667, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 0.9867, + "engine_top1": 0.4, + "entropy": 1.3898, + "n": 150, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.38, + "truth_eliminated": 0, + "width_median": 11.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.75, + "engine_top1": 0.3, + "entropy": 1.8436, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 5, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 1.5768, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.7333, + "engine_top1": 0.3333, + "entropy": 1.8368, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 16, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 5.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3333, + "entropy": 1.5645, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7833, + "truth_eliminated": 0, + "width_median": 15.0 + } + } + } + } + }, + "30": { + "baseline": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.151, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.85, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1646, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "calibration": [ + { + "bin": "[0.00,0.05)", + "n": 163, + "observed_rate": 0.0307, + "predicted_mean": 0.0257 + }, + { + "bin": "[0.05,0.10)", + "n": 124, + "observed_rate": 0.0484, + "predicted_mean": 0.0668 + }, + { + "bin": "[0.10,0.20)", + "n": 40, + "observed_rate": 0.175, + "predicted_mean": 0.1285 + }, + { + "bin": "[0.20,0.30)", + "n": 3, + "observed_rate": 0.6667, + "predicted_mean": 0.21 + }, + { + "bin": "[0.30,0.50)", + "n": 3, + "observed_rate": 0.0, + "predicted_mean": 0.41 + }, + { + "bin": "[0.50,0.70)", + "n": 1, + "observed_rate": 0.0, + "predicted_mean": 0.5 + }, + { + "bin": "[0.70,1.00)", + "n": 0, + "observed_rate": null, + "predicted_mean": null + } + ], + "full_fit": { + "A1": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 3.133, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 3.1392, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 32.0 + }, + "scale": 3.375292, + "temperature": 4.0 + }, + "A2": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.3, + "entropy": 3.004, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 7.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 3.1301, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "scale": 1.84633, + "temperature": 2.0 + }, + "A3": { + "percent": { + "coverage": 0.9, + "engine_top1": 0.45, + "entropy": 2.9852, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 12.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.45, + "entropy": 2.9242, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "scale": 1.506452, + "temperature": 2.0 + } + }, + "loo": { + "A1": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 0.95, + "engine_top1": 0.25, + "entropy": 3.1307, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 3.1414, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 32.0 + }, + "temperatures": [ + 2.0, + 4.0 + ], + "verdict": "pending_v5" + }, + "A2": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.8, + "engine_top1": 0.1, + "entropy": 3.0138, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 10.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.1341, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 34.0 + }, + "temperatures": [ + 1.0, + 2.0 + ], + "verdict": "pending_v5" + }, + "A3": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 0.9, + "engine_top1": 0.2, + "entropy": 3.0, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 13.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 2.9259, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "temperatures": [ + 1.0, + 2.0 + ], + "verdict": "pending_v5" + } + }, + "lr_table": { + "ashtakavarga_target_house_pressure_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 100.0, + "fires_truth": 12.0, + "log_lr_absent": -0.012009, + "log_lr_present": 0.339812, + "p05": 0.0424, + "p95": 0.7324, + "p_other": 0.028652, + "p_truth": 0.040248 + }, + "ashtakavarga_target_house_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2177.0, + "fires_truth": 207.0, + "log_lr_absent": -0.07072, + "log_lr_present": 0.041359, + "p05": -0.0024, + "p95": 0.0763, + "p_other": 0.617872, + "p_truth": 0.643963 + }, + "bhava_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 605.0, + "fires_truth": 59.0, + "log_lr_absent": -0.016859, + "log_lr_present": 0.077448, + "p05": -0.1477, + "p95": 0.2646, + "p_other": 0.171915, + "p_truth": 0.185759 + }, + "controlled_transit_jupiter_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 646.8333, + "fires_truth": 59.3333, + "log_lr_absent": -0.003692, + "log_lr_present": 0.016235, + "p05": -0.0804, + "p95": 0.1187, + "p_other": 0.183782, + "p_truth": 0.18679 + }, + "controlled_transit_saturn_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 548.1666, + "fires_truth": 51.3333, + "log_lr_absent": -0.007408, + "log_lr_present": 0.039215, + "p05": -0.1024, + "p95": 0.1725, + "p_other": 0.155792, + "p_truth": 0.162023 + }, + "ghati_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 511.0, + "fires_truth": 59.0, + "log_lr_absent": -0.048554, + "log_lr_present": 0.246003, + "p05": -0.0058, + "p95": 0.4429, + "p_other": 0.145248, + "p_truth": 0.185759 + }, + "hora_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 511.0, + "fires_truth": 59.0, + "log_lr_absent": -0.048554, + "log_lr_present": 0.246003, + "p05": -0.0058, + "p95": 0.4429, + "p_other": 0.145248, + "p_truth": 0.185759 + }, + "kp_asc_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 453.0, + "fires_truth": 38.1667, + "log_lr_absent": 0.008612, + "log_lr_present": -0.060288, + "p05": -0.4385, + "p95": 0.2757, + "p_other": 0.128794, + "p_truth": 0.121259 + }, + "kp_asc_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 382.6667, + "fires_truth": 60.3333, + "log_lr_absent": -0.095348, + "log_lr_present": 0.556533, + "p05": -0.104, + "p95": 1.0984, + "p_other": 0.108842, + "p_truth": 0.189886 + }, + "kp_asc_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 467.0, + "fires_truth": 47.4167, + "log_lr_absent": -0.019951, + "log_lr_present": 0.121359, + "p05": -0.2428, + "p95": 0.3892, + "p_other": 0.132766, + "p_truth": 0.149897 + }, + "kp_target_cusp_starlord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 660.25, + "fires_truth": 60.6667, + "log_lr_absent": -0.004107, + "log_lr_present": 0.017595, + "p05": -0.1232, + "p95": 0.1206, + "p_other": 0.187589, + "p_truth": 0.190918 + }, + "kp_target_cusp_starlord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 770.0833, + "fires_truth": 68.0, + "log_lr_absent": 0.006538, + "log_lr_present": -0.023707, + "p05": -0.2418, + "p95": 0.1482, + "p_other": 0.218747, + "p_truth": 0.213622 + }, + "kp_target_cusp_starlord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 499.8333, + "fires_truth": 43.1667, + "log_lr_absent": 0.006207, + "log_lr_present": -0.03832, + "p05": -0.2111, + "p95": 0.1437, + "p_other": 0.14208, + "p_truth": 0.136739 + }, + "kp_target_cusp_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 728.9167, + "fires_truth": 58.1667, + "log_lr_absent": 0.029684, + "log_lr_present": -0.122589, + "p05": -0.4276, + "p95": 0.0569, + "p_other": 0.207069, + "p_truth": 0.183179 + }, + "kp_target_cusp_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 656.75, + "fires_truth": 77.6667, + "log_lr_absent": -0.072592, + "log_lr_present": 0.266378, + "p05": -0.0871, + "p95": 0.5337, + "p_other": 0.186596, + "p_truth": 0.24355 + }, + "kp_target_cusp_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 717.0834, + "fires_truth": 68.0833, + "log_lr_absent": -0.012852, + "log_lr_present": 0.048711, + "p05": -0.1908, + "p95": 0.213, + "p_other": 0.203712, + "p_truth": 0.21388 + }, + "narayana_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 463.75, + "fires_truth": 34.5, + "log_lr_absent": 0.024954, + "log_lr_present": -0.181984, + "p05": -0.4537, + "p95": 0.0052, + "p_other": 0.131844, + "p_truth": 0.109907 + }, + "narayana_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 569.5834, + "fires_truth": 50.0, + "log_lr_absent": 0.004729, + "log_lr_present": -0.02485, + "p05": -0.143, + "p95": 0.0549, + "p_other": 0.161868, + "p_truth": 0.157895 + }, + "pranapada_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 501.0, + "fires_truth": 33.0, + "log_lr_absent": 0.042405, + "log_lr_present": -0.302256, + "p05": -0.7297, + "p95": -0.0528, + "p_other": 0.142411, + "p_truth": 0.105263 + }, + "shadbala_sthana_drik_naisargika_pressure_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2190.5833, + "fires_truth": 198.5, + "log_lr_absent": 0.010725, + "log_lr_present": -0.006582, + "p05": -0.0612, + "p95": 0.0496, + "p_other": 0.621726, + "p_truth": 0.617647 + }, + "shadbala_sthana_drik_naisargika_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1305.8333, + "fires_truth": 121.0833, + "log_lr_absent": -0.011563, + "log_lr_present": 0.019325, + "p05": -0.0825, + "p95": 0.1008, + "p_other": 0.370733, + "p_truth": 0.377967 + }, + "vim_ad_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 353.3333, + "fires_truth": 26.6667, + "log_lr_absent": 0.016391, + "log_lr_present": -0.160026, + "p05": -0.2859, + "p95": -0.0381, + "p_other": 0.10052, + "p_truth": 0.085655 + }, + "vim_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 540.1667, + "fires_truth": 52.1667, + "log_lr_absent": -0.013176, + "log_lr_present": 0.069688, + "p05": -0.0289, + "p95": 0.1841, + "p_other": 0.153522, + "p_truth": 0.164603 + }, + "vim_ad_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 815.3333, + "fires_truth": 69.8333, + "log_lr_absent": 0.015862, + "log_lr_present": -0.05451, + "p05": -0.1664, + "p95": 0.0185, + "p_other": 0.231584, + "p_truth": 0.219298 + }, + "vim_ad_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 682.1667, + "fires_truth": 86.1667, + "log_lr_absent": -0.099096, + "log_lr_present": 0.331067, + "p05": 0.1427, + "p95": 0.5103, + "p_other": 0.193806, + "p_truth": 0.269866 + }, + "vim_ad_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 534.9167, + "fires_truth": 45.0, + "log_lr_absent": 0.011279, + "log_lr_present": -0.065354, + "p05": -0.3476, + "p95": 0.1848, + "p_other": 0.152033, + "p_truth": 0.142415 + }, + "vim_ad_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1361.5, + "fires_truth": 116.1667, + "log_lr_absent": 0.03803, + "log_lr_present": -0.063496, + "p05": -0.1448, + "p95": 0.0135, + "p_other": 0.386525, + "p_truth": 0.362745 + }, + "vim_ad_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1246.8333, + "fires_truth": 124.4167, + "log_lr_absent": -0.054543, + "log_lr_present": 0.092461, + "p05": -0.0273, + "p95": 0.1918, + "p_other": 0.353995, + "p_truth": 0.388287 + }, + "vim_md_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 270.0, + "fires_truth": 25.0, + "log_lr_absent": -0.003925, + "log_lr_present": 0.045961, + "p05": -0.0973, + "p95": 0.2954, + "p_other": 0.076879, + "p_truth": 0.080495 + }, + "vim_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 479.6667, + "fires_truth": 45.6667, + "log_lr_absent": -0.009446, + "log_lr_present": 0.057839, + "p05": -0.0811, + "p95": 0.1992, + "p_other": 0.136359, + "p_truth": 0.144479 + }, + "vim_md_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 921.3333, + "fires_truth": 87.0, + "log_lr_absent": -0.014723, + "log_lr_present": 0.040414, + "p05": -0.0639, + "p95": 0.1361, + "p_other": 0.261655, + "p_truth": 0.272446 + }, + "vim_md_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 724.75, + "fires_truth": 86.6667, + "log_lr_absent": -0.08612, + "log_lr_present": 0.27632, + "p05": 0.0548, + "p95": 0.4848, + "p_other": 0.205887, + "p_truth": 0.271414 + }, + "vim_md_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 514.1667, + "fires_truth": 54.0, + "log_lr_absent": -0.02867, + "log_lr_present": 0.152826, + "p05": -0.1199, + "p95": 0.3922, + "p_other": 0.146147, + "p_truth": 0.170279 + }, + "vim_md_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 730.25, + "fires_truth": 64.3333, + "log_lr_absent": 0.00651, + "log_lr_present": -0.02527, + "p05": -0.436, + "p95": 0.224, + "p_other": 0.207447, + "p_truth": 0.20227 + }, + "vim_md_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1920.1667, + "fires_truth": 187.1667, + "log_lr_absent": -0.086129, + "log_lr_present": 0.066624, + "p05": -0.0489, + "p95": 0.1718, + "p_other": 0.545012, + "p_truth": 0.582559 + }, + "vim_pd_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 251.9167, + "fires_truth": 25.3333, + "log_lr_absent": -0.01059, + "log_lr_present": 0.127759, + "p05": -0.1884, + "p95": 0.3812, + "p_other": 0.071749, + "p_truth": 0.081527 + }, + "vim_pd_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 528.4167, + "fires_truth": 50.75, + "log_lr_absent": -0.01187, + "log_lr_present": 0.064632, + "p05": -0.1211, + "p95": 0.2197, + "p_other": 0.150189, + "p_truth": 0.160217 + }, + "vim_pd_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 817.25, + "fires_truth": 82.9167, + "log_lr_absent": -0.036708, + "log_lr_present": 0.11264, + "p05": -0.0431, + "p95": 0.2536, + "p_other": 0.232128, + "p_truth": 0.259804 + }, + "vim_pd_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 754.1667, + "fires_truth": 71.1667, + "log_lr_absent": -0.01177, + "log_lr_present": 0.042023, + "p05": -0.1441, + "p95": 0.229, + "p_other": 0.214232, + "p_truth": 0.223426 + }, + "vim_pd_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 536.3333, + "fires_truth": 50.9167, + "log_lr_absent": -0.009838, + "log_lr_present": 0.053005, + "p05": -0.1492, + "p95": 0.2272, + "p_other": 0.152435, + "p_truth": 0.160733 + }, + "vim_pd_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1289.0833, + "fires_truth": 101.25, + "log_lr_absent": 0.075055, + "log_lr_present": -0.145058, + "p05": -0.2501, + "p95": -0.0447, + "p_other": 0.365981, + "p_truth": 0.316563 + }, + "vim_pd_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1196.4167, + "fires_truth": 106.5833, + "log_lr_absent": 0.009972, + "log_lr_present": -0.019673, + "p05": -0.1355, + "p95": 0.1012, + "p_other": 0.339693, + "p_truth": 0.333075 + } + }, + "n": 20, + "noise_features": [ + "ashtakavarga_target_house_support_auxiliary", + "bhava_lagna_in_target_house", + "controlled_transit_jupiter_domain_house", + "controlled_transit_saturn_domain_house", + "ghati_lagna_in_target_house", + "hora_lagna_in_target_house", + "kp_asc_sublord_is_vim_ad", + "kp_asc_sublord_is_vim_md", + "kp_asc_sublord_is_vim_pd", + "kp_target_cusp_starlord_is_vim_ad", + "kp_target_cusp_starlord_is_vim_md", + "kp_target_cusp_starlord_is_vim_pd", + "kp_target_cusp_sublord_is_vim_ad", + "kp_target_cusp_sublord_is_vim_md", + "kp_target_cusp_sublord_is_vim_pd", + "narayana_ad_domain_house", + "narayana_md_domain_house", + "shadbala_sthana_drik_naisargika_pressure_auxiliary", + "shadbala_sthana_drik_naisargika_support_auxiliary", + "vim_ad_domain_house", + "vim_ad_domain_lord", + "vim_ad_domain_varga_d60", + "vim_ad_functional_benefic_auxiliary", + "vim_ad_functional_malefic_auxiliary", + "vim_md_arudha_auxiliary", + "vim_md_domain_house", + "vim_md_domain_lord", + "vim_md_domain_varga_d60", + "vim_md_functional_benefic_auxiliary", + "vim_md_functional_malefic_auxiliary", + "vim_pd_arudha_auxiliary", + "vim_pd_domain_house", + "vim_pd_domain_lord", + "vim_pd_domain_varga", + "vim_pd_domain_varga_d60", + "vim_pd_functional_malefic_auxiliary" + ], + "robustness": { + "A1": { + "flip_1": { + "percent": { + "coverage": 0.9444, + "engine_top1": 0.2778, + "entropy": 3.0241, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.0185, + "top1": 0.5093, + "truth_eliminated": 0, + "width_median": 32.0 + }, + "raw": { + "coverage": 0.9907, + "engine_top1": 0.2778, + "entropy": 3.0397, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.3889, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.8778, + "engine_top1": 0.2778, + "entropy": 2.752, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 22, + "tie": 0.0056, + "top1": 0.1778, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 0.9778, + "engine_top1": 0.2778, + "entropy": 2.7682, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.1167, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.95, + "engine_top1": 0.2, + "entropy": 3.1348, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1419, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 32.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.9833, + "engine_top1": 0.2, + "entropy": 3.1341, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1418, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6833, + "truth_eliminated": 0, + "width_median": 33.0 + } + } + }, + "A2": { + "flip_1": { + "percent": { + "coverage": 0.8519, + "engine_top1": 0.1111, + "entropy": 2.9474, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 16, + "tie": 0.0, + "top1": 0.5278, + "truth_eliminated": 0, + "width_median": 19.0 + }, + "raw": { + "coverage": 0.9815, + "engine_top1": 0.1111, + "entropy": 3.0352, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.4074, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.8611, + "engine_top1": 0.1111, + "entropy": 2.6772, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 25, + "tie": 0.0111, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 27.0 + }, + "raw": { + "coverage": 0.9444, + "engine_top1": 0.1111, + "entropy": 2.7639, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 10, + "tie": 0.0, + "top1": 0.1222, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.05, + "entropy": 3.0233, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 18.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.1348, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 34.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.8167, + "engine_top1": 0.1333, + "entropy": 3.0269, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 11, + "tie": 0.0, + "top1": 0.5167, + "truth_eliminated": 0, + "width_median": 12.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1333, + "entropy": 3.1354, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 33.0 + } + } + }, + "A3": { + "flip_1": { + "percent": { + "coverage": 0.8796, + "engine_top1": 0.2222, + "entropy": 2.9498, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 13, + "tie": 0.0, + "top1": 0.5185, + "truth_eliminated": 0, + "width_median": 23.0 + }, + "raw": { + "coverage": 0.9722, + "engine_top1": 0.2222, + "entropy": 2.9319, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.5093, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.8222, + "engine_top1": 0.2222, + "entropy": 2.6728, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 32, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 21.0 + }, + "raw": { + "coverage": 0.9444, + "engine_top1": 0.2222, + "entropy": 2.6317, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 10, + "tie": 0.0, + "top1": 0.1778, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.9, + "engine_top1": 0.15, + "entropy": 3.0061, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 2, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 18.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 2.9277, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 32.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.2167, + "entropy": 3.0173, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 9, + "tie": 0.0, + "top1": 0.5667, + "truth_eliminated": 0, + "width_median": 17.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2167, + "entropy": 2.9176, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 32.0 + } + } + } + } + }, + "60": { + "baseline": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9174, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 53.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9537, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 56.0 + } + }, + "calibration": [ + { + "bin": "[0.00,0.05)", + "n": 541, + "observed_rate": 0.0259, + "predicted_mean": 0.0204 + }, + { + "bin": "[0.05,0.10)", + "n": 89, + "observed_rate": 0.0674, + "predicted_mean": 0.0596 + }, + { + "bin": "[0.10,0.20)", + "n": 14, + "observed_rate": 0.0, + "predicted_mean": 0.1329 + }, + { + "bin": "[0.20,0.30)", + "n": 7, + "observed_rate": 0.0, + "predicted_mean": 0.2443 + }, + { + "bin": "[0.30,0.50)", + "n": 0, + "observed_rate": null, + "predicted_mean": null + }, + { + "bin": "[0.50,0.70)", + "n": 0, + "observed_rate": null, + "predicted_mean": null + }, + { + "bin": "[0.70,1.00)", + "n": 0, + "observed_rate": null, + "predicted_mean": null + } + ], + "full_fit": { + "A1": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.9069, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 54.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.9317, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "scale": 3.74725, + "temperature": 4.0 + }, + "A2": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.1, + "entropy": 3.8143, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 28.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.9313, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "scale": 2.400103, + "temperature": 2.0 + }, + "A3": { + "percent": { + "coverage": 0.8, + "engine_top1": 0.3, + "entropy": 3.7963, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 19.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 3.6958, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 56.0 + }, + "scale": 1.790734, + "temperature": 2.0 + } + }, + "loo": { + "A1": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 3.9106, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.65, + "truth_eliminated": 0, + "width_median": 57.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 3.9344, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.45, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "temperatures": [ + 4.0, + 8.0 + ], + "verdict": "pending_v5" + }, + "A2": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.8, + "engine_top1": 0.0, + "entropy": 3.8427, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 32.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.9353, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "temperatures": [ + 2.0, + 4.0 + ], + "verdict": "pending_v5" + }, + "A3": { + "debug_verdict_v4": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 0.8, + "engine_top1": 0.05, + "entropy": 3.8133, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 38.0 + }, + "raw": { + "coverage": 0.95, + "engine_top1": 0.05, + "entropy": 3.703, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "temperatures": [ + 2.0 + ], + "verdict": "pending_v5" + } + }, + "lr_table": { + "ashtakavarga_target_house_pressure_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 200.0, + "fires_truth": 12.0, + "log_lr_absent": -0.012945, + "log_lr_present": 0.372059, + "p05": -0.0234, + "p95": 0.7343, + "p_other": 0.027743, + "p_truth": 0.040248 + }, + "ashtakavarga_target_house_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 4437.0, + "fires_truth": 207.0, + "log_lr_absent": -0.084525, + "log_lr_present": 0.049994, + "p05": 0.0051, + "p95": 0.0811, + "p_other": 0.61256, + "p_truth": 0.643963 + }, + "bhava_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 1121.0, + "fires_truth": 59.0, + "log_lr_absent": -0.037239, + "log_lr_present": 0.181891, + "p05": -0.0095, + "p95": 0.362, + "p_other": 0.154865, + "p_truth": 0.185759 + }, + "controlled_transit_jupiter_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 1301.9166, + "fires_truth": 59.3333, + "log_lr_absent": -0.008515, + "log_lr_present": 0.037939, + "p05": -0.0689, + "p95": 0.1491, + "p_other": 0.179837, + "p_truth": 0.18679 + }, + "controlled_transit_saturn_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 1123.25, + "fires_truth": 51.3333, + "log_lr_absent": -0.008137, + "log_lr_present": 0.043177, + "p05": -0.1232, + "p95": 0.1808, + "p_other": 0.155176, + "p_truth": 0.162023 + }, + "ghati_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 1115.0, + "fires_truth": 59.0, + "log_lr_absent": -0.038218, + "log_lr_present": 0.187253, + "p05": -0.1753, + "p95": 0.4423, + "p_other": 0.154037, + "p_truth": 0.185759 + }, + "hora_lagna_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 1115.0, + "fires_truth": 59.0, + "log_lr_absent": -0.038218, + "log_lr_present": 0.187253, + "p05": -0.1753, + "p95": 0.4423, + "p_other": 0.154037, + "p_truth": 0.185759 + }, + "kp_asc_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 968.3333, + "fires_truth": 38.1667, + "log_lr_absent": 0.014367, + "log_lr_present": -0.098368, + "p05": -0.461, + "p95": 0.2379, + "p_other": 0.133793, + "p_truth": 0.121259 + }, + "kp_asc_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 890.25, + "fires_truth": 60.3333, + "log_lr_absent": -0.079315, + "log_lr_present": 0.434113, + "p05": -0.1927, + "p95": 0.9596, + "p_other": 0.123016, + "p_truth": 0.189886 + }, + "kp_asc_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 947.1667, + "fires_truth": 47.4167, + "log_lr_absent": -0.022133, + "log_lr_present": 0.135728, + "p05": -0.2531, + "p95": 0.4058, + "p_other": 0.130872, + "p_truth": 0.149897 + }, + "kp_target_cusp_starlord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 1399.5, + "fires_truth": 60.6667, + "log_lr_absent": 0.002955, + "log_lr_present": -0.012427, + "p05": -0.2017, + "p95": 0.1241, + "p_other": 0.193306, + "p_truth": 0.190918 + }, + "kp_target_cusp_starlord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 1587.0833, + "fires_truth": 68.0, + "log_lr_absent": 0.007115, + "log_lr_present": -0.025762, + "p05": -0.3339, + "p95": 0.2173, + "p_other": 0.219197, + "p_truth": 0.213622 + }, + "kp_target_cusp_starlord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 1120.8334, + "fires_truth": 43.1667, + "log_lr_absent": 0.021194, + "log_lr_present": -0.124335, + "p05": -0.3359, + "p95": 0.0417, + "p_other": 0.154842, + "p_truth": 0.136739 + }, + "kp_target_cusp_sublord_is_vim_ad": { + "draws": 200, + "extra": true, + "fires_other": 1510.5833, + "fires_truth": 58.1667, + "log_lr_absent": 0.031665, + "log_lr_present": -0.13014, + "p05": -0.3839, + "p95": 0.0282, + "p_other": 0.208638, + "p_truth": 0.183179 + }, + "kp_target_cusp_sublord_is_vim_md": { + "draws": 200, + "extra": true, + "fires_other": 1432.5, + "fires_truth": 77.6667, + "log_lr_absent": -0.058646, + "log_lr_present": 0.20776, + "p05": -0.114, + "p95": 0.4384, + "p_other": 0.197861, + "p_truth": 0.24355 + }, + "kp_target_cusp_sublord_is_vim_pd": { + "draws": 200, + "extra": true, + "fires_other": 1492.0834, + "fires_truth": 68.0833, + "log_lr_absent": -0.009868, + "log_lr_present": 0.037129, + "p05": -0.1873, + "p95": 0.2145, + "p_other": 0.206085, + "p_truth": 0.21388 + }, + "narayana_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 980.4167, + "fires_truth": 34.5, + "log_lr_absent": 0.02913, + "log_lr_present": -0.20905, + "p05": -0.5102, + "p95": 0.0, + "p_other": 0.135461, + "p_truth": 0.109907 + }, + "narayana_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 1163.8334, + "fires_truth": 50.0, + "log_lr_absent": 0.003429, + "log_lr_present": -0.018093, + "p05": -0.1467, + "p95": 0.0852, + "p_other": 0.160778, + "p_truth": 0.157895 + }, + "pranapada_in_target_house": { + "draws": 200, + "extra": true, + "fires_other": 1088.0, + "fires_truth": 33.0, + "log_lr_absent": 0.051659, + "log_lr_present": -0.35624, + "p05": -0.8298, + "p95": -0.0547, + "p_other": 0.150311, + "p_truth": 0.105263 + }, + "shadbala_sthana_drik_naisargika_pressure_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 4507.5, + "fires_truth": 198.5, + "log_lr_absent": 0.012221, + "log_lr_present": -0.007491, + "p05": -0.0736, + "p95": 0.0625, + "p_other": 0.622291, + "p_truth": 0.617647 + }, + "shadbala_sthana_drik_naisargika_support_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2684.0833, + "fires_truth": 121.0833, + "log_lr_absent": -0.011755, + "log_lr_present": 0.019651, + "p05": -0.1107, + "p95": 0.1179, + "p_other": 0.370612, + "p_truth": 0.377967 + }, + "vim_ad_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 646.6667, + "fires_truth": 26.6667, + "log_lr_absent": 0.004098, + "log_lr_present": -0.042733, + "p05": -0.2608, + "p95": 0.0895, + "p_other": 0.089395, + "p_truth": 0.085655 + }, + "vim_ad_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 1134.3333, + "fires_truth": 52.1667, + "log_lr_absent": -0.009408, + "log_lr_present": 0.049165, + "p05": -0.1036, + "p95": 0.1541, + "p_other": 0.156706, + "p_truth": 0.164603 + }, + "vim_ad_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 1690.5833, + "fires_truth": 69.8333, + "log_lr_absent": 0.018336, + "log_lr_present": -0.062676, + "p05": -0.25, + "p95": 0.0477, + "p_other": 0.233483, + "p_truth": 0.219298 + }, + "vim_ad_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 1587.5833, + "fires_truth": 86.1667, + "log_lr_absent": -0.067006, + "log_lr_present": 0.207639, + "p05": 0.0249, + "p95": 0.3774, + "p_other": 0.219266, + "p_truth": 0.269866 + }, + "vim_ad_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 1101.8333, + "fires_truth": 45.0, + "log_lr_absent": 0.011499, + "log_lr_present": -0.066582, + "p05": -0.3334, + "p95": 0.1863, + "p_other": 0.15222, + "p_truth": 0.142415 + }, + "vim_ad_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2757.5, + "fires_truth": 116.1667, + "log_lr_absent": 0.028653, + "log_lr_present": -0.04843, + "p05": -0.1616, + "p95": 0.0351, + "p_other": 0.380745, + "p_truth": 0.362745 + }, + "vim_ad_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2521.3333, + "fires_truth": 124.4167, + "log_lr_absent": -0.063554, + "log_lr_present": 0.109116, + "p05": -0.0164, + "p95": 0.2197, + "p_other": 0.348148, + "p_truth": 0.388287 + }, + "vim_md_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 560.0, + "fires_truth": 25.0, + "log_lr_absent": -0.003325, + "log_lr_present": 0.03879, + "p05": -0.1978, + "p95": 0.2129, + "p_other": 0.077433, + "p_truth": 0.080495 + }, + "vim_md_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 965.6667, + "fires_truth": 45.6667, + "log_lr_absent": -0.012837, + "log_lr_present": 0.079591, + "p05": -0.0629, + "p95": 0.1899, + "p_other": 0.133425, + "p_truth": 0.144479 + }, + "vim_md_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 1887.75, + "fires_truth": 87.0, + "log_lr_absent": -0.016019, + "log_lr_present": 0.044081, + "p05": -0.0756, + "p95": 0.1651, + "p_other": 0.260697, + "p_truth": 0.272446 + }, + "vim_md_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 1529.25, + "fires_truth": 86.6667, + "log_lr_absent": -0.079388, + "log_lr_present": 0.25077, + "p05": 0.0124, + "p95": 0.4643, + "p_other": 0.211215, + "p_truth": 0.271414 + }, + "vim_md_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 1072.9167, + "fires_truth": 54.0, + "log_lr_absent": -0.026228, + "log_lr_present": 0.13868, + "p05": -0.1143, + "p95": 0.3613, + "p_other": 0.148229, + "p_truth": 0.170279 + }, + "vim_md_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 1661.6667, + "fires_truth": 64.3333, + "log_lr_absent": 0.034719, + "log_lr_present": -0.126261, + "p05": -0.5659, + "p95": 0.137, + "p_other": 0.229492, + "p_truth": 0.20227 + }, + "vim_md_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 3854.25, + "fires_truth": 187.1667, + "log_lr_absent": -0.114057, + "log_lr_present": 0.090551, + "p05": -0.0316, + "p95": 0.2159, + "p_other": 0.532126, + "p_truth": 0.582559 + }, + "vim_pd_arudha_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 537.6667, + "fires_truth": 25.3333, + "log_lr_absent": -0.007784, + "log_lr_present": 0.092153, + "p05": -0.2921, + "p95": 0.3608, + "p_other": 0.07435, + "p_truth": 0.081527 + }, + "vim_pd_domain_house": { + "draws": 200, + "extra": false, + "fires_other": 1072.0833, + "fires_truth": 50.75, + "log_lr_absent": -0.014309, + "log_lr_present": 0.078548, + "p05": -0.1158, + "p95": 0.2466, + "p_other": 0.148114, + "p_truth": 0.160217 + }, + "vim_pd_domain_lord": { + "draws": 200, + "extra": false, + "fires_other": 1615.0833, + "fires_truth": 82.9167, + "log_lr_absent": -0.048446, + "log_lr_present": 0.152478, + "p05": -0.0105, + "p95": 0.3008, + "p_other": 0.223062, + "p_truth": 0.259804 + }, + "vim_pd_domain_varga": { + "draws": 200, + "extra": false, + "fires_other": 1596.0833, + "fires_truth": 71.1667, + "log_lr_absent": -0.003839, + "log_lr_present": 0.013459, + "p05": -0.1839, + "p95": 0.2082, + "p_other": 0.220439, + "p_truth": 0.223426 + }, + "vim_pd_domain_varga_d60": { + "draws": 200, + "extra": true, + "fires_other": 1102.9167, + "fires_truth": 50.9167, + "log_lr_absent": -0.009916, + "log_lr_present": 0.053435, + "p05": -0.124, + "p95": 0.2272, + "p_other": 0.152369, + "p_truth": 0.160733 + }, + "vim_pd_functional_benefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2581.5833, + "fires_truth": 101.25, + "log_lr_absent": 0.060156, + "log_lr_present": -0.11871, + "p05": -0.2421, + "p95": 0.0026, + "p_other": 0.356464, + "p_truth": 0.316563 + }, + "vim_pd_functional_malefic_auxiliary": { + "draws": 200, + "extra": false, + "fires_other": 2563.5833, + "fires_truth": 106.5833, + "log_lr_absent": 0.031846, + "log_lr_present": -0.060871, + "p05": -0.2179, + "p95": 0.0682, + "p_other": 0.35398, + "p_truth": 0.333075 + } + }, + "n": 20, + "noise_features": [ + "ashtakavarga_target_house_pressure_auxiliary", + "bhava_lagna_in_target_house", + "controlled_transit_jupiter_domain_house", + "controlled_transit_saturn_domain_house", + "ghati_lagna_in_target_house", + "hora_lagna_in_target_house", + "kp_asc_sublord_is_vim_ad", + "kp_asc_sublord_is_vim_md", + "kp_asc_sublord_is_vim_pd", + "kp_target_cusp_starlord_is_vim_ad", + "kp_target_cusp_starlord_is_vim_md", + "kp_target_cusp_starlord_is_vim_pd", + "kp_target_cusp_sublord_is_vim_ad", + "kp_target_cusp_sublord_is_vim_md", + "kp_target_cusp_sublord_is_vim_pd", + "narayana_ad_domain_house", + "narayana_md_domain_house", + "shadbala_sthana_drik_naisargika_pressure_auxiliary", + "shadbala_sthana_drik_naisargika_support_auxiliary", + "vim_ad_arudha_auxiliary", + "vim_ad_domain_house", + "vim_ad_domain_lord", + "vim_ad_domain_varga_d60", + "vim_ad_functional_benefic_auxiliary", + "vim_ad_functional_malefic_auxiliary", + "vim_md_arudha_auxiliary", + "vim_md_domain_house", + "vim_md_domain_lord", + "vim_md_domain_varga_d60", + "vim_md_functional_benefic_auxiliary", + "vim_md_functional_malefic_auxiliary", + "vim_pd_arudha_auxiliary", + "vim_pd_domain_house", + "vim_pd_domain_lord", + "vim_pd_domain_varga", + "vim_pd_domain_varga_d60", + "vim_pd_functional_benefic_auxiliary", + "vim_pd_functional_malefic_auxiliary" + ], + "robustness": { + "A1": { + "flip_1": { + "percent": { + "coverage": 0.9907, + "engine_top1": 0.1667, + "entropy": 3.8947, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.3889, + "truth_eliminated": 0, + "width_median": 75.0 + }, + "raw": { + "coverage": 0.9907, + "engine_top1": 0.1667, + "entropy": 3.9126, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.287, + "truth_eliminated": 0, + "width_median": 77.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.8944, + "engine_top1": 0.1667, + "entropy": 3.6107, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 19, + "tie": 0.0111, + "top1": 0.1, + "truth_eliminated": 0, + "width_median": 69.0 + }, + "raw": { + "coverage": 0.9278, + "engine_top1": 0.1667, + "entropy": 3.6264, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 13, + "tie": 0.0, + "top1": 0.0611, + "truth_eliminated": 0, + "width_median": 71.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.9109, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 54.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.9347, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.45, + "truth_eliminated": 0, + "width_median": 73.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.0667, + "entropy": 3.9114, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7833, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0667, + "entropy": 3.9338, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 61.0 + } + } + }, + "A2": { + "flip_1": { + "percent": { + "coverage": 0.7963, + "engine_top1": 0.0, + "entropy": 3.8452, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 22, + "tie": 0.0093, + "top1": 0.3704, + "truth_eliminated": 0, + "width_median": 62.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.9115, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3333, + "truth_eliminated": 0, + "width_median": 77.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.7611, + "engine_top1": 0.0, + "entropy": 3.5592, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 43, + "tie": 0.0, + "top1": 0.1333, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 0.9111, + "engine_top1": 0.0, + "entropy": 3.6239, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 16, + "tie": 0.0, + "top1": 0.0833, + "truth_eliminated": 0, + "width_median": 71.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.85, + "engine_top1": 0.0, + "entropy": 3.8459, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 3, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 37.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.9356, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.8667, + "engine_top1": 0.0833, + "entropy": 3.8484, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 8, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 40.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0833, + "entropy": 3.935, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 61.0 + } + } + }, + "A3": { + "flip_1": { + "percent": { + "coverage": 0.787, + "engine_top1": 0.0556, + "entropy": 3.8303, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 23, + "tie": 0.0093, + "top1": 0.3241, + "truth_eliminated": 0, + "width_median": 57.0 + }, + "raw": { + "coverage": 0.963, + "engine_top1": 0.0556, + "entropy": 3.7475, + "n": 108, + "refresh_mean": 0.0, + "squeezed": 4, + "tie": 0.0, + "top1": 0.2685, + "truth_eliminated": 0, + "width_median": 77.0 + } + }, + "flip_2": { + "percent": { + "coverage": 0.7111, + "engine_top1": 0.0556, + "entropy": 3.5437, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 52, + "tie": 0.0, + "top1": 0.1167, + "truth_eliminated": 0, + "width_median": 59.0 + }, + "raw": { + "coverage": 0.9056, + "engine_top1": 0.0556, + "entropy": 3.457, + "n": 180, + "refresh_mean": 0.0, + "squeezed": 17, + "tie": 0.0, + "top1": 0.0556, + "truth_eliminated": 0, + "width_median": 71.0 + } + }, + "jitter_7d": { + "percent": { + "coverage": 0.7, + "engine_top1": 0.05, + "entropy": 3.8116, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 6, + "tie": 0.0, + "top1": 0.3, + "truth_eliminated": 0, + "width_median": 32.0 + }, + "raw": { + "coverage": 0.95, + "engine_top1": 0.05, + "entropy": 3.6992, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 1, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "shift_1_3_months": { + "percent": { + "coverage": 0.7833, + "engine_top1": 0.15, + "entropy": 3.8162, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 13, + "tie": 0.0, + "top1": 0.45, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 3.6825, + "n": 60, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 61.0 + } + } + } + } + } + }, + "reconciliation": { + "cases": 60, + "rows": [ + { + "candidates": 11, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + } + ], + "unreconciled": 0 + }, + "seed": 20260929, + "variants": [ + "A1", + "A2", + "A3" + ], + "verdict": "pending_v5" +} diff --git a/docs/research/scoring_research_precision_followup_2026_09_29.json b/docs/research/scoring_research_precision_followup_2026_09_29.json new file mode 100644 index 00000000..a2ab15f5 --- /dev/null +++ b/docs/research/scoring_research_precision_followup_2026_09_29.json @@ -0,0 +1,1206 @@ +{ + "ask_count": 6, + "ayanamsa": "raman", + "case_count": 20, + "dataset": "v4", + "errors": [], + "fast": false, + "generated_for": "TASK-rectification-scoring-research-20260929", + "holdout": "references/real_case_calibration/minute_rectification_holdout_v4.json", + "minute_step": 2, + "node_mode": "mean", + "open_set_not_blind": true, + "part": "R-C", + "probe_today": "2026-09-16", + "probes": "production G0 probes generated once per case/radius from the production matrix; reused for every arm", + "radii": { + "10": { + "arms": { + "B0_true_precision": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0376, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0346, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "C0_spoken": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0396, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 12.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3, + "entropy": 2.0286, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.75, + "truth_eliminated": 0, + "width_median": 15.0 + } + }, + "C1_both": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7906, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7428, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "verdict": "pending_v5" + }, + "C1_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7905, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7799, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "verdict": "pending_v5" + }, + "C1_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7908, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.5, + "entropy": 1.7395, + "n": 2, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5, + "truth_eliminated": 0, + "width_median": 16.0 + }, + "verdict": "pending_v5" + }, + "C2_both": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 1.5789, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 1.5743, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "verdict": "pending_v5" + }, + "C2_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 1.5848, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 1.5849, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "verdict": "pending_v5" + }, + "C2_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 1.0, + "entropy": 1.5808, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 1.0, + "entropy": 1.5755, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 11.0 + }, + "verdict": "pending_v5" + } + }, + "askable_per_case": { + "cases_with_at_least_one": 2, + "cases_with_at_least_two": 1, + "distribution": { + "0": 18, + "1": 1, + "3": 1 + }, + "mean": 0.2 + }, + "n": 20 + }, + "30": { + "arms": { + "B0_true_precision": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.151, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.85, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.2, + "entropy": 3.1646, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.55, + "truth_eliminated": 0, + "width_median": 33.0 + } + }, + "C0_spoken": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 3.1519, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.05, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 34.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.25, + "entropy": 3.1585, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6, + "truth_eliminated": 0, + "width_median": 35.0 + } + }, + "C1_both": { + "debug_verdict_v4_vs_C0": { + "percent": "benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0509, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0578, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7143, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "verdict": "pending_v5" + }, + "C1_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0781, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0875, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7143, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "verdict": "pending_v5" + }, + "C1_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0526, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8571, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1429, + "entropy": 3.0565, + "n": 7, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.5714, + "truth_eliminated": 0, + "width_median": 33.0 + }, + "verdict": "pending_v5" + }, + "C2_both": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.3139, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.3175, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "verdict": "pending_v5" + }, + "C2_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "benefit", + "raw": "benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.3103, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.317, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 1.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "verdict": "pending_v5" + }, + "C2_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 1.0, + "entropy": 3.3137, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 1.0, + "entropy": 3.318, + "n": 1, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 31.0 + }, + "verdict": "pending_v5" + } + }, + "askable_per_case": { + "cases_with_at_least_one": 7, + "cases_with_at_least_two": 1, + "distribution": { + "0": 13, + "1": 6, + "3": 1 + }, + "mean": 0.45 + }, + "n": 20 + }, + "60": { + "arms": { + "B0_true_precision": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9174, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 53.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.05, + "entropy": 3.9537, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 56.0 + } + }, + "C0_spoken": { + "percent": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 3.9174, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.9, + "truth_eliminated": 0, + "width_median": 55.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.15, + "entropy": 3.9471, + "n": 20, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.45, + "truth_eliminated": 0, + "width_median": 61.0 + } + }, + "C1_both": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.5108, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.7, + "truth_eliminated": 0, + "width_median": 53.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.5319, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3, + "truth_eliminated": 0, + "width_median": 58.0 + }, + "verdict": "pending_v5" + }, + "C1_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.6783, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 44.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.7156, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 58.0 + }, + "verdict": "pending_v5" + }, + "C1_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.5111, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.8, + "truth_eliminated": 0, + "width_median": 53.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.1, + "entropy": 3.5316, + "n": 10, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.4, + "truth_eliminated": 0, + "width_median": 63.0 + }, + "verdict": "pending_v5" + }, + "C2_both": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.9205, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3333, + "truth_eliminated": 0, + "width_median": 83.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 3.9302, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3333, + "truth_eliminated": 0, + "width_median": 91.0 + }, + "verdict": "pending_v5" + }, + "C2_prior": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 4.0118, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6667, + "truth_eliminated": 0, + "width_median": 61.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.0, + "entropy": 4.036, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.6667, + "truth_eliminated": 0, + "width_median": 91.0 + }, + "verdict": "pending_v5" + }, + "C2_probe": { + "debug_verdict_v4_vs_C0": { + "percent": "no_benefit", + "raw": "no_benefit" + }, + "percent": { + "coverage": 1.0, + "engine_top1": 0.3333, + "entropy": 3.9219, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.3333, + "truth_eliminated": 0, + "width_median": 83.0 + }, + "raw": { + "coverage": 1.0, + "engine_top1": 0.3333, + "entropy": 3.9294, + "n": 3, + "refresh_mean": 0.0, + "squeezed": 0, + "tie": 0.0, + "top1": 0.0, + "truth_eliminated": 0, + "width_median": 91.0 + }, + "verdict": "pending_v5" + } + }, + "askable_per_case": { + "cases_with_at_least_one": 10, + "cases_with_at_least_two": 3, + "distribution": { + "0": 10, + "1": 7, + "2": 2, + "3": 1 + }, + "mean": 0.7 + }, + "n": 20 + } + }, + "reconciliation": { + "cases": 60, + "rows": [ + { + "candidates": 11, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "barack_obama_1961_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "angelina_jolie_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tiger_woods_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_taylor_1932_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "elizabeth_montgomery_1933_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "pablo_picasso_1881_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sigmund_freud_1856_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "salma_hayek_1966_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "steve_reich_1936_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "albert_brooks_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "paul_ryan_1970_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "tom_kennedy_1927_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "andrew_windsor_1960_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "sean_lennon_1975_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "joseph_kennedy_iii_1980_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "kurt_cobain_1967_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "john_robbins_1947_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "chad_everett_1937_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "bernd_eichinger_1949_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + }, + { + "candidates": 11, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 10 + }, + { + "candidates": 31, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 30 + }, + { + "candidates": 61, + "case_id": "iwao_takamoto_1925_aa_v4_holdout", + "changed": 0, + "details": [], + "radius": 60 + } + ], + "unreconciled": 0 + }, + "seed": 20260929, + "verdict": "pending_v5" +} diff --git a/docs/tasks/PROGRESS-rectification-scoring-research-20260929.md b/docs/tasks/PROGRESS-rectification-scoring-research-20260929.md new file mode 100644 index 00000000..531cf095 --- /dev/null +++ b/docs/tasks/PROGRESS-rectification-scoring-research-20260929.md @@ -0,0 +1,66 @@ +# PROGRESS · 生时校正打分方法研究:似然比权重 / 缺席证据 / 精度追问(脚手架阶段,2026-09-29) + +- 任务书:`docs/tasks/TASK-rectification-scoring-research-20260929.md` +- 执行:Claude(直接执行模式,fork 子代理),分支 `codex/rectification-scoring-research-20260929`,基线 `origin/staging` @ `6e393780`(任务书写作时 `5febb111`,其间只有文档提交)。**未推送。** +- 研究报告:`docs/research/rectification_scoring_research_2026_09_29.md` + `scoring_research_{lr,absence,precision_followup}_2026_09_29.json` +- 本轮范围(按调度方指令):在 v4 上完成 R-0 脚手架并把 R-A / R-B / R-C 流水线跑通;**所有判定字段写 `pending_v5`**,v4 数字只作调试参考。v5 交付后在 v5 上跑判定(复跑命令见研究报告末尾)。 + +## 开工 + +- `git fetch origin --prune` → worktree `.worktrees/rectification-scoring-research-20260929` 基于 `origin/staging`;任何 git 操作前 `git status -sb` 核对分支。 +- 先读:定论页(含 §7 勘误)、`rectification_offline_research_2026_09_26.md`、`rectification_typed_event_scoring_2026_09_29.md`、ERR-110 / ERR-111。 +- 环境:本机没有 `.venv`(所有 worktree 的 `.venv` 软链都指向不存在的 `/workspace/Jyotisha/.venv`),全部用系统 `python3` 3.13.5(swisseph 9.1.1、pytest、timezonefinder 齐)。快速门的前端步需要 Node 22:`frontend/node_modules` 软链到主检出,`npm` 用 `/exec-daemon/node` + `/exec-daemon/lib/node_modules/npm` 的 shim(系统 `node` 是 20.19,`/usr/bin/npm` 跑的 `tsx` 不在 PATH)。`PYTHONHASHSEED=0`(ERR-111)。 +- 另一子代理并行做 v5 扩集(`.worktrees/rectification-holdout-expansion-20260929`),未触碰。 + +## 交付物 + +| 文件 | 内容 | +| --- | --- | +| `scripts/research/scoring_research_lib.py` | R-0:特征记录 provider(线上同源 evidence + 额外观察)、逐分对账、特征矩阵、似然比估计(α = 1)、案例级 bootstrap、加权 provider、leave-one-case-out、缩放答案、Platt 温度、缺席年 / 精度升降 / 月份平移辅助 | +| `scripts/research/scoring_research.py` | R-A / R-B / R-C 流水线;`--limit / --radii / --parts / --fast / --cache-dir / --holdout / --dataset-label / --out-suffix` | +| `tests/test_scoring_research.py` | 10 条:纯函数 9 条 + v4 公开案例 ±10 对账 0 差 1 条 | +| `docs/research/rectification_scoring_research_2026_09_29.md` | 报告(结构按任务书 R-D;结论栏全部 `pending_v5`;复跑命令) | +| `docs/research/scoring_research_*_2026_09_29.json` | 三份结果(`dataset: v4, verdict: pending_v5`) | +| `docs/BUG_HISTORY.md` | BUG-1091 `investigating`(开工核对最大号 1089;1090 由扩集单占用,未冲突) | +| `docs/research/ACTIVE_FRONTS.md`、定论页 §9、`docs/tasks/README.md` | 索引与状态板 | + +生产代码、`sealed_holdout_rerun.py::PRODUCTION_FILES` 的 10 个冻结文件、任何常数:**零改动**(`git status` 只有上表文件)。 + +## R-0 · 对账(任务书硬红线 4) + +- 研究计分器 vs 线上 `score_candidates`(经 `build_event_contribution_matrix` 同一条矩阵链):**20 例 × 3 档 = 60 个 case-radius,分数差 0、规则 id 差 0**(`reconciliation.unreconciled = 0`,三份 JSON 均记录)。 +- 做法:provider 用 `precision_gate_lib.score_event_with_policy`(V0)+ 线上 `_controlled_transit_rules`(带缓存)/ `_ashtakavarga_auxiliary` / `_shadbala_verified_components_auxiliary`(取 context 里线上算好的结果),返回与 `_candidate_row` 相同的 evidence;额外观察只进 `FeatureStore`,不进 rule_ids / points,所以矩阵、`event_kind` 系数、`technique_layers`、出题都与线上一致。 +- 特征 42 个:线上规则 25 + 额外观察 17(KP 宫头子主 / 星主 × MD/AD/PD、KP 上升子主、D60 × MD/AD/PD、Pranapada / Hora / Ghati / Bhava lagna 落领域宫)。特征值 = 该规则在该事件采样日期上命中的比例。真值标签 = 候选属于真值所在签名簇。 +- 留一法按**案例**留(`loo_folds`),`scale` 与 `temperature` 也只从训练折估。 + +## R-A / R-B / R-C · 流水线状态 + +| 项 | 状态 | v4 调试参考(`raw` 口径,留一法) | 判定 | +| --- | --- | --- | --- | +| R-A A1 / A2 / A3 | 跑通(留一法 + 全集、六题回放、引擎头名、校准曲线、稳健性四项、LR 表 + 案例级 bootstrap 区间) | 头名 ±10 0.70 / 0.75 / 0.80(基线 0.80);±30 0.55 / 0.60 / 0.65(0.55);±60 0.45 / 0.55 / 0.50(0.40)。真值在区间除 ±60 A3 0.95 外全 1.00。`percent` 口径 softmax 后验过自信,挤出 3–5 例 | `pending_v5` | +| R-B 缺席 @0.5 / 1.0 / 2.0 分 | 跑通(含漏说 1–2 件 × 5 次) | 2.0 分(=卡片)把覆盖压到 0.85 / 0.70 / 0.40;0.5 / 1.0 分覆盖 1.00 但头名一律降 | `pending_v5`(v4 信号为负) | +| R-C C1 / C2 × prior / probe / both | 跑通 | 可追问件数每例均值 0.20 / 0.45 / 0.70;±30 上 C1_prior 7 例头名 0.71 | `pending_v5`(样本太少) | + +各特征在 v4 上的 LR 与「LR≈1」名单在报告 R-A 第 3 条与 JSON `radii.*.lr_table` / `noise_features`:**42 个特征里 31 / 36 / 38 个(±10 / ±30 / ±60)的 5–95% 区间跨 0**。 + +## 复跑与确定性(任务书硬红线 5) + +- `PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --cache-dir ` 连跑两次(758 s / 748 s,静态 context 走磁盘缓存),三份 JSON `cmp` **逐字节一致**;`--out-suffix rerun` 的副本比对后删除,不入库。 +- JSON 不写耗时字段,`sort_keys=True`;随机(漏说抽样、月份平移、答错抽样、bootstrap)全部按 `f"{SEED}:{case_id}:{radius}:…"` 种子。 + +## 快速门 + +- 开工基线(`origin/staging` @ `6e393780`,未改任何文件)与改后各跑一次 `python3 scripts/run_quality_gate.py --profile quick`(PATH 带 Node 22 shim): + - Python:**1001 passed, 1 skipped** 两次相同(改后含新增的 10 条 `test_scoring_research.py`——门禁的 pytest 集合按清单收集,新文件不在清单里,计数没变;定向 `python3 -m pytest tests/test_scoring_research.py -q` 10 passed)。 + - 前端 `npm test`:基线 4357 / pass 4301 / fail 25;改后 4357 / pass **4302** / fail **24**、cancelled 0。失败清单逐条比对:改后 24 条是基线 25 条的子集(少的一条 `D3 · the first stage line shows on send…` 是基线那次偶发红、改后绿,前端零改动);24 条全是需要 Docker / PostgreSQL / 部署环境的套件(Better Auth / billing / migration / staging sync 等),与 09-29 上一单验收记录的 24 条环境失败一致。 + - 门禁整体 exit 1 = 该 24 条环境失败所致,基线同样 exit 1;**无新增失败**。 +- 隐私:`tests/test_repo_privacy_markers.py` 在快速门 Python 步内通过(1001 passed 含之)。 + +## 缺口与说明 + +- **判定未做**:按任务书,R-A / R-B / R-C 的有无收益只能在 v5 上判;本轮 JSON 的 `debug_verdict_v4` 只是 `gate_verdict` 的机械输出,不作结论。 +- R-A 的 `percent` 口径需要先解决后验标定(单温度 Platt 不够,校准曲线高桶过自信);v5 上若仍过自信,考虑按训练折做等分位标定或直接用 `raw` 口径(缩放到线上极差)。 +- 六道题由线上矩阵出一次、各方案共用(隔离先验的影响);若某方案在 v5 上过门,实现前还应按方案矩阵重新出题再测一次。 +- 大运边界邻近项(`merge_transition_proximity`)按线上原样附加、未重加权。 +- R-B 的缺席年在 v4 上不可靠(传记摘录不是完整时间线);这条限制在真人口述上同样存在,报告已写明。 +- 未做:R-C 的产品层实现、任何实现单;`CHANGELOG.md` 无用户可见变化,未改。 diff --git a/docs/tasks/README.md b/docs/tasks/README.md index ec9247de..0c70bf01 100644 --- a/docs/tasks/README.md +++ b/docs/tasks/README.md @@ -256,7 +256,7 @@ | `TASK-mobile-chart-and-confirmed-edit-20260929.md` | `PROGRESS-mobile-chart-confirmed-edit-20260929.md` | **手机星盘 + 确认时间可改 + 天空定格**:iPhone 星盘被宽表撑出屏幕(BUG-1083);confirmed 状态改声明字段也重置排盘时间(推翻 BUG-264 约定);「那一刻的天空」改为重放汇聚→定格→右上角分享;去掉报告星图封面 | 已验收(Claude 2026-09-29),已推 staging | BUG-1083 | | `TASK-rectification-futile-collect-stop-20260929.md` | `PROGRESS-rectification-futile-collect-stop-20260929.md` | **生时校正停掉无效补经历循环(止血)**:打字经历不收窄(BUG-560 后果),流程却一路索要,真机 22 件整窗不动;采集只为开闸、门开后只问点选卡问完即出卡;交付正文去吻合率;多段时旁白误报「范围没变」;交付轮带采集题 | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1084~1087 | | `TASK-rectification-holdout-expansion-20260929.md` | — | **开放评价集扩到 ≥60 例(v5)**:所有打分判定都在 20 例上做,1 例 = 5pp,KP / 精度闸 / V1n 的 no_benefit 都只差 1–3 例。只用 Rodden AA 公开名人,事件人工核对出处,分层(年代 / 纬度 / 南半球 / 跨午夜 / UTC+8),v4 不动;基线成绩单 v4 子集须与已发布数字逐格一致。研究单的判定以 v5 为准 | 待领取 | — | -| `TASK-rectification-scoring-research-20260929.md` | — | **打分方法研究(离线)**:R-A 似然比校准权重(leave-one-case-out,KP / Pranapada / D60 作为特征由数据定权重、允许负权重)、R-B 缺席证据(口述时间线覆盖时段内的空白年扣分,必测漏说稳健性)、R-C 精度追问可行性(对已说事件追问月份,算在定向追问 2 条额度内)。判定必须在 v5 上做;有收益才立实现单并按 ERR-110 重冻结 | 待领取(脚手架可先在 v4 上开发) | — | +| `TASK-rectification-scoring-research-20260929.md` | `PROGRESS-rectification-scoring-research-20260929.md` | **打分方法研究(离线)**:R-A 似然比校准权重(leave-one-case-out,KP / Pranapada / D60 作为特征由数据定权重、允许负权重)、R-B 缺席证据(口述时间线覆盖时段内的空白年扣分,必测漏说稳健性)、R-C 精度追问可行性(对已说事件追问月份,算在定向追问 2 条额度内)。判定必须在 v5 上做;有收益才立实现单并按 ERR-110 重冻结 | 已实现待验收(Claude fork 2026-09-29:R-0 脚手架 + R-A / R-B / R-C 流水线在 v4 上跑通,对账 0 差,判定全部 `pending_v5`;生产代码零改动) | 分支 `codex/rectification-scoring-research-20260929`(BUG-1091 investigating) | | `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1088、1089 | | `TASK-report-reader-polish-20260929.md` | `PROGRESS-report-reader-polish-20260929.md` | **报告页打磨**:生成入口挪进页面主体(删标题栏按钮)、详情页到底部按钮(懒渲染一次到底)、导出按钮带文字、目录一级/二级分层、去掉「字段」与 RL/NL 缩写表头、状态列同义重复去重、报告表格淡底色、分块导出显示真实文件大小 | Claude 直接执行并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +7 B、快速门 Python 1000 passed),已推 staging,真机欠 | BUG-1092~1094 | | `TASK-report-english-edition-20260929.md` | — | **报告中英两版**:同一次引擎计算渲染 zh/en 两遍(不用模型翻译),瑜伽库 406 条补英文、模板与前端表头英文化;中文逐字节不变、英文零汉字、两版数字序列一致;阅读页 `?lang=en` 切换,导出当前语言 + 「问 AI 建议导出英文版」提示;旧报告不补英文 | 待领取(依赖 reader-polish 合入) | — | diff --git a/scripts/research/scoring_research.py b/scripts/research/scoring_research.py new file mode 100644 index 00000000..bf2487c0 --- /dev/null +++ b/scripts/research/scoring_research.py @@ -0,0 +1,529 @@ +#!/usr/bin/env python3 +"""Offline research runner for TASK-rectification-scoring-research-20260929. + +R-0 research scorer reconciled per minute against the production scorer, then + per-(candidate, event) rule features recorded (eleven production rules, + auxiliary rules, and unscored extras: KP cusp sub/star lords, Pranapada / + Hora / Ghati / Bhava lagna, D60). +R-A likelihood-ratio weights (Laplace α = 1) per radius, leave-one-CASE-out + and full-fit; variants A1 (production features, reward only), A2 (+extras, + reward only), A3 (+extras, naive-Bayes present/absent, negative allowed); + six-probe replay, engine top1, calibration curve, robustness (±7-day and + ±1–3-month date shifts, 1–2 flipped answers). +R-B absence evidence: gap years inside a domain's spoken span answered "no" + at 0.5 / 1.0 / 2.0 points (a card is 2.0), plus "user omitted 1–2 events" + robustness. +R-C precision follow-up: with all day events degraded to year ("spoken" + form), how many events would split candidates if the month were known, + and the gain from asking 1 / 2 of them (as a prior recompute and as an + answered probe). + +Every verdict field is written as ``pending_v5``: the task brief allows a +verdict only on the v5 open set. v4 numbers are debugging references. + +Run: + PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py --limit 2 --radii 10 --no-write + PYTHONHASHSEED=0 python3 scripts/research/scoring_research.py +""" + +from __future__ import annotations + +import argparse +import json +import pickle +import sys +import time as clock_module +import traceback +from collections import defaultdict +from pathlib import Path +from random import Random +from typing import Any, Sequence + +ROOT = Path(__file__).resolve().parents[2] +if str(ROOT) not in sys.path: + sys.path.insert(0, str(ROOT)) + +from scripts.active_rectification_event_engine import AYANAMSA, NODE_MODE, compute_candidate_static_contexts # noqa: E402 +from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix, scoreable_request # noqa: E402 +from scripts.research.minute_resolution_sweep import MINUTE_STEP, scoring_request_for # noqa: E402 +from scripts.research.offline_research_20260926_lib import exhaustive_flip_sets, flipped_answers, sample_flip_sets # noqa: E402 +from scripts.research.precision_gate_lib import gate_verdict, jitter_day_events # noqa: E402 +from scripts.research.probe_supply_after_six import ASK_COUNT, optimal_answer # noqa: E402 +import scripts.research.scoring_research_lib as lib # noqa: E402 + +HOLDOUT_V4 = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v4.json" +OUT_DIR = ROOT / "docs" / "research" +RADII = (10, 30, 60) +PRIORS = ("raw", "percent") +ABSENCE_POINTS = (0.5, 1.0, 2.0) # score points per "no"; a card is 2.0 +CARD_POINTS = 2.0 +OMIT_REPEATS = 5 +SHIFT_REPEATS = 3 +FLIP_REPEATS = 10 +SEED = 20260929 +ABSENCE_SOURCE = "absence_research" +FOLLOWUP_SOURCE = "precision_followup_research" + + +# --------------------------------------------------------------------------- +# per-case data +# --------------------------------------------------------------------------- + +class CaseData: + def __init__(self, case: dict[str, Any], radius: int, cache_dir: Path | None) -> None: + self.case = case + self.case_id = str(case["case_id"]) + self.radius = radius + self.true_time = str(case["birth"]["time"])[:5] + self.request = scoring_request_for({**case, "candidate_radius_minutes": radius}, radius) + self.contexts = self._contexts(cache_dir) + self.times = [ctx["candidate_at"].strftime("%H:%M") for ctx in self.contexts] + self.truth_times = lib.truth_cluster_times(self.contexts, self.true_time) + self.d60 = lib.d60_charts(self.contexts) + # production + prod_built = build_event_contribution_matrix(self.request, static_contexts=self.contexts) + self.prod_rows = score_from_matrix(self.request, prod_built) + self.prod_built = prod_built + self.probes = lib.production_probes(self.request, prod_built, self.times, self.true_time) + self.prod_public = lib.public_for(self.prod_rows, self.contexts) + # research recording pass (must reconcile to production) + self.store = lib.FeatureStore() + built = build_event_contribution_matrix( + self.request, row_provider=lib.make_recording_provider(self.contexts, self.store, self.d60), + static_contexts=self.contexts, + ) + rows = score_from_matrix(self.request, built) + self.reconciliation = lib.reconcile_rows(self.prod_rows, rows) + self.training_ids = lib.training_event_ids(self.request) + self.representatives = lib.representatives_for(self.contexts) + self.events_by_id = {str(e["id"]): e for e in self.request["events"]} + + def _contexts(self, cache_dir: Path | None) -> list[dict[str, Any]]: + if cache_dir is not None: + path = cache_dir / f"{self.case_id}_{self.radius}.pkl" + if path.exists(): + with path.open("rb") as handle: + return pickle.load(handle) + contexts = compute_candidate_static_contexts(self.request) + if cache_dir is not None: + cache_dir.mkdir(parents=True, exist_ok=True) + with (cache_dir / f"{self.case_id}_{self.radius}.pkl").open("wb") as handle: + pickle.dump(contexts, handle) + return contexts + + # features ------------------------------------------------------------- + def examples(self, *, include_extras: bool) -> list[tuple[dict[str, float], bool]]: + out: list[tuple[dict[str, float], bool]] = [] + truth = set(self.truth_times) + for event_id in self.training_ids: + matrix = lib.event_feature_matrix(self.store, event_id, self.times, include_extras=include_extras) + for time in self.times: + out.append((matrix.get(time, {}), time in truth)) + return out + + # scoring with a weight table ------------------------------------------ + def weighted_rows( + self, table: lib.LRTable, variant: str, scale: float, + store: lib.FeatureStore | None = None, request: dict[str, Any] | None = None, + ) -> tuple[list[dict[str, Any]], dict[str, Any]]: + request = request or self.request + built = build_event_contribution_matrix( + request, + row_provider=lib.make_weighted_provider(store or self.store, table, variant=variant, scale=scale), + static_contexts=self.contexts, + ) + return score_from_matrix(request, built), built + + def rescored_store(self, events: Sequence[dict[str, Any]]) -> tuple[lib.FeatureStore, dict[str, Any], list[dict[str, Any]]]: + """Re-record features for an altered event list (date shifts): (store, request, production rows).""" + request = scoring_request_for({**self.case, "events": list(events), "candidate_radius_minutes": self.radius}, self.radius) + store = lib.FeatureStore() + built = build_event_contribution_matrix( + request, row_provider=lib.make_recording_provider(self.contexts, store, self.d60), static_contexts=self.contexts, + ) + return store, request, score_from_matrix(request, built) + + def production_rows_for(self, events: Sequence[dict[str, Any]]) -> list[dict[str, Any]]: + request = scoring_request_for({**self.case, "events": list(events), "candidate_radius_minutes": self.radius}, self.radius) + built = build_event_contribution_matrix(request, static_contexts=self.contexts) + return score_from_matrix(request, built) + + +def evaluate(data: CaseData, rows: Sequence[dict[str, Any]], *, softmax_prior: bool, pre_answers=(), answers=None, temperature: float = 1.0) -> dict[str, dict[str, Any]]: + public = lib.public_for(rows, data.contexts) + out: dict[str, dict[str, Any]] = {} + for mode in PRIORS: + prior = lib.prior_for(public, mode=mode, softmax=softmax_prior, temperature=temperature) + result = lib.replay( + public=public, prior=prior, probes=data.probes, true_time=data.true_time, + window_times=data.times, pre_answers=pre_answers, answers=answers, + ) + result["engine_top1"] = lib.engine_top1(public, data.true_time) + out[mode] = result + return out + + +# --------------------------------------------------------------------------- +# R-A +# --------------------------------------------------------------------------- + +def fit_prior(training: Sequence[CaseData], table: lib.LRTable, variant: str) -> tuple[float, float]: + """(scale, temperature) from training cases only. + + scale: research within-case score range matched to the production median + range (so the raw-prior replay's ±2 / 8-lead constants mean the same); + temperature: single Platt temperature for the percent (softmax) prior. + """ + prod = [lib.within_case_range(d.prod_rows) for d in training] + rows_by_case = {d.case_id: d.weighted_rows(table, variant, 1.0)[0] for d in training} + research = [lib.within_case_range(rows) for rows in rows_by_case.values()] + p_med, r_med = lib.median(prod), lib.median(research) + scale = round(p_med / r_med, 6) if p_med and r_med else 1.0 + publics = [ + (lib.public_for([{**r, "score": float(r["score"]) * scale} for r in rows_by_case[d.case_id]], d.contexts), d.truth_times) + for d in training + ] + return scale, lib.fit_temperature(publics) + + +def run_ra(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]: + out: dict[str, Any] = {"variants": list(lib.VARIANTS), "alpha": lib.LAPLACE_ALPHA, "radii": {}} + for radius, cases in sorted(by_radius.items()): + block: dict[str, Any] = {"n": len(cases), "baseline": {}, "full_fit": {}, "loo": {}, "robustness": {}, "calibration": {}, "lr_table": {}, "noise_features": {}} + baseline_rows = [] + for data in cases: + row = evaluate(data, data.prod_rows, softmax_prior=False) + baseline_rows.append({"case_id": data.case_id, **{f"{m}": row[m] for m in PRIORS}}) + block["baseline"] = {m: lib.summarize([r[m] for r in baseline_rows]) for m in PRIORS} + per_case_examples = { + include: {d.case_id: d.examples(include_extras=include) for d in cases} for include in (False, True) + } + altered_stores: dict[tuple[str, str, int], tuple[lib.FeatureStore, dict[str, Any]]] = {} + for variant in lib.VARIANTS: + include_extras, _neg = lib.VARIANT_SPEC[variant] + examples = per_case_examples[include_extras] + # full fit ------------------------------------------------------- + full_table = lib.estimate_lr([ex for rows in examples.values() for ex in rows]) + full_scale, full_temperature = fit_prior(cases, full_table, variant) + full_rows = [] + for data in cases: + rows, _ = data.weighted_rows(full_table, variant, full_scale) + full_rows.append(evaluate(data, rows, softmax_prior=True, temperature=full_temperature)) + block["full_fit"][variant] = { + "scale": full_scale, + "temperature": full_temperature, + **{m: lib.summarize([r[m] for r in full_rows]) for m in PRIORS}, + } + if variant == "A3": + block["lr_table"] = { + name: { + "log_lr_present": full_table.present[name], + "log_lr_absent": full_table.absent[name], + **full_table.support[name], + "extra": lib.is_extra_feature(name), + } + for name in sorted(full_table.present) + } + ci = lib.bootstrap_lr(examples, resamples=50 if fast else 200) + for name, interval in ci.items(): + block["lr_table"].setdefault(name, {}).update(interval) + block["noise_features"] = sorted( + name for name, row in block["lr_table"].items() + if row.get("p05") is not None and row["p05"] <= 0.0 <= row["p95"] + ) + # leave-one-case-out ------------------------------------------- + loo_rows = [] + robustness = defaultdict(list) + calibration_points: list[tuple[float, bool]] = [] + temperatures: list[float] = [] + for held in cases: + training = [d for d in cases if d.case_id != held.case_id] + table = lib.estimate_lr([ex for d in training for ex in examples[d.case_id]]) + scale, temperature = fit_prior(training, table, variant) + temperatures.append(temperature) + rows, _ = held.weighted_rows(table, variant, scale) + result = evaluate(held, rows, softmax_prior=True, temperature=temperature) + loo_rows.append(result) + if variant == "A3": + public = lib.public_for(rows, held.contexts) + posterior = lib.percent_softmax({str(r["time"])[:5]: float(r["score"]) for r in public}, temperature) + truth = set(held.truth_times) + for rep in public: + stamp = str(rep["time"])[:5] + members = {str(t)[:5] for t in (rep.get("cluster_times") or [stamp])} + calibration_points.append((posterior.get(stamp, 0.0) / 100.0, bool(members & truth))) + # robustness: date shifts (features re-recorded with fold weights) + if not fast or variant == "A3": + events = list(held.case.get("events") or []) + shifted_sets = { + "jitter_7d": [jitter_day_events(events, 7, case_id=held.case_id)], + "shift_1_3_months": [ + lib.shift_day_events_by_months(events, seed=f"{SEED}:{held.case_id}:{radius}:{k}") + for k in range(1 if fast else SHIFT_REPEATS) + ], + } + for label, variants_of_events in shifted_sets.items(): + for index, altered in enumerate(variants_of_events): + cache_key = (held.case_id, label, index) + if cache_key not in altered_stores: + altered_stores[cache_key] = held.rescored_store(altered)[:2] + store, request_alt = altered_stores[cache_key] + rows_alt, _ = held.weighted_rows(table, variant, scale, store=store, request=request_alt) + robustness[label].append(evaluate(held, rows_alt, softmax_prior=True, temperature=temperature)) + # answer flips on the LOO rows + asked = list(held.probes)[:ASK_COUNT] + optimal = [optimal_answer(p, held.true_time) for p in asked] + for k in (1, 2): + sets = exhaustive_flip_sets(optimal, k) if k == 1 else sample_flip_sets( + optimal, k, 3 if fast else FLIP_REPEATS, case_id=held.case_id, radius=radius, + ) + for picked in sets: + robustness[f"flip_{k}"].append(evaluate( + held, rows, softmax_prior=True, answers=flipped_answers(optimal, picked), temperature=temperature, + )) + block["loo"][variant] = {m: lib.summarize([r[m] for r in loo_rows]) for m in PRIORS} + block["loo"][variant]["temperatures"] = sorted(set(temperatures)) + block["loo"][variant]["debug_verdict_v4"] = { + m: gate_verdict(block["baseline"][m], block["loo"][variant][m]) for m in PRIORS + } + block["loo"][variant]["verdict"] = "pending_v5" + if robustness: + block["robustness"][variant] = { + label: {m: lib.summarize([r[m] for r in rows_]) for m in PRIORS} + for label, rows_ in robustness.items() + } + if variant == "A3": + block["calibration"] = lib.calibration_bins(calibration_points) + out["radii"][str(radius)] = block + return out + + +# --------------------------------------------------------------------------- +# R-B +# --------------------------------------------------------------------------- + +def absence_probes(data: CaseData, events: Sequence[dict[str, Any]]) -> tuple[list[dict[str, Any]], dict[str, list[int]]]: + gaps = lib.absent_years(events) + probes: list[dict[str, Any]] = [] + for domain, years in gaps.items(): + for year in years: + probe = lib.split_probe( + data.representatives, birth_date=str(data.request["birth_date"]), domain=domain, + year=year, month=None, source=ABSENCE_SOURCE, + ) + if probe is not None: + probes.append(probe) + return probes, gaps + + +def run_rb(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]: + out: dict[str, Any] = {"points": list(ABSENCE_POINTS), "card_points": CARD_POINTS, "omit_repeats": OMIT_REPEATS, "radii": {}} + for radius, cases in sorted(by_radius.items()): + rows_by_arm: dict[str, list[dict[str, Any]]] = defaultdict(list) + gap_stats = [] + omit_rows: dict[str, list[dict[str, Any]]] = defaultdict(list) + for data in cases: + training = [data.events_by_id[i] for i in data.training_ids] + spoken = [{**e, "date": e.get("date_start"), "domain": e["domain"]} for e in training] + probes, gaps = absence_probes(data, spoken) + gap_stats.append({ + "case_id": data.case_id, "absent_years": sum(len(v) for v in gaps.values()), + "domains_with_gaps": len(gaps), "absence_probes": len(probes), + }) + base = evaluate(data, data.prod_rows, softmax_prior=False) + rows_by_arm["B0"].append(base) + for points in ABSENCE_POINTS: + pre = [(probe, "no", points / CARD_POINTS) for probe in probes] + rows_by_arm[f"absence@{points}"].append(evaluate(data, data.prod_rows, softmax_prior=False, pre_answers=pre)) + # omission robustness: user did not mention 1–2 of the training events + for k in (1, 2): + for rep in range(1 if fast else OMIT_REPEATS): + kept = lib.drop_events(spoken, k, seed=f"{SEED}:{data.case_id}:{radius}:{k}:{rep}") + probes_k, _ = absence_probes(data, kept) + for points in ABSENCE_POINTS: + pre = [(probe, "no", points / CARD_POINTS) for probe in probes_k] + omit_rows[f"omit_{k}@{points}"].append(evaluate(data, data.prod_rows, softmax_prior=False, pre_answers=pre)) + block: dict[str, Any] = { + "n": len(cases), + "gap_stats": { + "mean_absent_years": round(sum(g["absent_years"] for g in gap_stats) / len(gap_stats), 2), + "mean_absence_probes": round(sum(g["absence_probes"] for g in gap_stats) / len(gap_stats), 2), + "cases_with_probes": sum(1 for g in gap_stats if g["absence_probes"]), + }, + "arms": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in rows_by_arm.items()}, + "omission": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in omit_rows.items()}, + } + for arm in block["arms"]: + if arm == "B0": + continue + block["arms"][arm]["debug_verdict_v4"] = {m: gate_verdict(block["arms"]["B0"][m], block["arms"][arm][m]) for m in PRIORS} + block["arms"][arm]["verdict"] = "pending_v5" + out["radii"][str(radius)] = block + return out + + +# --------------------------------------------------------------------------- +# R-C +# --------------------------------------------------------------------------- + +def run_rc(by_radius: dict[int, list[CaseData]], *, fast: bool) -> dict[str, Any]: + out: dict[str, Any] = {"radii": {}} + for radius, cases in sorted(by_radius.items()): + rows_by_arm: dict[str, list[dict[str, Any]]] = defaultdict(list) + askable_counts = [] + for data in cases: + originals = {str(e["id"]): e for e in data.case.get("events") or []} + raw_events = list(data.case.get("events") or []) + spoken = lib.degrade_to_year(raw_events) + spoken_rows = data.production_rows_for(spoken) + rows_by_arm["B0_true_precision"].append(evaluate(data, data.prod_rows, softmax_prior=False)) + rows_by_arm["C0_spoken"].append(evaluate(data, spoken_rows, softmax_prior=False)) + # which spoken (year) events would split candidates if the month were known? + training = set(data.training_ids) + id_map = {str(e["id"]): str(rid) for rid, e in zip( + [str(x["id"]) for x in data.request["events"]], data.case.get("events") or [], + )} + candidates = [] + for event in raw_events: + if str(event.get("precision")) != "day": + continue + request_id = id_map.get(str(event["id"])) + if request_id not in training: + continue + year, month = int(str(event["date"])[:4]), int(str(event["date"])[5:7]) + probe = lib.split_probe( + data.representatives, birth_date=str(data.request["birth_date"]), domain=str(event["domain"]), + year=year, month=month, source=FOLLOWUP_SOURCE, + ) + if probe is None: + continue + candidates.append((float(probe.get("information_gain") or 0), str(event["id"]), probe)) + candidates.sort(key=lambda item: (-item[0], item[1])) + askable_counts.append(len(candidates)) + for k in (1, 2): + picked = candidates[:k] + if len(picked) < k: + continue + ids = {item[1] for item in picked} + upgraded = lib.upgrade_to_month(spoken, ids, originals) + rows_by_arm[f"C{k}_prior"].append(evaluate(data, data.production_rows_for(upgraded), softmax_prior=False)) + pre = [(item[2], "yes", 1.0) for item in picked] + rows_by_arm[f"C{k}_probe"].append(evaluate(data, spoken_rows, softmax_prior=False, pre_answers=pre)) + rows_by_arm[f"C{k}_both"].append(evaluate(data, data.production_rows_for(upgraded), softmax_prior=False, pre_answers=pre)) + block: dict[str, Any] = { + "n": len(cases), + "askable_per_case": { + "mean": round(sum(askable_counts) / len(askable_counts), 2) if askable_counts else None, + "distribution": {str(k): askable_counts.count(k) for k in sorted(set(askable_counts))}, + "cases_with_at_least_one": sum(1 for c in askable_counts if c >= 1), + "cases_with_at_least_two": sum(1 for c in askable_counts if c >= 2), + }, + "arms": {arm: {m: lib.summarize([r[m] for r in rows]) for m in PRIORS} for arm, rows in rows_by_arm.items()}, + } + for arm in block["arms"]: + if arm.startswith("C0") or arm.startswith("B0"): + continue + block["arms"][arm]["debug_verdict_v4_vs_C0"] = {m: gate_verdict(block["arms"]["C0_spoken"][m], block["arms"][arm][m]) for m in PRIORS} + block["arms"][arm]["verdict"] = "pending_v5" + out["radii"][str(radius)] = block + return out + + +# --------------------------------------------------------------------------- +# main +# --------------------------------------------------------------------------- + +def load_cases(path: Path) -> tuple[dict[str, Any], list[dict[str, Any]]]: + payload = json.loads(path.read_text(encoding="utf-8")) + return payload, list(payload["cases"]) + + +def write_json(path: Path, payload: dict[str, Any]) -> None: + path.write_text(json.dumps(payload, ensure_ascii=False, indent=1, sort_keys=True) + "\n", encoding="utf-8") + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--holdout", default=str(HOLDOUT_V4)) + parser.add_argument("--dataset-label", default="v4") + parser.add_argument("--limit", type=int, default=0) + parser.add_argument("--radii", nargs="+", type=int, default=list(RADII)) + parser.add_argument("--parts", nargs="+", default=["ra", "rb", "rc"]) + parser.add_argument("--cache-dir", default="") + parser.add_argument("--fast", action="store_true", help="fewer repeats (development)") + parser.add_argument("--no-write", action="store_true") + parser.add_argument("--out-suffix", default="2026_09_29") + args = parser.parse_args() + started = clock_module.time() + holdout_path = Path(args.holdout) + meta, cases = load_cases(holdout_path) + if args.limit: + cases = cases[: args.limit] + cache_dir = Path(args.cache_dir) if args.cache_dir else None + by_radius: dict[int, list[CaseData]] = defaultdict(list) + errors: list[dict[str, Any]] = [] + reconciliation: list[dict[str, Any]] = [] + for case in cases: + for radius in args.radii: + try: + data = CaseData(case, radius, cache_dir) + except Exception as exc: # noqa: BLE001 + errors.append({"case_id": case.get("case_id"), "radius": radius, "error": f"{type(exc).__name__}: {exc}", "trace": traceback.format_exc()}) + continue + reconciliation.append({"case_id": data.case_id, "radius": radius, **data.reconciliation}) + by_radius[radius].append(data) + print(f"loaded {case.get('case_id')} ({clock_module.time() - started:.0f}s)", flush=True) + unreconciled = [row for row in reconciliation if row["changed"]] + header = { + "generated_for": "TASK-rectification-scoring-research-20260929", + "dataset": args.dataset_label, + "holdout": str(holdout_path.relative_to(ROOT)) if holdout_path.is_relative_to(ROOT) else str(holdout_path), + "case_count": len(cases), + "radii": list(args.radii), + "minute_step": MINUTE_STEP, + "ayanamsa": AYANAMSA, + "node_mode": NODE_MODE, + "ask_count": ASK_COUNT, + "probe_today": lib.TODAY.isoformat(), + "probes": "production G0 probes generated once per case/radius from the production matrix; reused for every arm", + "open_set_not_blind": True, + "verdict": "pending_v5", + "reconciliation": {"cases": len(reconciliation), "unreconciled": len(unreconciled), "rows": reconciliation}, + "errors": errors, + "seed": SEED, + "fast": bool(args.fast), + } + if unreconciled: + print(json.dumps({"UNRECONCILED": unreconciled}, ensure_ascii=False, indent=1)) + return 2 + results: dict[str, dict[str, Any]] = {} + if "ra" in args.parts: + results["lr"] = {**header, "part": "R-A", **run_ra(by_radius, fast=args.fast)} + print("R-A done", f"{clock_module.time() - started:.0f}s", flush=True) + if "rb" in args.parts: + results["absence"] = {**header, "part": "R-B", **run_rb(by_radius, fast=args.fast)} + print("R-B done", f"{clock_module.time() - started:.0f}s", flush=True) + if "rc" in args.parts: + results["precision_followup"] = {**header, "part": "R-C", **run_rc(by_radius, fast=args.fast)} + print("R-C done", f"{clock_module.time() - started:.0f}s", flush=True) + for key, payload in results.items(): + payload["elapsed_seconds"] = round(clock_module.time() - started, 1) + if not args.no_write: + write_json(OUT_DIR / f"scoring_research_{key}_{args.out_suffix}.json", {k: v for k, v in payload.items() if k != "elapsed_seconds"}) + brief: dict[str, Any] = {"reconciliation_unreconciled": len(unreconciled), "errors": len(errors)} + for key, payload in results.items(): + brief[key] = {} + for radius, block in payload["radii"].items(): + if key == "lr": + brief[key][radius] = { + "baseline": {m: (block["baseline"][m]["top1"], block["baseline"][m]["coverage"], block["baseline"][m]["width_median"], block["baseline"][m]["engine_top1"]) for m in PRIORS}, + **{v: {m: (block["loo"][v][m]["top1"], block["loo"][v][m]["coverage"], block["loo"][v][m]["width_median"], block["loo"][v][m]["engine_top1"]) for m in PRIORS} for v in lib.VARIANTS}, + } + else: + brief[key][radius] = {arm: {m: (s[m]["top1"], s[m]["coverage"], s[m]["width_median"]) for m in PRIORS} for arm, s in block["arms"].items()} + print(json.dumps(brief, ensure_ascii=False, indent=1)) + return 0 if not errors else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/research/scoring_research_lib.py b/scripts/research/scoring_research_lib.py new file mode 100644 index 00000000..b1d824c5 --- /dev/null +++ b/scripts/research/scoring_research_lib.py @@ -0,0 +1,820 @@ +"""Scaffold (R-0) for TASK-rectification-scoring-research-20260929. + +Offline only. Nothing here changes a production default or a frozen scoring +file. The module records, for every (case, candidate minute, event, sample +date), which scoring rules fire — the eleven production rules from +`active_rectification_event_engine._score_event` plus the auxiliary +transit / Ashtakavarga / Shadbala rules — and a set of *extra* observations +that production computes but does not score (KP cusp sub-lords, Pranapada / +Hora / Ghati / Bhava lagna, D60). Likelihood-ratio weights are estimated per +feature with leave-one-**case**-out folds, and a weighted row provider feeds +the same contribution-matrix / probe / six-question replay machinery that the +09-14 / 09-26 / 09-29 studies used. + +Recording provider vs production: the provider returns exactly the production +evidence (same rule ids, same points), so `build_event_contribution_matrix` +yields the same matrix as the production path. `reconcile_rows` asserts that +per case before any experiment (task hard line 4). +""" + +from __future__ import annotations + +import math +import statistics +import sys +from collections import defaultdict +from dataclasses import dataclass, field +from datetime import date, datetime +from pathlib import Path +from random import Random +from typing import Any, Callable, Sequence + +ROOT = Path(__file__).resolve().parents[2] +if str(ROOT) not in sys.path: + sys.path.insert(0, str(ROOT)) + +from scripts.active_rectification_event_engine import ( # noqa: E402 + DOMAIN_CONFIG, + OBSERVATION_ONLY_LAYERS, + _active_narayana, + _active_vimshottari, + _ashtakavarga_auxiliary, + _controlled_transit_rules, + _event_datetime, + _relative_house, + _shadbala_verified_components_auxiliary, + _varga_chart, + _varga_house, +) +import scripts.rectification.event_probes as event_probes # noqa: E402 +from scripts.rectification.candidate_contrast import cluster_contexts_by_signature # noqa: E402 +from scripts.rectification.case_holdout import holdout_event_ids # noqa: E402 +from scripts.rectification.contracts import is_scoreable_event # noqa: E402 +from scripts.rectification.refinement_packet import window_scan # noqa: E402 +from scripts.rectification.scoring_service import ( # noqa: E402 + build_event_contribution_matrix, + score_from_matrix, +) +from scripts.research.cluster_width_lib import ( # noqa: E402 + SEPARATION_LEAD, + delivery_from_public, + merge_adjacent_traced, + metrics_bundle, + public_from_clusters, + raw_signature_clusters, + shannon_entropy, + still_valid_public, +) +from scripts.research.precision_gate_lib import ( # noqa: E402 + VargaPolicy, + score_event_with_policy, +) +from scripts.research.probe_supply_after_six import ( # noqa: E402 + ASK_COUNT, + SCORE_DELTA, + apply_answer, + inherit_direction, + optimal_answer, + outcome_groups, +) +import varga # noqa: E402 + +TODAY = date(2026, 9, 16) # same probe "today" as the 09-16 / 09-29 replays +LAPLACE_ALPHA = 1.0 +EPS = 1e-9 +EXTRA_PREFIXES = ("kp_", "pranapada_", "hora_lagna_", "ghati_lagna_", "bhava_lagna_") +D60_SUFFIX = "_domain_varga_d60" +IGNORED_RULE_PREFIXES = ("event_kind:", "event_kind_profile:", "no_domain_activation") +CONSTANT_RULES = frozenset({"occupation_auxiliary_not_primary", "appearance_auxiliary_not_primary"}) +VARIANTS = ("A1", "A2", "A3") +VARIANT_SPEC = { + # (uses extra features, allows negative / absence weights) + "A1": (False, False), + "A2": (True, False), + "A3": (True, True), +} + + +# --------------------------------------------------------------------------- +# small helpers +# --------------------------------------------------------------------------- + +def hhmm(value: object) -> str | None: + text = str(value or "")[:5] + return text if len(text) == 5 and text[2] == ":" else None + + +def clock(value: str) -> int: + return int(value[:2]) * 60 + int(value[3:5]) + + +def is_extra_feature(name: str) -> bool: + return name.startswith(EXTRA_PREFIXES) or name.endswith(D60_SUFFIX) + + +def is_scoring_rule(name: str) -> bool: + if name.startswith(IGNORED_RULE_PREFIXES) or name in CONSTANT_RULES: + return False + return not is_extra_feature(name) + + +def median(values: Sequence[float]) -> float | None: + items = [float(item) for item in values if item is not None] + return statistics.median(items) if items else None + + +# --------------------------------------------------------------------------- +# R-0 · feature recording provider (production-identical evidence + extras) +# --------------------------------------------------------------------------- + +@dataclass +class FeatureStore: + """(event_id, sample_date) -> HH:MM -> {rules, extras}.""" + + rows: dict[tuple[str, str], dict[str, dict[str, Any]]] = field(default_factory=dict) + + def put(self, event_id: str, sample: str, time: str, rules: Sequence[str], extras: Sequence[str]) -> None: + self.rows.setdefault((event_id, sample), {})[time] = { + "rules": sorted(set(rules)), + "extras": sorted(set(extras)), + } + + def samples_for(self, event_id: str) -> list[str]: + return sorted(sample for (item, sample) in self.rows if item == event_id) + + +def d60_charts(contexts: Sequence[dict[str, Any]]) -> dict[str, dict[str, Any] | None]: + out: dict[str, dict[str, Any] | None] = {} + for context in contexts: + stamp = context["candidate_at"].strftime("%H:%M") + try: + computed = varga.calc_all_vargas( + context["planet_longitudes"], float(context["ascendant_longitude"]), divisions=[60], + ) + out[stamp] = _varga_chart(computed, "D60") + except (KeyError, TypeError, ValueError): + out[stamp] = None + return out + + +def extra_features( + context: dict[str, Any], + target_houses: tuple[int, ...], + vimshottari: tuple[str, str, str], + d60: dict[str, Any] | None, +) -> list[str]: + feature = context.get("feature") if isinstance(context.get("feature"), dict) else {} + ascendant_index = int(context["ascendant_index"]) + kp_houses = ((feature.get("kp_cusps") or {}).get("houses") or {}) + target_sub = {str((kp_houses.get(str(h)) or {}).get("sub_lord") or "") for h in target_houses} + target_star = {str((kp_houses.get(str(h)) or {}).get("nakshatra_lord") or "") for h in target_houses} + asc_sub = str((kp_houses.get("1") or {}).get("sub_lord") or "") + out: list[str] = [] + for lord, label in zip(vimshottari, ("md", "ad", "pd")): + if lord in target_sub: + out.append(f"kp_target_cusp_sublord_is_vim_{label}") + if lord in target_star: + out.append(f"kp_target_cusp_starlord_is_vim_{label}") + if asc_sub and lord == asc_sub: + out.append(f"kp_asc_sublord_is_vim_{label}") + if isinstance(d60, dict) and _varga_house(d60, lord) in target_houses: + out.append(f"vim_{label}{D60_SUFFIX}") + for key, name in ( + ("pranapada_sign_index", "pranapada"), + ("hora_sign_index", "hora_lagna"), + ("ghati_sign_index", "ghati_lagna"), + ("bhava_sign_index", "bhava_lagna"), + ): + index = feature.get(key) + if isinstance(index, int) and _relative_house(index, ascendant_index) in target_houses: + out.append(f"{name}_in_target_house") + return out + + +def recorded_row( + request: dict[str, Any], + context: dict[str, Any], + *, + store: FeatureStore, + transit_cache: dict[tuple[Any, ...], dict[str, Any]], + d60_by_time: dict[str, dict[str, Any] | None], +) -> dict[str, Any]: + """Production `_candidate_row` for one legacy (single event, one date) request, plus extras.""" + candidate_at = context["candidate_at"] + chart = context["chart"] + planet_longitudes = context["planet_longitudes"] + ascendant_index = context["ascendant_index"] + arudha_padas = context["arudha_padas"] + varga_charts = context["varga_charts"] + moon_longitude = planet_longitudes["Moon"] + stamp = candidate_at.strftime("%H:%M") + policy = VargaPolicy(name="V0", window_minutes=0.0) + evidence: list[dict[str, Any]] = [] + missing: list[str] = [] + for event in request["events"]: + event_at = _event_datetime(event) + prefixes, target_houses = DOMAIN_CONFIG[event["domain"]] + selected = {prefix: varga_charts.get(prefix) for prefix in prefixes} + if any(item is None for item in selected.values()): + missing.extend(prefixes) + continue + try: + vimshottari = _active_vimshottari( + candidate_at.date().isoformat(), moon_longitude, event_at, context.get("vimshottari_timeline"), + ) + except (KeyError, TypeError, ValueError): + missing.append("Vimshottari_MD_AD_PD") + continue + try: + narayana = _active_narayana( + ascendant_index, planet_longitudes, candidate_at, event_at, context.get("narayana_periods"), + ) + except (KeyError, TypeError, ValueError): + missing.append("Narayana_MD_AD") + continue + if narayana[0] is None or narayana[1] is None: + missing.append("Narayana_MD_AD") + continue + row = score_event_with_policy( + candidate_time=stamp, event=event, natal_chart=chart, + varga_by_prefix={k: v for k, v in selected.items() if isinstance(v, dict)}, + vimshottari=vimshottari, narayana=narayana, arudha_padas=arudha_padas, policy=policy, + ) + weight = 1.0 # legacy requests carry precision "day"; the matrix applies the real weight + transit_rules = _controlled_transit_rules( + request, event, ascendant_index, target_houses, transit_cache, + ) + if transit_rules: + row["rule_ids"].extend(transit_rules) + row["points"] = round(row["points"] + 0.25 * len(transit_rules) * weight, 4) + av_rules, av_points = _ashtakavarga_auxiliary( + chart, ascendant_index, target_houses, context.get("ashtakavarga_result"), + ) + if av_rules: + row["rule_ids"].extend(av_rules) + row["points"] = round(row["points"] + av_points * weight, 4) + sb_rules, sb_points = _shadbala_verified_components_auxiliary( + chart, candidate_at.hour + candidate_at.minute / 60, vimshottari, context.get("shadbala_result"), + ) + if sb_rules: + row["rule_ids"].extend(sb_rules) + row["points"] = round(row["points"] + sb_points * weight, 4) + evidence.append(row) + store.put( + str(event["id"]), str(event["date"]), stamp, row["rule_ids"], + extra_features(context, target_houses, vimshottari, d60_by_time.get(stamp)), + ) + return { + "time": stamp, + "score": round(sum(item["points"] for item in evidence), 4), + "evidence": evidence, + "missing_layers": sorted(set( + missing + [layer for layer in context["feature"]["blocked_layers"] if layer not in OBSERVATION_ONLY_LAYERS] + )), + } + + +def make_recording_provider( + contexts: Sequence[dict[str, Any]], + store: FeatureStore, + d60_by_time: dict[str, dict[str, Any] | None], +) -> Callable[[dict[str, Any]], list[dict[str, Any]]]: + transit_cache: dict[tuple[Any, ...], dict[str, Any]] = {} + + def provider(request: dict[str, Any]) -> list[dict[str, Any]]: + return [ + recorded_row(request, context, store=store, transit_cache=transit_cache, d60_by_time=d60_by_time) + for context in contexts + ] + + return provider + + +def score_map(rows: Sequence[dict[str, Any]]) -> dict[str, float]: + return {str(row["time"])[:5]: float(row.get("score") or 0) for row in rows} + + +def reconcile_rows(production: Sequence[dict[str, Any]], research: Sequence[dict[str, Any]]) -> dict[str, Any]: + """Per-minute score and rule-id comparison. `changed == 0` is the R-0 gate.""" + prod = {str(r["time"])[:5]: r for r in production} + rese = {str(r["time"])[:5]: r for r in research} + changed: list[dict[str, Any]] = [] + for time, row in prod.items(): + other = rese.get(time) + if other is None: + changed.append({"time": time, "reason": "missing"}) + continue + if abs(float(row.get("score") or 0) - float(other.get("score") or 0)) > EPS: + changed.append({"time": time, "reason": "score", "production": row.get("score"), "research": other.get("score")}) + continue + prod_rules = { + (item["event_id"], tuple(sorted(item["rule_ids"]))) for item in row.get("evidence") or [] + } + rese_rules = { + (item["event_id"], tuple(sorted(item["rule_ids"]))) for item in other.get("evidence") or [] + } + if prod_rules != rese_rules: + changed.append({"time": time, "reason": "rule_ids"}) + return {"candidates": len(prod), "changed": len(changed), "details": changed[:5]} + + +# --------------------------------------------------------------------------- +# feature matrices and likelihood-ratio weights +# --------------------------------------------------------------------------- + +def event_feature_matrix( + store: FeatureStore, + event_id: str, + times: Sequence[str], + *, + include_extras: bool, +) -> dict[str, dict[str, float]]: + """HH:MM -> feature -> fraction of sample dates on which it fired.""" + samples = store.samples_for(event_id) + out: dict[str, dict[str, float]] = {} + if not samples: + return {time: {} for time in times} + for time in times: + counts: dict[str, int] = defaultdict(int) + for sample in samples: + cell = store.rows.get((event_id, sample), {}).get(time) + if not cell: + continue + names = [name for name in cell["rules"] if is_scoring_rule(name)] + if include_extras: + names += list(cell["extras"]) + for name in names: + counts[name] += 1 + out[time] = {name: round(count / len(samples), 6) for name, count in counts.items()} + return out + + +def training_event_ids(request: dict[str, Any]) -> list[str]: + holdout = holdout_event_ids(request["events"]) + return [str(e["id"]) for e in request["events"] if is_scoreable_event(e) and str(e["id"]) not in holdout] + + +@dataclass +class LRTable: + alpha: float + positives: float + negatives: float + present: dict[str, float] # log LR of the feature firing + absent: dict[str, float] # log LR of the feature not firing (naive Bayes complement) + support: dict[str, dict[str, float]] + + def weight(self, name: str, *, allow_negative: bool) -> tuple[float, float]: + """(present weight, absent weight) under the variant policy.""" + present = self.present.get(name, 0.0) + absent = self.absent.get(name, 0.0) + if allow_negative: + return present, absent + return max(present, 0.0), 0.0 + + +def estimate_lr(examples: Sequence[tuple[dict[str, float], bool]], *, alpha: float = LAPLACE_ALPHA) -> LRTable: + """Per-feature log likelihood ratios from (feature fractions, is-truth-cluster) examples.""" + positives = float(sum(1 for _values, truth in examples if truth)) + negatives = float(sum(1 for _values, truth in examples if not truth)) + names: set[str] = set() + pos_sum: dict[str, float] = defaultdict(float) + neg_sum: dict[str, float] = defaultdict(float) + for values, truth in examples: + for name, value in values.items(): + names.add(name) + if truth: + pos_sum[name] += float(value) + else: + neg_sum[name] += float(value) + present: dict[str, float] = {} + absent: dict[str, float] = {} + support: dict[str, dict[str, float]] = {} + for name in sorted(names): + p_pos = (pos_sum[name] + alpha) / (positives + 2 * alpha) + p_neg = (neg_sum[name] + alpha) / (negatives + 2 * alpha) + present[name] = round(math.log(p_pos / p_neg), 6) + absent[name] = round(math.log((1 - p_pos) / (1 - p_neg)), 6) + support[name] = { + "fires_truth": round(pos_sum[name], 4), + "fires_other": round(neg_sum[name], 4), + "p_truth": round(p_pos, 6), + "p_other": round(p_neg, 6), + } + return LRTable(alpha=alpha, positives=positives, negatives=negatives, present=present, absent=absent, support=support) + + +def bootstrap_lr( + per_case_examples: dict[str, list[tuple[dict[str, float], bool]]], + *, + resamples: int = 200, + seed: int = 20260929, + alpha: float = LAPLACE_ALPHA, +) -> dict[str, dict[str, float]]: + """Case-level bootstrap 5–95 % interval of the present-log-LR per feature.""" + ids = sorted(per_case_examples) + rng = Random(f"{seed}:lr-bootstrap") + draws: dict[str, list[float]] = defaultdict(list) + for _ in range(resamples): + picked = [rng.choice(ids) for _ in ids] + table = estimate_lr([ex for cid in picked for ex in per_case_examples[cid]], alpha=alpha) + for name, value in table.present.items(): + draws[name].append(value) + out: dict[str, dict[str, float]] = {} + for name, values in draws.items(): + ordered = sorted(values) + lo = ordered[int(0.05 * (len(ordered) - 1))] + hi = ordered[int(0.95 * (len(ordered) - 1))] + out[name] = {"p05": round(lo, 4), "p95": round(hi, 4), "draws": len(ordered)} + return out + + +def weighted_points(values: dict[str, float], table: LRTable, *, allow_negative: bool, universe: Sequence[str]) -> float: + total = 0.0 + for name in universe: + present, absent = table.weight(name, allow_negative=allow_negative) + x = float(values.get(name, 0.0)) + total += x * present + (1.0 - x) * absent + return total + + +def make_weighted_provider( + store: FeatureStore, + table: LRTable, + *, + variant: str, + scale: float = 1.0, + baseline_rows_by_sample: dict[tuple[str, str], dict[str, dict[str, Any]]] | None = None, +) -> Callable[[dict[str, Any]], list[dict[str, Any]]]: + """Row provider that scores each (event, sample date) as Σ feature weights. + + Production rule ids are kept on the evidence so the matrix's kind factor + and technique layers behave as in production; only the points change. + """ + include_extras, allow_negative = VARIANT_SPEC[variant] + universe = [ + name for name in sorted(table.present) + if (include_extras or not is_extra_feature(name)) + ] + + def provider(request: dict[str, Any]) -> list[dict[str, Any]]: + rows: list[dict[str, Any]] = [] + for event in request["events"]: + key = (str(event["id"]), str(event["date"])) + cells = store.rows.get(key, {}) + for time in sorted(cells, key=clock): + cell = cells[time] + names = [name for name in cell["rules"] if is_scoring_rule(name)] + if include_extras: + names += list(cell["extras"]) + values = {name: 1.0 for name in names} + points = round(scale * weighted_points(values, table, allow_negative=allow_negative, universe=universe), 4) + rows.append({ + "time": time, + "score": points, + "evidence": [{ + "event_id": event["id"], "domain": event["domain"], "candidate_time": time, + "rule_ids": list(cell["rules"]), "points": points, + }], + "missing_layers": [], + }) + return rows + + return provider + + +def loo_folds(case_ids: Sequence[str]) -> list[tuple[str, list[str]]]: + ids = list(case_ids) + return [(held, [other for other in ids if other != held]) for held in ids] + + +# --------------------------------------------------------------------------- +# truth labels, priors, replay +# --------------------------------------------------------------------------- + +def truth_cluster_times(contexts: Sequence[dict[str, Any]], true_time: str) -> list[str]: + for cluster in raw_signature_clusters(contexts): + times = [str(item)[:5] for item in cluster.get("times") or []] + if true_time[:5] in times: + return sorted(times, key=clock) + return [true_time[:5]] + + +def public_for(rows: Sequence[dict[str, Any]], contexts: Sequence[dict[str, Any]]) -> list[dict[str, Any]]: + raw = raw_signature_clusters(contexts) + by_time = {stamp: row for row in rows if (stamp := hhmm(row.get("time")))} + merged, _trace = merge_adjacent_traced(raw, by_time) + public = public_from_clusters(merged, rows) + for row in public: + row["score"] = float(row.get("score") or 0) + return public + + +def percent_proportional(scores: dict[str, float]) -> dict[str, float]: + total = sum(max(value, 0.0) for value in scores.values()) + if total <= 0: + return {key: 0.0 for key in scores} + return {key: float(round(max(value, 0.0) / total * 100)) for key, value in scores.items()} + + +def softmax_probabilities(scores: dict[str, float], temperature: float = 1.0) -> dict[str, float]: + if not scores: + return {} + temp = float(temperature) if temperature and temperature > 0 else 1.0 + peak = max(scores.values()) + weights = {key: math.exp((value - peak) / temp) for key, value in scores.items()} + total = sum(weights.values()) + return {key: weight / total for key, weight in weights.items()} + + +def percent_softmax(scores: dict[str, float], temperature: float = 1.0) -> dict[str, float]: + return {key: float(round(p * 100)) for key, p in softmax_probabilities(scores, temperature).items()} + + +TEMPERATURE_GRID = (0.25, 0.5, 1.0, 2.0, 4.0, 8.0, 16.0, 32.0, 64.0, 128.0, 256.0) + + +def truth_mass(public: Sequence[dict[str, Any]], probabilities: dict[str, float], truth_times: Sequence[str]) -> float: + truth = {str(t)[:5] for t in truth_times} + mass = 0.0 + for row in public: + stamp = str(row["time"])[:5] + members = {str(t)[:5] for t in (row.get("cluster_times") or [stamp])} + if members & truth: + mass += probabilities.get(stamp, 0.0) + return mass + + +def fit_temperature( + training: Sequence[tuple[Sequence[dict[str, Any]], Sequence[str]]], + grid: Sequence[float] = TEMPERATURE_GRID, +) -> float: + """Platt-style single temperature: maximise Σ log P(truth cluster) on training cases only.""" + best_t, best_ll = 1.0, None + for temp in grid: + ll = 0.0 + for public, truth_times in training: + base = {str(r["time"])[:5]: float(r.get("score") or 0) for r in public} + mass = truth_mass(public, softmax_probabilities(base, temp), truth_times) + ll += math.log(max(mass, 1e-9)) + if best_ll is None or ll > best_ll + 1e-12: + best_t, best_ll = float(temp), ll + return best_t + + +def prior_for(public: Sequence[dict[str, Any]], *, mode: str, softmax: bool, temperature: float = 1.0) -> dict[str, float]: + base = {str(row["time"])[:5]: float(row.get("score") or 0) for row in public} + if mode == "raw": + return base + return percent_softmax(base, temperature) if softmax else percent_proportional(base) + + +def apply_scaled_answer( + scores: dict[str, float], + conflicts: dict[str, int], + eliminated: set[str], + probe: dict[str, Any], + answer: str, + candidate_times: Sequence[str], + factor: float, +) -> tuple[dict[str, float], dict[str, int], set[str]]: + """`apply_answer` at card strength (factor == 1); weaker evidence scales the + delta and never counts a strong conflict (so it can never eliminate).""" + if factor >= 1.0 - EPS: + return apply_answer(scores, conflicts, eliminated, probe, answer, candidate_times) + yes, no = outcome_groups(probe) + next_scores = dict(scores) + for time in candidate_times: + if time in eliminated: + continue + if time in yes: + raw = "support" + elif time in no: + raw = "conflict" + else: + raw = inherit_direction(time, yes, no) + if answer == "no": + raw = {"support": "conflict", "conflict": "support", "neutral": "neutral"}[raw] + next_scores[time] = next_scores.get(time, 0.0) + SCORE_DELTA[raw] * factor + return next_scores, dict(conflicts), set(eliminated) + + +def replay( + *, + public: Sequence[dict[str, Any]], + prior: dict[str, float], + probes: Sequence[dict[str, Any]], + true_time: str, + window_times: Sequence[str], + pre_answers: Sequence[tuple[dict[str, Any], str, float]] = (), + answers: Sequence[str | None] | None = None, +) -> dict[str, Any]: + """Prior → optional pre-answers (probe, answer, factor) → six probes answered from truth.""" + reps = [str(row["time"])[:5] for row in public] + scores = {time: float(prior.get(time, 0.0)) for time in reps} + conflicts = {time: 0 for time in reps} + eliminated: set[str] = set() + for probe, answer, factor in pre_answers: + scores, conflicts, eliminated = apply_scaled_answer(scores, conflicts, eliminated, probe, answer, reps, factor) + asked = list(probes)[:ASK_COUNT] + given = list(answers) if answers is not None else [optimal_answer(p, true_time) for p in asked] + answered = 0 + for probe, answer in zip(asked, given): + if answer is None: + continue + answered += 1 + scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, answer, reps) + posterior = [{**row, "score": scores.get(str(row["time"])[:5], row.get("score") or 0)} for row in public] + valid = still_valid_public(posterior, scores, eliminated, lead=SEPARATION_LEAD) + delivery = delivery_from_public(valid) + alive = [row for row in posterior if str(row["time"])[:5] not in eliminated] + metrics = metrics_bundle( + public=alive, true_time=true_time, window_times=window_times, + delivery_times=delivery["times"], delivery_width=delivery["width"], independent=True, + entropy_scores=[scores.get(str(row["time"])[:5], 0.0) for row in alive], + ) + truth_eliminated = any( + true_time in (row.get("cluster_times") or [str(row.get("time"))[:5]]) and str(row["time"])[:5] in eliminated + for row in public + ) + return { + "top1": bool(metrics["top1_hit"]), + "coverage": bool(metrics["coverage"]), + "width": delivery["width"], + "tie": bool(metrics["tie"]), + "squeezed": bool(metrics["truth_squeezed"]), + "eliminated": len(eliminated), + "truth_eliminated": bool(truth_eliminated), + "questions": answered, + "entropy": round(shannon_entropy(max(s, 0.0) for t, s in scores.items() if t not in eliminated), 4), + } + + +def engine_top1(public: Sequence[dict[str, Any]], true_time: str) -> bool: + from scripts.research.cluster_width_lib import top1_from_public + return bool(top1_from_public(public, true_time)) + + +def within_case_range(rows: Sequence[dict[str, Any]]) -> float: + scores = [float(row.get("score") or 0) for row in rows] + return (max(scores) - min(scores)) if scores else 0.0 + + +def summarize(rows: Sequence[dict[str, Any]]) -> dict[str, Any]: + """Same keys as precision_gate_sweep.summarize so `gate_verdict` can read it.""" + if not rows: + return {"n": 0, "top1": None, "coverage": None, "width_median": None, "tie": None, "squeezed": 0, + "engine_top1": None, "refresh_mean": 0.0, "entropy": None, "truth_eliminated": 0} + n = len(rows) + return { + "n": n, + "top1": round(sum(1 for r in rows if r.get("top1")) / n, 4), + "coverage": round(sum(1 for r in rows if r.get("coverage")) / n, 4), + "width_median": median([r.get("width") for r in rows]), + "tie": round(sum(1 for r in rows if r.get("tie")) / n, 4), + "squeezed": sum(1 for r in rows if r.get("squeezed")), + "engine_top1": round(sum(1 for r in rows if r.get("engine_top1")) / n, 4), + "refresh_mean": 0.0, + "entropy": round(sum(float(r.get("entropy") or 0) for r in rows) / n, 4), + "truth_eliminated": sum(1 for r in rows if r.get("truth_eliminated")), + } + + +# --------------------------------------------------------------------------- +# probes on representatives (typed-event convention), absence, precision follow-up +# --------------------------------------------------------------------------- + +def production_probes(request: dict[str, Any], built: dict[str, Any], times: Sequence[str], true_time: str) -> list[dict[str, Any]]: + payload = {**request, "refresh_probes": False, "asked_probe_keys": []} + return event_probes.discriminating_event_probes( + payload, built, scan=window_scan(built), candidate_times=list(times), + representative_time=true_time, today=TODAY, + ) + + +@dataclass +class Representatives: + reps: list[dict[str, Any]] + clusters: list[dict[str, Any]] + set_version: str + + +def representatives_for(contexts: Sequence[dict[str, Any]]) -> Representatives: + full = [item for item in contexts if event_probes._context_time(item)] + full.sort(key=lambda item: event_probes._clock(str(event_probes._context_time(item)))) + clusters = cluster_contexts_by_signature(full) + reps = [c["representative"] for c in clusters if event_probes._scoreable(c["representative"])] + if len(reps) < 2: + reps = [item for item in full if event_probes._scoreable(item)] + version = event_probes.candidate_set_version([c["times"] for c in clusters]) + return Representatives(reps=reps, clusters=clusters, set_version=version) + + +def split_probe( + representatives: Representatives, + *, + birth_date: str, + domain: str, + year: int, + month: int | None, + source: str, +) -> dict[str, Any] | None: + canonical = event_probes.canonical_domain(domain) + if canonical not in event_probes.DOMAIN_CATALOG: + return None + return event_probes._evaluate_contexts( + representatives.reps, birth_date=birth_date, domain=canonical, year=year, month=month, + source=source, clusters=representatives.clusters, set_version=representatives.set_version, + ) + + +def event_year(event: dict[str, Any]) -> int | None: + raw = str(event.get("date_start") or event.get("date") or "") + return int(raw[:4]) if raw[:4].isdigit() else None + + +def event_month(event: dict[str, Any]) -> int | None: + raw = str(event.get("date_start") or event.get("date") or "") + return int(raw[5:7]) if len(raw) >= 7 and raw[4] == "-" else None + + +def absent_years(events: Sequence[dict[str, Any]]) -> dict[str, list[int]]: + """Per domain: years strictly inside the span of that domain's dated events with no event.""" + years: dict[str, set[int]] = defaultdict(set) + for event in events: + year = event_year(event) + if year is None: + continue + years[str(event.get("domain") or "")].add(year) + out: dict[str, list[int]] = {} + for domain, known in years.items(): + if len(known) < 2: + continue + lo, hi = min(known), max(known) + gaps = [year for year in range(lo + 1, hi) if year not in known] + if gaps: + out[domain] = gaps + return dict(sorted(out.items())) + + +def drop_events(events: Sequence[dict[str, Any]], count: int, *, seed: str) -> list[dict[str, Any]]: + rng = Random(seed) + if count <= 0 or count >= len(events): + return list(events) + dropped = set(rng.sample(range(len(events)), count)) + return [event for index, event in enumerate(events) if index not in dropped] + + +def degrade_to_year(events: Sequence[dict[str, Any]]) -> list[dict[str, Any]]: + """The 'spoken' form of a case: every day/month event becomes year precision.""" + out = [] + for event in events: + item = dict(event) + if str(item.get("precision")) in {"day", "month"}: + item["precision"] = "year" + item["date"] = str(item.get("date") or "")[:4] + out.append(item) + return out + + +def upgrade_to_month(events: Sequence[dict[str, Any]], event_ids: set[str], originals: dict[str, dict[str, Any]]) -> list[dict[str, Any]]: + out = [] + for event in events: + item = dict(event) + original = originals.get(str(item.get("id"))) + if str(item.get("id")) in event_ids and original is not None and len(str(original.get("date") or "")) >= 7: + item["precision"] = "month" + item["date"] = str(original["date"])[:7] + out.append(item) + return out + + +def shift_day_events_by_months(events: Sequence[dict[str, Any]], *, seed: str, choices: Sequence[int] = (-3, -2, -1, 1, 2, 3)) -> list[dict[str, Any]]: + rng = Random(seed) + out = [] + for event in events: + item = dict(event) + raw = str(item.get("date") or "") + if str(item.get("precision")) == "day" and len(raw) >= 10: + year, month, day = int(raw[:4]), int(raw[5:7]), int(raw[8:10]) + index = year * 12 + (month - 1) + rng.choice(list(choices)) + new_year, new_month = divmod(index, 12) + item["date"] = f"{new_year:04d}-{new_month + 1:02d}-{min(day, 28):02d}" + item["shift_months"] = index - (year * 12 + month - 1) + out.append(item) + return out + + +def calibration_bins(points: Sequence[tuple[float, bool]], edges: Sequence[float] = (0.0, 0.05, 0.1, 0.2, 0.3, 0.5, 0.7, 1.0001)) -> list[dict[str, Any]]: + out = [] + for lo, hi in zip(edges, edges[1:]): + inside = [(p, t) for p, t in points if lo <= p < hi] + out.append({ + "bin": f"[{lo:.2f},{min(hi, 1.0):.2f})", + "n": len(inside), + "predicted_mean": round(sum(p for p, _ in inside) / len(inside), 4) if inside else None, + "observed_rate": round(sum(1 for _, t in inside if t) / len(inside), 4) if inside else None, + }) + return out + + +__all__ = [name for name in dir() if not name.startswith("__")] diff --git a/tests/test_scoring_research.py b/tests/test_scoring_research.py new file mode 100644 index 00000000..1e382742 --- /dev/null +++ b/tests/test_scoring_research.py @@ -0,0 +1,159 @@ +"""Regression tests for the 2026-09-29 scoring-method research scaffold. + +Pure helpers use fictional inputs. The one integration test reconciles the +research recording scorer against the production scorer on a public v4 case +(task hard line 4: zero difference before any experiment). +""" +from __future__ import annotations + +import json +from pathlib import Path + +import pytest + +import scripts.research.scoring_research_lib as lib + +ROOT = Path(__file__).resolve().parents[1] +HOLDOUT_V4 = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v4.json" + + +def test_estimate_lr_rewards_discriminating_feature_and_leaves_noise_near_zero() -> None: + examples = [] + for _ in range(20): + examples.append(({"good": 1.0, "noise": 1.0}, True)) + examples.append(({"good": 1.0, "noise": 0.0}, True)) + for _ in range(80): + examples.append(({"noise": 1.0}, False)) + examples.append(({"noise": 0.0}, False)) + table = lib.estimate_lr(examples, alpha=1.0) + assert table.positives == 40 and table.negatives == 160 + assert table.present["good"] > 1.5 + assert table.absent["good"] < -1.5 + assert abs(table.present["noise"]) < 0.05 + # reward-only policy clips the negative absent term; naive-Bayes keeps it + assert table.weight("good", allow_negative=False) == (table.present["good"], 0.0) + assert table.weight("good", allow_negative=True) == (table.present["good"], table.absent["good"]) + + +def test_loo_folds_never_contain_the_held_out_case() -> None: + folds = lib.loo_folds(["a", "b", "c"]) + assert [held for held, _ in folds] == ["a", "b", "c"] + for held, training in folds: + assert held not in training and len(training) == 2 + + +def test_event_feature_matrix_is_the_fraction_of_sample_dates() -> None: + store = lib.FeatureStore() + store.put("e1", "1990-01-15", "10:00", ["vim_md_domain_house", "event_kind:career"], ["pranapada_in_target_house"]) + store.put("e1", "1990-02-15", "10:00", ["vim_ad_domain_lord"], []) + store.put("e1", "1990-01-15", "10:02", ["no_domain_activation"], []) + store.put("e1", "1990-02-15", "10:02", ["no_domain_activation"], []) + matrix = lib.event_feature_matrix(store, "e1", ["10:00", "10:02"], include_extras=True) + assert matrix["10:00"] == {"vim_md_domain_house": 0.5, "vim_ad_domain_lord": 0.5, "pranapada_in_target_house": 0.5} + assert matrix["10:02"] == {} + without = lib.event_feature_matrix(store, "e1", ["10:00"], include_extras=False) + assert "pranapada_in_target_house" not in without["10:00"] + + +def test_scoring_rule_filter_drops_event_kind_and_constant_rules() -> None: + assert lib.is_scoring_rule("vim_md_domain_varga") + assert lib.is_scoring_rule("controlled_transit_jupiter_domain_house") + assert not lib.is_scoring_rule("event_kind:career") + assert not lib.is_scoring_rule("event_kind_profile:career:change") + assert not lib.is_scoring_rule("no_domain_activation") + assert not lib.is_scoring_rule("occupation_auxiliary_not_primary") + assert not lib.is_scoring_rule("kp_asc_sublord_is_vim_md") # extras are a separate namespace + assert lib.is_extra_feature("vim_md_domain_varga_d60") + + +def test_weighted_provider_scores_each_sample_from_the_store() -> None: + store = lib.FeatureStore() + store.put("e1", "1990-01-15", "10:00", ["vim_md_domain_house"], ["kp_asc_sublord_is_vim_md"]) + store.put("e1", "1990-01-15", "10:02", [], []) + table = lib.LRTable( + alpha=1.0, positives=1, negatives=1, + present={"vim_md_domain_house": 1.0, "kp_asc_sublord_is_vim_md": 0.5}, + absent={"vim_md_domain_house": -0.25, "kp_asc_sublord_is_vim_md": -0.1}, + support={}, + ) + request = {"events": [{"id": "e1", "date": "1990-01-15", "domain": "career"}]} + a1 = lib.make_weighted_provider(store, table, variant="A1")(request) + a2 = lib.make_weighted_provider(store, table, variant="A2")(request) + a3 = lib.make_weighted_provider(store, table, variant="A3", scale=2.0)(request) + assert [(r["time"], r["score"]) for r in a1] == [("10:00", 1.0), ("10:02", 0.0)] + assert [(r["time"], r["score"]) for r in a2] == [("10:00", 1.5), ("10:02", 0.0)] + assert [(r["time"], r["score"]) for r in a3] == [("10:00", 3.0), ("10:02", -0.7)] + assert a1[0]["evidence"][0]["rule_ids"] == ["vim_md_domain_house"] + + +def _probe() -> dict: + return {"expected_outcomes": [ + {"answer_class": "yes", "supports": ["10:00"], "conflicts": ["10:04"]}, + {"answer_class": "no", "supports": ["10:04"], "conflicts": ["10:00"]}, + ]} + + +def test_scaled_answer_halves_the_delta_and_never_counts_a_conflict() -> None: + scores = {"10:00": 10.0, "10:02": 10.0, "10:04": 10.0} + conflicts = {time: 0 for time in scores} + weak, weak_conflicts, weak_elim = lib.apply_scaled_answer(scores, conflicts, set(), _probe(), "no", list(scores), 0.5) + assert weak == {"10:00": 9.0, "10:02": 10.0, "10:04": 11.0} + assert weak_conflicts == conflicts and weak_elim == set() + full, full_conflicts, _ = lib.apply_scaled_answer(scores, conflicts, set(), _probe(), "no", list(scores), 1.0) + assert full == {"10:00": 8.0, "10:02": 10.0, "10:04": 12.0} + assert full_conflicts["10:00"] == 1 + + +def test_absent_years_are_gaps_strictly_inside_each_domain_span() -> None: + events = [ + {"domain": "career", "date": "2017-05-01", "precision": "day"}, + {"domain": "career", "date": "2019", "precision": "year"}, + {"domain": "career", "date": "2025-02", "precision": "month"}, + {"domain": "relocation", "date": "2000", "precision": "year"}, + ] + assert lib.absent_years(events) == {"career": [2018, 2020, 2021, 2022, 2023, 2024]} + + +def test_degrade_and_upgrade_precision_helpers() -> None: + events = [{"id": "a", "domain": "career", "date": "2017-05-20", "precision": "day"}, + {"id": "b", "domain": "career", "date": "2019", "precision": "year"}] + spoken = lib.degrade_to_year(events) + assert spoken[0]["precision"] == "year" and spoken[0]["date"] == "2017" + assert spoken[1] == events[1] + upgraded = lib.upgrade_to_month(spoken, {"a"}, {e["id"]: e for e in events}) + assert upgraded[0]["precision"] == "month" and upgraded[0]["date"] == "2017-05" + shifted = lib.shift_day_events_by_months(events, seed="fictional") + assert shifted[0]["shift_months"] in {-3, -2, -1, 1, 2, 3} + assert shifted[1] == events[1] + assert lib.drop_events(events, 1, seed="x") != events and len(lib.drop_events(events, 1, seed="x")) == 1 + + +def test_softmax_and_proportional_percent_sum_to_about_one_hundred() -> None: + scores = {"10:00": 12.0, "10:02": 11.0, "10:04": 9.0} + assert abs(sum(lib.percent_proportional(scores).values()) - 100) <= 2 + soft = lib.percent_softmax(scores) + assert abs(sum(soft.values()) - 100) <= 2 and soft["10:00"] > soft["10:02"] > soft["10:04"] + assert lib.calibration_bins([(0.9, True), (0.02, False)])[-1]["observed_rate"] == 1.0 + + +@pytest.mark.skipif(not HOLDOUT_V4.exists(), reason="v4 open set not present") +def test_recording_scorer_reconciles_with_production_on_a_public_case() -> None: + from scripts.active_rectification_event_engine import compute_candidate_static_contexts + from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix + from scripts.research.minute_resolution_sweep import scoring_request_for + + case = json.loads(HOLDOUT_V4.read_text(encoding="utf-8"))["cases"][0] + request = scoring_request_for({**case, "candidate_radius_minutes": 10}, 10) + contexts = compute_candidate_static_contexts(request) + production = score_from_matrix(request, build_event_contribution_matrix(request, static_contexts=contexts)) + store = lib.FeatureStore() + built = build_event_contribution_matrix( + request, row_provider=lib.make_recording_provider(contexts, store, lib.d60_charts(contexts)), static_contexts=contexts, + ) + research = score_from_matrix(request, built) + outcome = lib.reconcile_rows(production, research) + assert outcome["changed"] == 0, outcome + assert outcome["candidates"] == len(contexts) == 11 + # extras were recorded alongside the production rules + recorded = {name for cells in store.rows.values() for cell in cells.values() for name in cell["extras"]} + assert any(name.startswith("kp_") for name in recorded) or any(name.endswith("_in_target_house") for name in recorded)