docs: publish calibrated public-case evidence
This commit is contained in:
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,33 @@
|
||||
# Public Real-Case Calibration Release Boundary
|
||||
|
||||
Date: 2026-07-12
|
||||
|
||||
This release publishes reproducible public-case calibration safeguards, not a
|
||||
claim of predictive accuracy. It contains no user birth data, private feedback,
|
||||
or private event history.
|
||||
|
||||
## Included Evidence
|
||||
|
||||
- V2.1 scoring correction: `public_real_case_23_case_v21_corrected_observation_2026_07_11.json`
|
||||
- Date-control pilot: `public_real_case_negative_control_pilot_2026_07_11.json`
|
||||
- Annual-control pilot: `public_real_case_annual_control_pilot_2026_07_11.json`
|
||||
- Human-readable correction and control reports in this directory.
|
||||
|
||||
## Current Interpretation Boundary
|
||||
|
||||
- V2.1 removes duplicate MD/AD scoring and labels legacy precision-like fields
|
||||
as deprecated.
|
||||
- SAV/BAV remains descriptive, non-scoring evidence.
|
||||
- The date-control and annual-control pilots do not support exact-day or
|
||||
exact-month claims from the current replay score.
|
||||
- The calibration corpus uses known positive events and partially adjudicated
|
||||
control dates. It does not establish specificity, balanced accuracy, or
|
||||
general predictive accuracy.
|
||||
- External same-chart parity is separate. Local replay must not be described as
|
||||
VedAstro, PyJHora/JHora, or jyotishganit verified until all required raw
|
||||
oracle fields are imported and compared.
|
||||
|
||||
## Excluded Local Material
|
||||
|
||||
Earlier V1/V2 snapshots, probe outputs, comparison intermediates, and planning
|
||||
files remain local. They are retained for audit but are not release evidence.
|
||||
@@ -0,0 +1,45 @@
|
||||
# 真实案例负样本日期排序 Pilot(2026-07-11)
|
||||
|
||||
## 设计
|
||||
|
||||
对 Trump 就职、DiCaprio 奥斯卡、Markle 婚姻三个 AA 案例,分别取真实事件日前后 `30/60/90/120` 天,共 8 个控制日期。控制日期只保证没有发生该项精确目标事件,不保证没有其他人生事件。
|
||||
|
||||
统一使用 V2.1;SAV/BAV 仍为非评分证据。并列采用保守排名:控制日期与真实日期同分时,控制日期排在真实日期前。
|
||||
|
||||
## 汇总
|
||||
|
||||
- 真实日期 Top-1:`0/3`。
|
||||
- 真实日期 Top-3:`0/3`。
|
||||
- Mean Reciprocal Rank:`0.1407`。
|
||||
- 平均真实日分数边际:`-1.3333`。
|
||||
- 24 个控制日期中,`10/24 = 41.67%` 达到 activation 阈值。
|
||||
- `10/24 = 41.67%` 达到 strong 阈值。
|
||||
|
||||
## 逐案
|
||||
|
||||
| 案例 | 真实分 | 最高控制分 | 真实日排名/9 | 结论 |
|
||||
|---|---:|---:|---:|---|
|
||||
| Trump 2017 就职 | 3 | 7 | 5 | 两个更早控制日期反而 strong |
|
||||
| DiCaprio 2016 奥斯卡 | 1 | 1 | 9 | 九个日期全部同分,完全无日期区分力 |
|
||||
| Markle 2018 婚姻 | 7 | 7 | 9 | 真实日和八个控制日期全部 strong |
|
||||
|
||||
## 裁决
|
||||
|
||||
当前评分器主要识别持续数月或更长的 Dasha、分盘与慢行星背景,不能从该背景中确定具体月日。进一步使用 `±1年/±2年` 的 12 个年度控制日期后,真实日 Top-1/Top-3 也只有 `33.33%`:Trump、DiCaprio 均排最后,只有 Markle 婚姻排第一。
|
||||
|
||||
由此新增硬门:
|
||||
|
||||
- `exact_day`:blocked。
|
||||
- `exact_month_from_current_replay_score`:blocked。
|
||||
- 当前最大支持精度:`unvalidated_broad_window`。
|
||||
- 事业 timing:blocked。
|
||||
- 婚姻宽窗口:partial candidate,仍需更多样本。
|
||||
|
||||
月级或日级输出只有在 PD/PrAD、Mudda/Varshaphala、精确 KP cusp、快速过境加入后,并在新的控制日期排名中通过,才能解除门控。
|
||||
|
||||
本 pilot 仍不能计算完整 balanced accuracy,因为控制日期未被独立核验为“所有同领域事件均未发生”。但它足以反证当前分数具有精确日期识别能力。
|
||||
|
||||
机器报告:
|
||||
|
||||
- `docs/benchmark/public_real_case_negative_control_pilot_2026_07_11.json`
|
||||
- `docs/benchmark/public_real_case_annual_control_pilot_2026_07_11.json`
|
||||
@@ -0,0 +1,42 @@
|
||||
# 真实案例 V2.1 计分修正与 SAV/BAV 审计(2026-07-11)
|
||||
|
||||
## 修正内容
|
||||
|
||||
1. MD 与 AD 为同一颗星时,不再重复执行整套宫位、落宫和 karaka 加分。
|
||||
2. 保留旧字段兼容,但新增真实指标名:
|
||||
- `known_event_activation_rate`
|
||||
- `strong_activation_rate`
|
||||
3. `positive_event_recall` 与 `exact_label_rate` 标记为 deprecated。
|
||||
4. 23 案例全部加入 D1 SAV/BAV、事件宫 SAV、事件日 Jupiter/Saturn 过境 SAV/BAV;本轮不参与评分。
|
||||
|
||||
## V2.1 观察结果
|
||||
|
||||
- 总体:`7 strong + 10 weak + 6 miss`。
|
||||
- known-event activation:`17/23 = 0.7391`。
|
||||
- strong activation:`7/23 = 0.3043`。
|
||||
- 事业:activation `8/12 = 0.6667`,strong `2/12 = 0.1667`。
|
||||
- 婚姻:activation `9/11 = 0.8182`,strong `5/11 = 0.4545`。
|
||||
|
||||
旧 23 案例 strong activation 为 `9/23 = 0.3913`。去重后降为 `7/23 = 0.3043`,确认重复 MD/AD 计分曾抬高强命中数量。
|
||||
|
||||
## SAV/BAV 描述性结果
|
||||
|
||||
| 分组 | 事件宫 SAV 均值 | Jupiter/Saturn 过境 SAV 均值 | 过境星自身 BAV 均值 |
|
||||
|---|---:|---:|---:|
|
||||
| 已激活 strong/weak | 28.809 | 28.735 | 3.882 |
|
||||
| miss | 30.375 | 27.583 | 3.917 |
|
||||
|
||||
当前正样本中,miss 的事件宫 SAV 均值反而高于已激活组,过境 BAV 几乎没有差异。因此:
|
||||
|
||||
- SAV 不能直接作为“高分即发生事件”的加分器。
|
||||
- SAV 更适合作为本命承载力背景,与大运、宫主、BAV 和过境共同分析。
|
||||
- 是否具有日期区分能力,必须用同人物同年度负样本验证。
|
||||
|
||||
## 技术边界
|
||||
|
||||
- 本仓 SAV 总数和七曜 BAV 总数不变量已由 `tests/test_ashtakavarga_invariants.py` 守门。
|
||||
- 本轮 23/23 案例 SAV/BAV 状态为 `used_non_scoring`。
|
||||
- 尚未完成 JHora/PyJHora 的逐星座 SAV/BAV raw parity。
|
||||
- 本结果使用已见正事件,只是校正观察,不构成 V2.1 晋级验证。
|
||||
|
||||
机器报告:`docs/benchmark/public_real_case_23_case_v21_corrected_observation_2026_07_11.json`。
|
||||
Reference in New Issue
Block a user