fix(rectification): use candidate dates for cross-midnight dasha scoring
Add date-isolated caches and regression coverage, align scoring identity, and freeze full research reruns while preserving historical artifacts. Record unresolved cache/receipt identity and end-to-end acceptance gaps for branch review only. Co-Authored-By: Claude Code <noreply@anthropic.com>
This commit is contained in:
@@ -8,7 +8,9 @@ This file is the small index for the current engineering fronts that still drive
|
||||
- [固定口径重跑与契约](sealed_holdout_rerun_2026_09_20.md):已曝光 v3 的当前实现成绩,不是新的独立盲测;确认门关闭。
|
||||
- [v5 独立封存采集协议](sealed_holdout_v5_protocol_2026_09_20.md):本轮只出协议,不采事件。
|
||||
- 前置结论:[分钟分辨率两轮研究](rectification_minute_resolution_closure_2026_09_14.md)。不重走已证伪的加权重与簇合并假设。
|
||||
- 执行状态:[本轮进度](../tasks/PROGRESS-rectification-validation-20260920.md)。
|
||||
- 验证补缺轮:[进度](../tasks/PROGRESS-rectification-validation-20260920.md)。
|
||||
- 后续候选级跨午夜日期修复:[进度与独立验收](../tasks/PROGRESS-rectification-cross-midnight-20260920.md)。当前评测转至 `reported_offset_cross_midnight_2026_09_20.json` / `sealed_holdout_rerun_cross_midnight_2026_09_20.json`;旧 JSON / freeze 原字节留存,旧文档及评测器在 `history/rectification_pre_cross_midnight_2026_09_20/`。
|
||||
- 尚未闭环:BUG-982 跨午夜簇跨度、BUG-983 凌晨日期锚点、BUG-984 时段缓存/回执身份;[验收补单](../tasks/TASK-rectification-cross-midnight-dasha-fix-20260920.md) 不代表批准实施。
|
||||
|
||||
## Shortest-Path Closure Order (2026-06-29)
|
||||
|
||||
|
||||
+12201
File diff suppressed because it is too large
Load Diff
+64
@@ -0,0 +1,64 @@
|
||||
# 申报偏差敏感性评测(2026-09-20)
|
||||
|
||||
## 结论与边界
|
||||
|
||||
**默认半径 ±15 分钟时,申报误差绝对值超过 15 分钟,记录真值就不在候选窗内;按整数分钟,第 16 分钟开始出窗。** 这是搜索窗几何边界,不需要知道真实用户偏差分布。
|
||||
|
||||
本轮完整评测预先固定 15 个偏移 × 3 档半径 × 20 例,共 900 组合。修正版全部运行完成,validator 排除 0 例;实测结果以下列最终 JSON 为准,不能用中途旧版本初扫代替。
|
||||
|
||||
**区间覆盖不等于分钟命中,模拟偏差敏感性不等于真实用户准确率。** 不运行六题真值方向回放,不以用户认可作真值,不调参、不改生产打分、不打开确认门。
|
||||
|
||||
## 口径
|
||||
|
||||
| 项目 | 口径 |
|
||||
| --- | --- |
|
||||
| 数据集 | `references/real_case_calibration/minute_rectification_holdout_v3.json`,20 例公开 AA,每例 3 件事件,已曝光 |
|
||||
| ayanamsa / node mode | `raman` / `mean` |
|
||||
| 申报偏移 | 0、±3、±5、±8、±10、±15、±20、±30 分钟 |
|
||||
| 搜索半径 | ±15 / ±30 / ±60 分钟;±120 仅说明几何边界,未评分 |
|
||||
| 候选步长 | 1 分钟,含两端;不是生产 2 分钟抽样 |
|
||||
| 12 文件历史身份 | `b15d9ea15227cd58`(完整值见 JSON) |
|
||||
| 实际评分路径 | 原生 event contribution matrix;不是 T2 的 shadow fact ranker |
|
||||
| 额外身份 | JSON 的 `production_scoring_sha256` 与 `research_implementation_sha256` 分别绑定扩展评分文件和研究适配文件;显式文件清单不是完整传递依赖锁 |
|
||||
| 头名与排名 | 分数降序,再用既有 `_opaque_winner` 同源 SHA-256 规则排同分;不选择最接近真值的候选 |
|
||||
| 交付区间 | 初始未答题:signature clusters 中落后头名不足 8 分的仍有效簇外包范围;无新回答、无淘汰,不是完整问答链交付效果 |
|
||||
| 跨午夜 | 候选、误差及区间使用完整日期;离线矩阵按候选日期分组,合并后全窗排序/聚类 |
|
||||
| 独立性 | `is_blind_evaluation=false`,已曝光集;评分/区间定稿后才揭示标签,仅证明程序标签隔离 |
|
||||
|
||||
## 实测表
|
||||
|
||||
每格依次为 **真值在窗比例 / 哈希头名命中 / 初始交付区间覆盖**,分母均为 20 例。列为申报偏移分钟,行为搜索半径。统一口径:v3、raman/mean、1 分钟步长;历史打分身份 `b15d9ea15227cd58`,扩展原生评分身份 `115c3fbcffdaff49`,研究适配身份 `8d0c613dc0900adb`。
|
||||
|
||||
| 半径 | -30 | -20 | -15 | -10 | -8 | -5 | -3 | 0 | +3 | +5 | +8 | +10 | +15 | +20 | +30 |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
| ±15 | 0/0/0 | 0/0/0 | 100/10/95 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/95 | 100/15/95 | 0/0/0 | 0/0/0 |
|
||||
| ±30 | 100/10/95 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/95 | 100/15/95 |
|
||||
| ±60 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 | 100/5/100 |
|
||||
|
||||
表内全部数字单位为 %。这只是 20 例公开已曝光案例的条件敏感性,不附会总体置信度。即使真值在候选窗中,±15 的 -15/+10/+15 档及 ±30 的 -30/+20/+30 档仍各有一例未被初始交付范围覆盖;窗口包含与交付覆盖必须分开。±60 的所有采样偏移虽保持覆盖,但真正选中的头名只有 5%,不能用宽区间覆盖声称分钟准确。
|
||||
|
||||
修正版与错误初版的汇总比例恰好相同,不表示初版评分正确;身份、逐候选计算与跨日回归以修正版为准。
|
||||
|
||||
## 放宽能救回什么
|
||||
|
||||
- 真值在候选窗的条件始终是 `abs(offset) <= radius`。本次偏移最大绝对值 30,因此半径 30 / 60 / 120 在几何上都能包含本次所有偏移标签。
|
||||
- 对 ±15 已出窗的 ±20 / ±30 两组偏移,放宽到 ±30 可把**候选窗包含率**从 0 恢复到 100%;±60 / ±120 对这些特定偏移不再增加几何包含率。
|
||||
- 半径 ±30 自第 31 个整数分钟偏移出窗,±60 自第 61 个、±120 自第 121 个。放宽不是无条件保证,更不代表排名命中或交付覆盖同样恢复。
|
||||
- ±120 未跑评分,不报告其头名命中率或交付覆盖率。真实用户申报偏差分布未知,不估算现实用户中有多少人能被放宽救回。
|
||||
|
||||
## 初扫作废与独立复核
|
||||
|
||||
首版正确枚举了跨日候选,但矩阵内 transition proximity 仍共用窗口起日,独立审查发现跨日候选得分偏移。因此旧初扫作废,修正版标记 `replay_revision=candidate_date_grouped_v2`,并登记 `supersedes=initial_sweep_invalidated_cross_midnight_transition_anchor`。
|
||||
|
||||
修复仅在本次新脚本 `score_window()` 做按日期分组适配,未改冻结打分文件。实引擎回归验证跨日窗口的分组计算与逐候选独立重算相同;聚类和交付范围在合并后全窗运行。**生产日期处理未在本单修复,不能称为生产端到端回放。**
|
||||
|
||||
## 可复算产物与验收
|
||||
|
||||
```bash
|
||||
python scripts/research/reported_offset_sweep.py --json
|
||||
python -m pytest tests/test_reported_offset_research.py tests/test_rectification_validation_integrity_gate.py -q
|
||||
```
|
||||
|
||||
`reported_offset_2026_09_20.json` 保留逐例四项必要输出:真值是否在窗、真值排名、头名分钟误差、交付区间覆盖,以及每个格子的分母、全部参数和实现哈希。仅保存公开集序号,不保存出生分钟、坐标或事件正文。
|
||||
|
||||
排名不能与 T2 的竞争排名 top-1/top-3 混比:T2 按旧协议允许并列多个分钟同时算第一;本表把同分完全按哈希打破。最终测试数字与已知基线失败见 [进度](../tasks/PROGRESS-rectification-validation-20260920.md)。真实分布采集与新未曝光样本见 [v5 协议](sealed_holdout_v5_protocol_2026_09_20.md)。
|
||||
+43
@@ -0,0 +1,43 @@
|
||||
{
|
||||
"record_version": "exposed-v3-fixed-protocol-rerun-v1",
|
||||
"frozen_at_utc": "2026-09-20T03:27:26.670869+00:00",
|
||||
"dataset_path": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"algorithm_version": "birth-time-event-fact-ranker-v4-shadow",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"historical_frozen_sha256": "f41c298dd6cdcebe7a93e632f7191954be7987f6012ca2a34cecb6e447fbf196",
|
||||
"evaluator_sha256": "3bf35de5f8a0b9588808178d47ba2457c398cf2fd2d2fcb03c5c71420f998da5",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"candidate_radius_minutes": [
|
||||
10
|
||||
],
|
||||
"minute_step": 1,
|
||||
"release_metrics": {
|
||||
"top_1_rate_minimum": 0.6,
|
||||
"top_3_rate_minimum": 0.85,
|
||||
"mean_absolute_minute_error_maximum": 2.0,
|
||||
"false_confirmation_rate_maximum": 0.05,
|
||||
"correct_insufficient_evidence_rejection_rate_minimum": 0.9
|
||||
},
|
||||
"results_previously_seen": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"must_not_use_for_tuning": true,
|
||||
"tie_breaker": "sha256(benchmark_id:case_id:candidate_time); published minute is never passed to the ranker",
|
||||
"metric_rank_definition": "competition_rank_1_plus_strictly_higher_scores_legacy_protocol",
|
||||
"extra_metric_rank_definition": "score_desc_then_existing_opaque_sha256_total_order"
|
||||
}
|
||||
+626
@@ -0,0 +1,626 @@
|
||||
{
|
||||
"scope": "fixed_protocol_previously_exposed_v3_rerun",
|
||||
"evaluated_on": "2026-09-20",
|
||||
"frozen_record": {
|
||||
"record_version": "exposed-v3-fixed-protocol-rerun-v1",
|
||||
"frozen_at_utc": "2026-09-20T03:27:26.670869+00:00",
|
||||
"dataset_path": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"algorithm_version": "birth-time-event-fact-ranker-v4-shadow",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"historical_frozen_sha256": "f41c298dd6cdcebe7a93e632f7191954be7987f6012ca2a34cecb6e447fbf196",
|
||||
"evaluator_sha256": "3bf35de5f8a0b9588808178d47ba2457c398cf2fd2d2fcb03c5c71420f998da5",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"candidate_radius_minutes": [
|
||||
10
|
||||
],
|
||||
"minute_step": 1,
|
||||
"release_metrics": {
|
||||
"top_1_rate_minimum": 0.6,
|
||||
"top_3_rate_minimum": 0.85,
|
||||
"mean_absolute_minute_error_maximum": 2.0,
|
||||
"false_confirmation_rate_maximum": 0.05,
|
||||
"correct_insufficient_evidence_rejection_rate_minimum": 0.9
|
||||
},
|
||||
"results_previously_seen": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"must_not_use_for_tuning": true,
|
||||
"tie_breaker": "sha256(benchmark_id:case_id:candidate_time); published minute is never passed to the ranker",
|
||||
"metric_rank_definition": "competition_rank_1_plus_strictly_higher_scores_legacy_protocol",
|
||||
"extra_metric_rank_definition": "score_desc_then_existing_opaque_sha256_total_order"
|
||||
},
|
||||
"implementation_hash_matches_at_replay": true,
|
||||
"dataset_hash_matches_at_replay": true,
|
||||
"source_audit_status": "corrected_known_date_errors",
|
||||
"validation_status": "blocked_awaiting_public_aa_cases",
|
||||
"valid_public_aa_cases": 20,
|
||||
"excluded_cases": [],
|
||||
"trial_count": 20,
|
||||
"trials": [
|
||||
{
|
||||
"case_ordinal": 1,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 5,
|
||||
"opaque_true_rank": 8,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 2,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 2,
|
||||
"opaque_true_rank": 8,
|
||||
"minute_error": 3,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 3,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 12,
|
||||
"minute_error": 5,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 4,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 1,
|
||||
"minute_error": 0,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 5,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 10,
|
||||
"minute_error": 1,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 6,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 7,
|
||||
"opaque_true_rank": 12,
|
||||
"minute_error": 6,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 7,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 5,
|
||||
"opaque_true_rank": 7,
|
||||
"minute_error": 9,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 8,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 5,
|
||||
"opaque_true_rank": 8,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 9,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 3,
|
||||
"minute_error": 1,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"missing_mandatory_layers",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"missing_mandatory_layers",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 10,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 8,
|
||||
"minute_error": 5,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 11,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 6,
|
||||
"opaque_true_rank": 17,
|
||||
"minute_error": 5,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 12,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 4,
|
||||
"minute_error": 5,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 13,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 5,
|
||||
"opaque_true_rank": 7,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 14,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 17,
|
||||
"opaque_true_rank": 19,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 15,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 7,
|
||||
"minute_error": 9,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"missing_mandatory_layers",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"missing_mandatory_layers",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 16,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 12,
|
||||
"opaque_true_rank": 12,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 17,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 6,
|
||||
"opaque_true_rank": 6,
|
||||
"minute_error": 6,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 18,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 11,
|
||||
"opaque_true_rank": 14,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 19,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 6,
|
||||
"minute_error": 4,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_ordinal": 20,
|
||||
"candidate_count": 21,
|
||||
"event_count": 3,
|
||||
"true_rank": 1,
|
||||
"opaque_true_rank": 7,
|
||||
"minute_error": 10,
|
||||
"would_confirm": false,
|
||||
"false_confirmation": false,
|
||||
"insufficient_evidence_rejected": true,
|
||||
"full_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
],
|
||||
"sparse_trial_reasons": [
|
||||
"unique_minute_not_found",
|
||||
"insufficient_events",
|
||||
"insufficient_domains",
|
||||
"neighbor_stability_not_passed",
|
||||
"leave_one_event_out_not_passed",
|
||||
"top_candidate_feature_not_unique",
|
||||
"fact_ranker_v4_holdout_not_ready"
|
||||
]
|
||||
}
|
||||
],
|
||||
"metrics": {
|
||||
"top_1_rate": 0.45,
|
||||
"top_3_rate": 0.5,
|
||||
"mean_absolute_minute_error": 6.45,
|
||||
"false_confirmation_rate": 0.0,
|
||||
"correct_insufficient_evidence_rejection_rate": 1.0,
|
||||
"confirmation_coverage_rate": 0.0
|
||||
},
|
||||
"metric_gates_passed": false,
|
||||
"opaque_exact_top_1_rate": 0.05,
|
||||
"opaque_exact_top_3_rate": 0.1,
|
||||
"official_valid_independent_blind": false,
|
||||
"official_blind_trial_count": 0,
|
||||
"is_blind_evaluation": false,
|
||||
"truth_hidden_from_ranker": true,
|
||||
"results_previously_seen": true,
|
||||
"verified_minute_claim_allowed": false,
|
||||
"status": "blocked_independent_blind_evidence",
|
||||
"boundary": "Three-event low-information protocol, not a mathematical accuracy lower bound and not representative of real sessions. Historical v3/v4 exposure cannot be undone by freezing today's scorer. No tuning or release claims."
|
||||
}
|
||||
+72
@@ -0,0 +1,72 @@
|
||||
# 当前冻结实现的 v3 固定口径重跑(2026-09-20)
|
||||
|
||||
## 结论
|
||||
|
||||
**已补上当前实现可归属、可复算的成绩;没有补成新的独立官方盲测。** 既有 v3/v4 人物与成绩已经曝光,重新冻结今天的打分文件不能消除曝光。任务书 T2 中「首次口径干净的官方盲测」在既有 BUG-427 / BUG-428 红线下仍为 **blocked**,不得改名通过。确认门保持关闭。
|
||||
|
||||
这是每例只有 3 件事的低信息量(任务书称「下限」)口径,不代表真实会话;事件少不等于数学上保证成绩更低,因此也不是准确率下界。
|
||||
|
||||
## 固定口径与可审计文件
|
||||
|
||||
| 项目 | 本次实际值 |
|
||||
| --- | --- |
|
||||
| 数据集 | `references/real_case_calibration/minute_rectification_holdout_v3.json` |
|
||||
| 数据集 SHA-256 | `45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea` |
|
||||
| 计算规格 | ayanamsa `raman` / node mode `mean` |
|
||||
| 半径 / 步长 | ±10 分钟 / 1 分钟,含两端;真值为窗心 |
|
||||
| 算法 | `birth-time-event-fact-ranker-v4-shadow`,不是生产 V9 问答链 |
|
||||
| 当前打分身份 | `b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18`(前缀 `b15d9ea15227cd58`) |
|
||||
| 样本 / 事件 | 20 例 / 每例 3 件;validator 排除 0 例 |
|
||||
| 资料审计状态 | `corrected_known_date_errors`,不等同全新独立人工重审 |
|
||||
| 排名标签隔离 | 排序和稀疏证据判定完成后才读取真值标签 |
|
||||
| 独立官方盲测 | `official_valid_independent_blind=false`;有效官方盲测试次仍为 0 |
|
||||
|
||||
冻结记录:`sealed_holdout_rerun_2026_09_20.freeze.json`。先以独占创建模式记录 12 个文件、数据集与评测器字节哈希,再执行评分;评分前验证全部身份,任何漂移即拒绝。历史 v3 `frozen_scoring` 原值 `f41c298dd6cdcebe…` 未改。
|
||||
|
||||
脱敏逐例结果:`sealed_holdout_rerun_2026_09_20.json`,只含公开集序号、排名、误差、门禁原因,不含出生分钟、坐标或事件正文。复算:
|
||||
|
||||
```bash
|
||||
python scripts/research/sealed_holdout_rerun.py --json
|
||||
```
|
||||
|
||||
## 五项预注册指标与门槛
|
||||
|
||||
本表统一口径:v3 / `raman` / `mean` / ±10 / 1 分钟 / `b15d9ea15227cd58`。
|
||||
|
||||
| 指标 | 实测 | 预注册门槛 | 数值对照 |
|
||||
| --- | ---: | ---: | --- |
|
||||
| top-1(旧协议竞争排名,含并列) | 0.45 | ≥0.60 | 未通过 |
|
||||
| top-3(旧协议竞争排名,含并列) | 0.50 | ≥0.85 | 未通过 |
|
||||
| 头名平均绝对分钟误差(哈希打破并列) | 6.45 分钟 | ≤2.00 分钟 | 未通过 |
|
||||
| 误确认率(全样本分母) | 0.00 | ≤0.05 | 通过,但无确认覆盖,不能证明确认精度 |
|
||||
| 信息不足正确拒绝率(每例仅首件事件) | 1.00 | ≥0.90 | 通过 |
|
||||
| 确认覆盖率(补充) | 0.00 | 本单禁止抬高 | 保持关闭 |
|
||||
| 作为独立盲测发布成绩的资格 | 不适用 | 新封存、未曝光、独立审计 | blocked,不可当发布指标 |
|
||||
|
||||
**重要口径纠正:旧 top-1/top-3 不是哈希选中那个分钟的命中率。** 现有评测的 `true_rank = 1 + 严格高于真值分数的候选数` 会把同分的多个分钟全算第一;本次为保证与旧协议可比保留它,同时增加哈希全序的真实选中指标,不暗改预注册定义。
|
||||
|
||||
| 附加指标(同一上述口径) | 实测 |
|
||||
| --- | ---: |
|
||||
| `_opaque_winner` 真正选中的头名命中 | 0.05 |
|
||||
| 同一哈希全序中的 top-3 命中 | 0.10 |
|
||||
|
||||
这两个数字不能与旧 0.45/0.50 混称同一种准确率。没有据此改权重、换半径、换案例或剔除失败例。
|
||||
|
||||
## 契约改了什么、没有改什么
|
||||
|
||||
- `current_tree_scorer` 绑定本次真实哈希及新报告,新增固定口径试次数与成绩;旧 `current_tree_unfrozen_diagnostic` 历史块保留。
|
||||
- `source_audit_status` 对齐数据集;`evaluated_on` 表示最新固定口径重跑日期,并新增作用域说明。历史顶层指标仍有 `metrics_produced_by` 的旧日期与来源,未伪装成本次结果。
|
||||
- `status`、`valid_public_aa_cases`、`required_cases`、`top_1_rate`、`confirmation_coverage_rate`、`sealed_benchmark_id` 六个 runtime 键的值与类型全不变;尤其 `status=not_ready`、确认覆盖率为零。
|
||||
- `official_eval_implementation_hash_matches=false`、`official_eval_trial_count=0` 保留;新增重跑记录不冒充已重新获得独立性。
|
||||
- 未改任何冻结打分文件、生产判定、现有 `_candidate_moments()` 或 `request_from_case()`。
|
||||
|
||||
## 防复发与既有断言变更
|
||||
|
||||
新增 `tests/test_sealed_holdout_contract_freshness.py` 钉死当前树身份、审计状态、冻结记录、报告汇总与独立性声明;冻结身份不符必须在计算前拒绝。`tests/test_rectification_validation_integrity_gate.py` 将其和偏差评测回归纳入现有 `test_rectification_*.py` 快速门,不改 workflow。
|
||||
|
||||
| 原断言 | 新断言 | 原因 |
|
||||
| --- | --- | --- |
|
||||
| Python / 前端:`current_tree_scorer.implementation_sha256` 等于旧 post-audit sidecar 的哈希 | 等于新固定口径报告 `frozen_record.implementation_sha256`,另验证当前树与新报告 | 契约当前树块的含义就是当前实现;继续绑定旧报告会强制过期。新增身份断言更严格,不弱化门禁 |
|
||||
| 官方标志 false、官方试次 0、六键等于历史报告、确认门全套行为断言 | 原值保留,另加固定口径重跑身份/试次和独立盲测 false | 固定口径重跑不是新的独立盲测,不借元数据迁移打开确认门 |
|
||||
|
||||
本机 Python 3.11.7 可跑;`python3` Windows 别名退出 49。该任务定向命令的最终测试清单与基线失败对照由 `PROGRESS-rectification-validation-20260920.md` 统一记录。特别是既有 v2 冻结身份测试不属本次刷新范围,不得为了通过而改其历史封存或断言。
|
||||
@@ -0,0 +1,55 @@
|
||||
{
|
||||
"scope": "immutable_pre_fix_artifacts_not_current_results",
|
||||
"baseline_commit": "932f2fff",
|
||||
"reason": "BUG-981 production candidate-date fix; prior T1 used offline date grouping and T2 shadow scorer; numerical changes not assumed",
|
||||
"files": [
|
||||
{
|
||||
"source_path": "docs/research/reported_offset_2026_09_20.md",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/docs/research/reported_offset_2026_09_20.md",
|
||||
"sha256": "1033fe68ff4b2802ab321ef2f76e558dc21c3c534486954bf03af8b7f114f6df",
|
||||
"size_bytes": 6018
|
||||
},
|
||||
{
|
||||
"source_path": "docs/research/reported_offset_2026_09_20.json",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/docs/research/reported_offset_2026_09_20.json",
|
||||
"sha256": "9878f2b50c957a470fafcb2ed0a9eb16405c3f6fa422ea50f467e84b0189ace1",
|
||||
"size_bytes": 324672
|
||||
},
|
||||
{
|
||||
"source_path": "docs/research/sealed_holdout_rerun_2026_09_20.md",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/docs/research/sealed_holdout_rerun_2026_09_20.md",
|
||||
"sha256": "dab1bf59748efab481aee394ab25822ee51b0665bf66c769afdfb4d53a492d29",
|
||||
"size_bytes": 5818
|
||||
},
|
||||
{
|
||||
"source_path": "docs/research/sealed_holdout_rerun_2026_09_20.json",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/docs/research/sealed_holdout_rerun_2026_09_20.json",
|
||||
"sha256": "40df220bfcb51a683b31fbd626896e47a5aed7ee4f30ff669543691f7ce53c1e",
|
||||
"size_bytes": 20052
|
||||
},
|
||||
{
|
||||
"source_path": "docs/research/sealed_holdout_rerun_2026_09_20.freeze.json",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/docs/research/sealed_holdout_rerun_2026_09_20.freeze.json",
|
||||
"sha256": "d17651cb50224acdd1af7a4692c777ed169b216a962a0675371254ce561d0921",
|
||||
"size_bytes": 1898
|
||||
},
|
||||
{
|
||||
"source_path": "references/rectification_sealed_holdout.v1.json",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/references/rectification_sealed_holdout.v1.json",
|
||||
"sha256": "863806992ee7162e959bb319fd82bd75f13fadea6218e197dd4e78fd2ffd0279",
|
||||
"size_bytes": 3705
|
||||
},
|
||||
{
|
||||
"source_path": "scripts/research/reported_offset_sweep.py",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/scripts/research/reported_offset_sweep.py.txt",
|
||||
"sha256": "87f87859bb399da1e860ab5e750e68088401a5a82fbba10ba709c17e5266167c",
|
||||
"size_bytes": 11001
|
||||
},
|
||||
{
|
||||
"source_path": "scripts/research/sealed_holdout_rerun.py",
|
||||
"archive_path": "docs/research/history/rectification_pre_cross_midnight_2026_09_20/scripts/research/sealed_holdout_rerun.py.txt",
|
||||
"sha256": "3bf35de5f8a0b9588808178d47ba2457c398cf2fd2d2fcb03c5c71420f998da5",
|
||||
"size_bytes": 7482
|
||||
}
|
||||
]
|
||||
}
|
||||
+21
@@ -0,0 +1,21 @@
|
||||
{
|
||||
"scope": "superseded_preparation_audit",
|
||||
"reason": "Existing report overwrite denied; retained initial freeze and completed 20-case preparation replay unchanged. Stopped 900 sweep before result write. Final report paths changed; evaluator identity re-frozen before new final replay.",
|
||||
"files": [
|
||||
{
|
||||
"path": "docs/research/sealed_holdout_rerun_cross_midnight_2026_09_20.freeze.json",
|
||||
"sha256": "6638e52191fed63ed3a2e99da149ba3269ea4367d8e60b8408f2b05165ca280e",
|
||||
"status": "superseded_preparation_not_final_report"
|
||||
},
|
||||
{
|
||||
"path": "docs/research/reported_offset_cross_midnight_2026_09_20.freeze.json",
|
||||
"sha256": "80004557170595e7ae222ecf538482e2f1f279f3858896e2553ba0d565a45eea",
|
||||
"status": "superseded_preparation_not_final_report"
|
||||
},
|
||||
{
|
||||
"path": "docs/research/sealed_holdout_rerun_2026_09_20.cross_midnight.rerun.json",
|
||||
"sha256": "f4b0ed8a8ad17604940f9022339854571501d95af863820f58501a4981d4e4e1",
|
||||
"status": "superseded_preparation_not_final_report"
|
||||
}
|
||||
]
|
||||
}
|
||||
+79
@@ -0,0 +1,79 @@
|
||||
{
|
||||
"sealed_benchmark_id": "minute_rectification_fact_ranker_v4_holdout_v3",
|
||||
"status": "not_ready",
|
||||
"valid_public_aa_cases": 20,
|
||||
"required_cases": 20,
|
||||
"top_1_rate": 0.15,
|
||||
"confirmation_coverage_rate": 0.0,
|
||||
"previous_pilot_id": "minute_rectification_holdout_v2",
|
||||
"dataset_path": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"report_path": "references/real_case_calibration/minute_rectification_holdout_v3_report.json",
|
||||
"source_audit_status": "corrected_known_date_errors",
|
||||
"evaluated_on": "2026-09-20",
|
||||
"evaluated_on_scope": "latest_fixed_protocol_rerun_not_historical_runtime_metrics",
|
||||
"historical_source_audit_status": "invalidated_after_replay",
|
||||
"report_status": "invalidated_after_source_audit",
|
||||
"trial_count": 20,
|
||||
"top_3_rate": 0.25,
|
||||
"mean_absolute_minute_error": 6.95,
|
||||
"metrics_produced_by": {
|
||||
"implementation_sha256": "f41c298dd6cdcebe7a93e632f7191954be7987f6012ca2a34cecb6e447fbf196",
|
||||
"algorithm_version": "birth-time-event-fact-ranker-v4-shadow",
|
||||
"implementation_hash_matches_at_replay": true,
|
||||
"source_report": "references/real_case_calibration/minute_rectification_holdout_v3_report.json",
|
||||
"evaluated_on": "2026-07-21"
|
||||
},
|
||||
"current_tree_scorer": {
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"matches_metrics_scorer": false,
|
||||
"official_eval_implementation_hash_matches": false,
|
||||
"official_eval_trial_count": 0,
|
||||
"source_report": "docs/research/sealed_holdout_rerun_2026_09_20.json",
|
||||
"fixed_protocol_rerun_trial_count": 20,
|
||||
"fixed_protocol_rerun_hash_matches": true,
|
||||
"metrics": {
|
||||
"top_1_rate": 0.45,
|
||||
"top_3_rate": 0.5,
|
||||
"mean_absolute_minute_error": 6.45,
|
||||
"false_confirmation_rate": 0.0,
|
||||
"correct_insufficient_evidence_rejection_rate": 1.0,
|
||||
"confirmation_coverage_rate": 0.0
|
||||
}
|
||||
},
|
||||
"current_tree_fixed_protocol_rerun": {
|
||||
"evaluated_on": "2026-09-20",
|
||||
"report_path": "docs/research/sealed_holdout_rerun_2026_09_20.json",
|
||||
"freeze_record_path": "docs/research/sealed_holdout_rerun_2026_09_20.freeze.json",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"scorer_frozen_before_rerun": true,
|
||||
"source_audit_status": "corrected_known_date_errors",
|
||||
"trial_count": 20,
|
||||
"official_valid_independent_blind": false,
|
||||
"is_blind_evaluation": false,
|
||||
"truth_hidden_from_ranker": true,
|
||||
"results_previously_seen": true,
|
||||
"must_not_claim_as_release_metrics": true,
|
||||
"must_not_use_for_tuning": true,
|
||||
"verified_minute_claim_allowed": false,
|
||||
"metric_gates_passed": false,
|
||||
"events_per_case": 3,
|
||||
"boundary": "First current-hash fixed-protocol rerun in this task, not a first independent official blind evaluation. Three-event low-information protocol is not representative of real sessions and is not a mathematical accuracy lower bound. Historical v3/v4 score exposure requires a fresh sealed set."
|
||||
},
|
||||
"current_tree_unfrozen_diagnostic": {
|
||||
"is_blind_evaluation": false,
|
||||
"results_already_seen": true,
|
||||
"scorer_frozen": false,
|
||||
"source_audit_passed": false,
|
||||
"must_not_claim_as_release_metrics": true,
|
||||
"must_not_use_for_tuning": true,
|
||||
"report_path": "references/real_case_calibration/minute_rectification_holdout_v3_post_audit_diagnostic_report.json",
|
||||
"implementation_sha256": "99730c84c6434f52669a05e4d9a4a87df0a218e3237b5436b933252f2028384e",
|
||||
"trial_count": 20,
|
||||
"top_1_rate": 0.45,
|
||||
"top_3_rate": 0.5,
|
||||
"mean_absolute_minute_error": 5.8,
|
||||
"confirmation_coverage_rate": 0.0,
|
||||
"metric_gates_passed": false,
|
||||
"verified_minute_claim_allowed": false
|
||||
}
|
||||
}
|
||||
+211
@@ -0,0 +1,211 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Reported-centre sensitivity on exposed public cases, without oracle answers."""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from datetime import datetime, timedelta
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.active_rectification_event_engine import (
|
||||
AYANAMSA, NODE_MODE, compute_candidate_static_contexts,
|
||||
)
|
||||
from scripts.minute_rectification_blind_eval import _opaque_winner, implementation_sha256
|
||||
from scripts.minute_rectification_holdout_validator import validate
|
||||
from scripts.rectification.candidate_contrast import select_signature_representatives
|
||||
from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix
|
||||
from scripts.research.cluster_width_lib import SEPARATION_LEAD, still_valid_public
|
||||
from scripts.research.probe_supply_after_six import request_from_case
|
||||
from scripts.research.sealed_holdout_rerun import DATASET, file_sha256, opaque_order
|
||||
|
||||
OFFSETS = (-30, -20, -15, -10, -8, -5, -3, 0, 3, 5, 8, 10, 15, 20, 30)
|
||||
RADII = (15, 30, 60)
|
||||
MINUTE_STEP = 1
|
||||
PRODUCTION_FILES = [
|
||||
"scripts/rectification/scoring_service.py",
|
||||
"scripts/rectification/dasha_transition_proximity.py",
|
||||
"scripts/rectification/candidate_contrast.py",
|
||||
"scripts/rectification/case_holdout.py",
|
||||
"scripts/rectification/contracts.py",
|
||||
"scripts/rectification/event_probes.py",
|
||||
]
|
||||
RESEARCH_FILES = [
|
||||
"scripts/research/reported_offset_sweep.py",
|
||||
"scripts/research/cluster_width_lib.py",
|
||||
"scripts/research/probe_supply_after_six.py",
|
||||
"scripts/research/sealed_holdout_rerun.py",
|
||||
"scripts/minute_rectification_blind_eval.py",
|
||||
]
|
||||
|
||||
|
||||
def shifted_window(case: dict[str, Any], offset: int, radius: int) -> tuple[dict[str, Any], list[datetime]]:
|
||||
if radius not in (*RADII, 120):
|
||||
raise ValueError("unsupported_radius")
|
||||
true_at = datetime.fromisoformat(f"{case['birth']['date']}T{case['birth']['time']}:00")
|
||||
reported_at = true_at + timedelta(minutes=offset)
|
||||
start = reported_at - timedelta(minutes=radius)
|
||||
end = reported_at + timedelta(minutes=radius)
|
||||
candidates = [start + timedelta(minutes=i) for i in range(0, radius * 2 + 1, MINUTE_STEP)]
|
||||
# Reuse only the established event normalization, then replace the window.
|
||||
# The ranker sees neither the truth label nor the simulated offset.
|
||||
request = request_from_case(case)
|
||||
request.update({
|
||||
"birth_date": start.date().isoformat(),
|
||||
"start_time": start.strftime("%H:%M"),
|
||||
"end_time": end.strftime("%H:%M"),
|
||||
"minute_step": MINUTE_STEP,
|
||||
"ayanamsa": AYANAMSA,
|
||||
"node_mode": NODE_MODE,
|
||||
})
|
||||
return request, candidates
|
||||
|
||||
|
||||
def score_window(request: dict[str, Any], contexts: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
"""Keep auxiliary transition anchors on each candidate's actual date.
|
||||
|
||||
The production matrix's transition-proximity helper accepts one birth date
|
||||
per call, unlike the static chart layer which reads candidate_at. Grouping
|
||||
is only an offline adapter; it does not change the production scorer.
|
||||
"""
|
||||
by_date: dict[str, list[dict[str, Any]]] = {}
|
||||
for context in contexts:
|
||||
by_date.setdefault(context["candidate_at"].date().isoformat(), []).append(context)
|
||||
by_time = {}
|
||||
for candidate_date, group in by_date.items():
|
||||
dated_request = {**request, "birth_date": candidate_date}
|
||||
built = build_event_contribution_matrix(dated_request, static_contexts=group)
|
||||
by_time.update({row["time"]: row for row in score_from_matrix(dated_request, built)})
|
||||
return [by_time[context["candidate_at"].strftime("%H:%M")] for context in contexts]
|
||||
|
||||
|
||||
def delivery_moments(public: list[dict[str, Any]], candidates: list[datetime]) -> list[datetime]:
|
||||
"""Initial delivery envelope in date-aware order; no truth-based narrowing."""
|
||||
valid = still_valid_public(public, {}, lead=SEPARATION_LEAD)
|
||||
clocks = {str(value)[:5] for row in valid for value in (row.get("cluster_times") or [row["time"]])}
|
||||
return [candidate for candidate in candidates if candidate.strftime("%H:%M") in clocks]
|
||||
|
||||
|
||||
def reveal_metrics(
|
||||
rows: list[dict[str, Any]], candidates: list[datetime], delivery: list[datetime],
|
||||
truth: datetime, benchmark_id: str, case_id: str,
|
||||
) -> dict[str, Any]:
|
||||
ordered = opaque_order(benchmark_id, case_id, rows)
|
||||
predicted_clock = _opaque_winner(benchmark_id, case_id, rows)
|
||||
by_clock = {candidate.strftime("%H:%M"): candidate for candidate in candidates}
|
||||
if len(by_clock) != len(candidates):
|
||||
raise ValueError("ambiguous_candidate_clock")
|
||||
predicted = by_clock[predicted_clock]
|
||||
within = min(candidates) <= truth <= max(candidates)
|
||||
in_candidates = truth in candidates
|
||||
rank = next((i for i, row in enumerate(ordered, 1) if by_clock[row["time"]] == truth), None)
|
||||
covered = bool(delivery) and min(delivery) <= truth <= max(delivery)
|
||||
return {
|
||||
"truth_in_window": within,
|
||||
"truth_in_candidates": in_candidates,
|
||||
"true_rank": rank,
|
||||
"top_1_hit": predicted == truth,
|
||||
"top_1_minute_error": abs((predicted - truth).total_seconds()) / 60,
|
||||
"delivery_covers_truth": covered,
|
||||
"candidate_count": len(candidates),
|
||||
"delivery_width_minutes": int((max(delivery) - min(delivery)).total_seconds() / 60) + 1 if delivery else 0,
|
||||
}
|
||||
|
||||
|
||||
def summarize(trials: list[dict[str, Any]], radii: tuple[int, ...], offsets: tuple[int, ...]) -> list[dict[str, Any]]:
|
||||
result = []
|
||||
for radius in radii:
|
||||
for offset in offsets:
|
||||
group = [row for row in trials if row["radius_minutes"] == radius and row["offset_minutes"] == offset]
|
||||
count = len(group)
|
||||
result.append({
|
||||
"radius_minutes": radius, "offset_minutes": offset, "trial_count": count,
|
||||
**{key: round(sum(row[source] for row in group) / count, 4) if count else None for key, source in (
|
||||
("truth_in_window_rate", "truth_in_window"),
|
||||
("top_1_rate", "top_1_hit"),
|
||||
("delivery_coverage_rate", "delivery_covers_truth"),
|
||||
("mean_absolute_minute_error", "top_1_minute_error"),
|
||||
)},
|
||||
})
|
||||
return result
|
||||
|
||||
|
||||
def run(dataset: Path = DATASET, radii: tuple[int, ...] = RADII, offsets: tuple[int, ...] = OFFSETS) -> dict[str, Any]:
|
||||
manifest = json.loads(dataset.read_text(encoding="utf-8"))
|
||||
frozen_files = manifest["frozen_scoring"]["files"]
|
||||
scoring_files = sorted(set(frozen_files + PRODUCTION_FILES))
|
||||
starting_hash = implementation_sha256(scoring_files)
|
||||
validation = validate(dataset)
|
||||
invalid = validation["invalid_cases"]
|
||||
trials = []
|
||||
for index, case in enumerate(manifest["cases"], 1):
|
||||
if case["case_id"] in invalid:
|
||||
continue
|
||||
# Reuse static chart calculations, not window-dependent scores or ranks.
|
||||
contexts_by_moment: dict[datetime, dict[str, Any]] = {}
|
||||
for radius in radii:
|
||||
for offset in offsets:
|
||||
request, candidates = shifted_window(case, offset, radius)
|
||||
missing = [moment for moment in candidates if moment not in contexts_by_moment]
|
||||
if missing:
|
||||
contexts_by_moment.update(zip(missing, compute_candidate_static_contexts(request, candidates=missing), strict=True))
|
||||
contexts = [contexts_by_moment[moment] for moment in candidates]
|
||||
rows = score_window(request, contexts)
|
||||
public = select_signature_representatives(rows, contexts)
|
||||
delivery = delivery_moments(public, candidates)
|
||||
# All scores and the delivery envelope are locked before reveal.
|
||||
truth = datetime.fromisoformat(f"{case['birth']['date']}T{case['birth']['time']}:00")
|
||||
trials.append({
|
||||
"case_ordinal": index, "offset_minutes": offset, "radius_minutes": radius,
|
||||
**reveal_metrics(rows, candidates, delivery, truth, manifest["benchmark_id"], case["case_id"]),
|
||||
})
|
||||
if implementation_sha256(scoring_files) != starting_hash:
|
||||
raise ValueError("scorer_changed_during_sweep")
|
||||
return {
|
||||
"scope": "reported_offset_sensitivity_not_product_accuracy",
|
||||
"specification": {
|
||||
"ayanamsa": AYANAMSA, "node_mode": NODE_MODE,
|
||||
"radii_minutes": list(radii), "offsets_minutes": list(offsets), "minute_step": MINUTE_STEP,
|
||||
"dataset": dataset.relative_to(ROOT).as_posix(), "dataset_sha256": file_sha256(dataset),
|
||||
"dataset_benchmark_id": manifest["benchmark_id"],
|
||||
"implementation_sha256": implementation_sha256(frozen_files),
|
||||
"implementation_sha256_prefix": implementation_sha256(frozen_files)[:16],
|
||||
"production_scoring_files": scoring_files, "production_scoring_sha256": starting_hash,
|
||||
"research_files": RESEARCH_FILES,
|
||||
"research_implementation_sha256": implementation_sha256(RESEARCH_FILES),
|
||||
"hash_scope": "explicit_identity_file_sets_not_a_transitive_dependency_lock",
|
||||
"scorer": "native_event_contribution_matrix_not_shadow_fact_ranker",
|
||||
"delivery": "initial_signature_clusters_peak_gap_lt_8_envelope_no_answers_no_elimination",
|
||||
"rank": "score_desc_then_sha256(benchmark_id:case_id:candidate_time)",
|
||||
"grid": "every_minute_inclusive_not_production_two_minute_sampling",
|
||||
"cross_midnight": "date_aware_candidates_distance_and_matrix_grouped_by_candidate_date",
|
||||
"evaluator_sha256": file_sha256(Path(__file__)),
|
||||
"replay_revision": "candidate_date_grouped_v2",
|
||||
"supersedes": "initial_sweep_invalidated_cross_midnight_transition_anchor",
|
||||
"is_blind_evaluation": False, "truth_hidden_from_ranker": True,
|
||||
"results_previously_seen": True, "must_not_use_for_tuning": True,
|
||||
},
|
||||
"excluded_cases": invalid, "case_count": validation["valid_public_aa_cases"],
|
||||
"trial_count": len(trials), "trials": trials,
|
||||
"summary": summarize(trials, radii, offsets),
|
||||
"widening_geometry": {
|
||||
"scored_radii": list(radii),
|
||||
"radius_120_scored": 120 in radii,
|
||||
"truth_in_window_condition": "abs(reported_offset_minutes) <= radius_minutes",
|
||||
"radius_15_first_integer_minute_outside": 16,
|
||||
"all_tested_offsets_within_radius_30_60_120": max(map(abs, offsets)) <= 30,
|
||||
"boundary": "Geometry rescues candidate inclusion only, not ranking or delivery coverage; no population frequency without a reported-offset distribution.",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--json", action="store_true")
|
||||
args = parser.parse_args()
|
||||
print(json.dumps(run(), ensure_ascii=False, indent=2))
|
||||
+153
@@ -0,0 +1,153 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Freeze and replay the exposed v3 corpus; never claim a fresh blind holdout."""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import hashlib
|
||||
import json
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.active_rectification_event_engine import AYANAMSA, NODE_MODE
|
||||
from scripts.minute_rectification_blind_eval import (
|
||||
_candidate_moments, _clock_distance, _opaque_winner, _request,
|
||||
implementation_sha256, summarize_trials,
|
||||
)
|
||||
from scripts.minute_rectification_fact_blind_eval_v4 import _would_confirm
|
||||
from scripts.minute_rectification_fact_ranker_v4 import (
|
||||
ALGORITHM_VERSION, rank_fact_rows, score_fact_ranker_v4,
|
||||
)
|
||||
from scripts.minute_rectification_feature_facts_v4 import build_feature_fact_rows
|
||||
from scripts.minute_rectification_holdout_validator import validate
|
||||
|
||||
DATASET = ROOT / "references/real_case_calibration/minute_rectification_holdout_v3.json"
|
||||
FREEZE = ROOT / "docs/research/sealed_holdout_rerun_2026_09_20.freeze.json"
|
||||
REPORT = ROOT / "docs/research/sealed_holdout_rerun_2026_09_20.json"
|
||||
|
||||
|
||||
def file_sha256(path: Path) -> str:
|
||||
return hashlib.sha256(path.read_bytes()).hexdigest()
|
||||
|
||||
|
||||
def opaque_order(benchmark_id: str, case_id: str, rows: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
"""Extend the existing opaque winner to a total, truth-independent ranking."""
|
||||
return sorted(rows, key=lambda row: (
|
||||
-row["score"],
|
||||
hashlib.sha256(f"{benchmark_id}:{case_id}:{row['time']}".encode()).hexdigest(),
|
||||
))
|
||||
|
||||
|
||||
def freeze_record(dataset: Path = DATASET) -> dict[str, Any]:
|
||||
manifest = json.loads(dataset.read_text(encoding="utf-8"))
|
||||
files = manifest["frozen_scoring"]["files"]
|
||||
return {
|
||||
"record_version": "exposed-v3-fixed-protocol-rerun-v1",
|
||||
"frozen_at_utc": datetime.now(timezone.utc).isoformat(),
|
||||
"dataset_path": dataset.relative_to(ROOT).as_posix(),
|
||||
"dataset_sha256": file_sha256(dataset),
|
||||
"algorithm_version": ALGORITHM_VERSION,
|
||||
"implementation_sha256": implementation_sha256(files),
|
||||
"files": files,
|
||||
"historical_frozen_sha256": manifest["frozen_scoring"]["implementation_sha256"],
|
||||
"evaluator_sha256": file_sha256(Path(__file__)),
|
||||
"ayanamsa": AYANAMSA,
|
||||
"node_mode": NODE_MODE,
|
||||
"candidate_radius_minutes": sorted({case["candidate_radius_minutes"] for case in manifest["cases"]}),
|
||||
"minute_step": 1,
|
||||
"release_metrics": manifest["release_metrics"],
|
||||
"results_previously_seen": True,
|
||||
"official_valid_independent_blind": False,
|
||||
"must_not_use_for_tuning": True,
|
||||
"tie_breaker": manifest["frozen_scoring"]["tie_breaker"],
|
||||
"metric_rank_definition": "competition_rank_1_plus_strictly_higher_scores_legacy_protocol",
|
||||
"extra_metric_rank_definition": "score_desc_then_existing_opaque_sha256_total_order",
|
||||
}
|
||||
|
||||
|
||||
def run(freeze_path: Path = FREEZE, dataset: Path = DATASET) -> dict[str, Any]:
|
||||
frozen = json.loads(freeze_path.read_text(encoding="utf-8"))
|
||||
actual = freeze_record(dataset)
|
||||
for key in actual:
|
||||
if key != "frozen_at_utc" and actual[key] != frozen.get(key):
|
||||
raise ValueError(f"frozen_record_mismatch:{key}")
|
||||
manifest = json.loads(dataset.read_text(encoding="utf-8"))
|
||||
validation = validate(dataset)
|
||||
invalid = validation["invalid_cases"]
|
||||
trials = []
|
||||
for index, case in enumerate(manifest["cases"], 1):
|
||||
if case["case_id"] in invalid:
|
||||
continue
|
||||
request = _request(case, case["events"])
|
||||
candidates = _candidate_moments(case)
|
||||
facts = build_feature_fact_rows(request, candidates=candidates)
|
||||
rows, _ = rank_fact_rows(facts, request["events"])
|
||||
result = score_fact_ranker_v4(facts, request["events"])
|
||||
predicted = _opaque_winner(manifest["benchmark_id"], case["case_id"], rows)
|
||||
ordered = opaque_order(manifest["benchmark_id"], case["case_id"], rows)
|
||||
sparse_request = _request(case, case["events"][:1])
|
||||
sparse_facts = build_feature_fact_rows(sparse_request, candidates=candidates)
|
||||
sparse_result = score_fact_ranker_v4(sparse_facts, sparse_request["events"])
|
||||
# Truth is revealed only after both full and sparse ranking/decisions.
|
||||
truth = case["birth"]["time"]
|
||||
truth_score = next(row["score"] for row in rows if row["time"] == truth)
|
||||
would_confirm = _would_confirm(result)
|
||||
trials.append({
|
||||
"case_ordinal": index,
|
||||
"candidate_count": len(rows),
|
||||
"event_count": len(request["events"]),
|
||||
"true_rank": 1 + sum(row["score"] > truth_score for row in rows),
|
||||
"opaque_true_rank": next(i for i, row in enumerate(ordered, 1) if row["time"] == truth),
|
||||
"minute_error": _clock_distance(predicted, truth),
|
||||
"would_confirm": would_confirm,
|
||||
"false_confirmation": would_confirm and predicted != truth,
|
||||
"insufficient_evidence_rejected": not _would_confirm(sparse_result),
|
||||
"full_trial_reasons": result["reasons"],
|
||||
"sparse_trial_reasons": sparse_result["reasons"],
|
||||
})
|
||||
aggregate = summarize_trials(trials, manifest["release_metrics"])
|
||||
count = len(trials)
|
||||
return {
|
||||
"scope": "fixed_protocol_previously_exposed_v3_rerun",
|
||||
"evaluated_on": datetime.now(timezone.utc).date().isoformat(),
|
||||
"frozen_record": frozen,
|
||||
"implementation_hash_matches_at_replay": True,
|
||||
"dataset_hash_matches_at_replay": True,
|
||||
"source_audit_status": manifest["source_audit_status"],
|
||||
"validation_status": validation["status"],
|
||||
"valid_public_aa_cases": validation["valid_public_aa_cases"],
|
||||
"excluded_cases": invalid,
|
||||
"trial_count": count,
|
||||
"trials": trials,
|
||||
**aggregate,
|
||||
"opaque_exact_top_1_rate": sum(row["opaque_true_rank"] == 1 for row in trials) / count if count else None,
|
||||
"opaque_exact_top_3_rate": sum(row["opaque_true_rank"] <= 3 for row in trials) / count if count else None,
|
||||
"official_valid_independent_blind": False,
|
||||
"official_blind_trial_count": 0,
|
||||
"is_blind_evaluation": False,
|
||||
"truth_hidden_from_ranker": True,
|
||||
"results_previously_seen": True,
|
||||
"verified_minute_claim_allowed": False,
|
||||
"status": "blocked_independent_blind_evidence",
|
||||
"boundary": "Three-event low-information protocol, not a mathematical accuracy lower bound and not representative of real sessions. Historical v3/v4 exposure cannot be undone by freezing today's scorer. No tuning or release claims.",
|
||||
}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--freeze", action="store_true", help="Create a new record before replay; never overwrite an existing record")
|
||||
parser.add_argument("--freeze-path", type=Path, default=FREEZE)
|
||||
parser.add_argument("--json", action="store_true")
|
||||
args = parser.parse_args()
|
||||
if args.freeze:
|
||||
with args.freeze_path.open("x", encoding="utf-8", newline="\n") as handle:
|
||||
json.dump(freeze_record(), handle, ensure_ascii=False, indent=2)
|
||||
handle.write("\n")
|
||||
print(args.freeze_path.relative_to(ROOT).as_posix())
|
||||
else:
|
||||
print(json.dumps(run(args.freeze_path), ensure_ascii=False, indent=2))
|
||||
@@ -2,6 +2,12 @@
|
||||
|
||||
Purpose: read this file before substantial project work. It exists to stop repeat mistakes caused by multiple Codex windows, WorkBuddy mirrors, local drafts, backup folders, and partial cloud-git visibility.
|
||||
|
||||
## 2026-09-20 · 跨午夜生产修复的身份与本机预检
|
||||
|
||||
- 任务书的“V9 `engine_version` 全仓只写不比”不完整:成功回执来自实际 `algorithmVersion`,已有 score cache 通过版本接口/环境覆盖比较算法身份。只 bump 前端默认串不能保证成功结果与旧缓存可区分;本单同步 bump 既有后端算法身份,不新增历史会话打开门,不重标旧缓存。
|
||||
- 新工作树预检仍为远端 verified、适配器/碎片检查成功、focused 镜像路径断言失败;快速门基线与本轮均在 source inventory 因缺 `mcp` 失败。BUG-089 的 `mcp>=1.0,<2` 约束两处仍在,本机问题是未安装,不是再次升到 MCP 2。
|
||||
- 广域校正 Python 基线已有 4 个失败,必须逐条对照,不借机修改历史冻结值或削弱旧断言。验收记录见 `docs/tasks/PROGRESS-rectification-cross-midnight-20260920.md`。
|
||||
|
||||
## 2026-09-20 · 校正验证补缺的本机验收复现
|
||||
|
||||
- 基线与实现工作树均使用 Python 3.11.7(无项目 `.venv`;`python3` launcher 退出 49)。开工预检远端 verified,但同一碎片镜像路径断言失败;不得把同步成功写成预检通过。
|
||||
|
||||
@@ -1,4 +1,6 @@
|
||||
# 申报偏差敏感性评测(2026-09-20)
|
||||
# 申报偏差敏感性评测(2026-09-20,BUG-981 后原生单矩阵重跑)
|
||||
|
||||
> 当前正式结果:`reported_offset_cross_midnight_2026_09_20.json`,冻结:`reported_offset_cross_midnight_2026_09_20.final.freeze.json`。旧 `reported_offset_2026_09_20.json` 原字节保留,不再充当当前身份报告。修订前本 Markdown、旧 JSON/冻结/契约/评测器逐字节保留在 `history/rectification_pre_cross_midnight_2026_09_20/`,由 manifest 固定哈希。
|
||||
|
||||
## 结论与边界
|
||||
|
||||
@@ -6,7 +8,7 @@
|
||||
|
||||
本轮完整评测预先固定 15 个偏移 × 3 档半径 × 20 例,共 900 组合。修正版全部运行完成,validator 排除 0 例;实测结果以下列最终 JSON 为准,不能用中途旧版本初扫代替。
|
||||
|
||||
**区间覆盖不等于分钟命中,模拟偏差敏感性不等于真实用户准确率。** 不运行六题真值方向回放,不以用户认可作真值,不调参、不改生产打分、不打开确认门。
|
||||
**区间覆盖不等于分钟命中,模拟偏差敏感性不等于真实用户准确率。** 不运行六题真值方向回放,不以用户认可作真值,不调参、不换案例、不打开确认门。本轮生产代理单独修正候选日期与版本标记;本研究使用修复后的原生路径,不增加评分规则。
|
||||
|
||||
## 口径
|
||||
|
||||
@@ -22,12 +24,12 @@
|
||||
| 额外身份 | JSON 的 `production_scoring_sha256` 与 `research_implementation_sha256` 分别绑定扩展评分文件和研究适配文件;显式文件清单不是完整传递依赖锁 |
|
||||
| 头名与排名 | 分数降序,再用既有 `_opaque_winner` 同源 SHA-256 规则排同分;不选择最接近真值的候选 |
|
||||
| 交付区间 | 初始未答题:signature clusters 中落后头名不足 8 分的仍有效簇外包范围;无新回答、无淘汰,不是完整问答链交付效果 |
|
||||
| 跨午夜 | 候选、误差及区间使用完整日期;离线矩阵按候选日期分组,合并后全窗排序/聚类 |
|
||||
| 跨午夜 | 候选、误差及区间使用完整日期;修复后的原生单矩阵内按候选日期计分,全窗排序/聚类;不再运行日期分组适配 |
|
||||
| 独立性 | `is_blind_evaluation=false`,已曝光集;评分/区间定稿后才揭示标签,仅证明程序标签隔离 |
|
||||
|
||||
## 实测表
|
||||
|
||||
每格依次为 **真值在窗比例 / 哈希头名命中 / 初始交付区间覆盖**,分母均为 20 例。列为申报偏移分钟,行为搜索半径。统一口径:v3、raman/mean、1 分钟步长;历史打分身份 `b15d9ea15227cd58`,扩展原生评分身份 `115c3fbcffdaff49`,研究适配身份 `8d0c613dc0900adb`。
|
||||
每格依次为 **真值在窗比例 / 哈希头名命中 / 初始交付区间覆盖**,分母均为 20 例。列为申报偏移分钟,行为搜索半径。统一口径:v3、raman/mean、1 分钟步长;历史 12 文件打分身份 `b15d9ea15227cd58`,扩展原生评分身份 `7fffd1db612af1f2`,研究评测身份 `2931d4661bf53dde`。
|
||||
|
||||
| 半径 | -30 | -20 | -15 | -10 | -8 | -5 | -3 | 0 | +3 | +5 | +8 | +10 | +15 | +20 | +30 |
|
||||
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
|
||||
@@ -46,11 +48,25 @@
|
||||
- 半径 ±30 自第 31 个整数分钟偏移出窗,±60 自第 61 个、±120 自第 121 个。放宽不是无条件保证,更不代表排名命中或交付覆盖同样恢复。
|
||||
- ±120 未跑评分,不报告其头名命中率或交付覆盖率。真实用户申报偏差分布未知,不估算现实用户中有多少人能被放宽救回。
|
||||
|
||||
## 初扫作废与独立复核
|
||||
## 历史初扫、日期分组版与本轮原生重跑
|
||||
|
||||
首版正确枚举了跨日候选,但矩阵内 transition proximity 仍共用窗口起日,独立审查发现跨日候选得分偏移。因此旧初扫作废,修正版标记 `replay_revision=candidate_date_grouped_v2`,并登记 `supersedes=initial_sweep_invalidated_cross_midnight_transition_anchor`。
|
||||
历史初扫正确枚举跨日候选,但矩阵内 transition proximity 共用窗口起日,因此初扫作废。其后已修成 `candidate_date_grouped_v2` 离线日期分组版,旧同名 JSON 是这个修正版,不是错误初扫。修正版旧身份:扩展原生评分 `115c3fbcffdaff49`,研究适配 `8d0c613dc0900adb`,旧数值不能被误说成必然错误。
|
||||
|
||||
修复仅在本次新脚本 `score_window()` 做按日期分组适配,未改冻结打分文件。实引擎回归验证跨日窗口的分组计算与逐候选独立重算相同;聚类和交付范围在合并后全窗运行。**生产日期处理未在本单修复,不能称为生产端到端回放。**
|
||||
本轮 `replay_revision=native_candidate_date_v3`:生产 helper 已修,`score_window()` 改为一次原生矩阵调用;回归同时锁定调用次数为 1、与逐候选独立真实引擎重算相等。不再保留运行中的日期分组替代路径。旧数字因实现/路径身份更新不再代表当前报告,是否变化只由完整比对决定,不能预设数字必变。
|
||||
|
||||
**这里只保证 matrix scoring path 是修复后的原生单次调用,不是完整生产端到端回放。** `shifted_window()` 使用完整日期,`delivery_moments()` 按完整日期定义研究交付外包区间;它们仍是离线协议。协调独立审查另报 BUG-982/983:生产跨午夜簇/span 的钟点排序与前端早凌晨申报窗日期锚点仍有独立缺陷。本研究不覆盖这些路径,不能凭区间覆盖数字声称线上所有跨日问题已修。
|
||||
|
||||
## 本轮预冻结身份
|
||||
|
||||
| 项目 | 值 |
|
||||
| --- | --- |
|
||||
| 冻结时间(先于评分) | `2026-09-20T04:56:55.260470+00:00` |
|
||||
| 原 12 文件 | `b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18` |
|
||||
| 扩展生产 18 文件 | `7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51` |
|
||||
| 研究/校验 7 文件 | `2931d4661bf53dde5c57f6492de65be4c00877293b7277353279a05eda5af215` |
|
||||
| 冻结文件 SHA-256 | `1f281b6c3e9b5a7dc2de91fb2ae1fd94da38f9b8a7bb1f930f670e9a39865074` |
|
||||
|
||||
原 12 文件未变,但实际变化的 `scoring_service.py` 与 `dasha_transition_proximity.py` 都进入扩展身份。冻结亦绑定数据集、所有组合/参数、每文件字节哈希、归档 manifest 和评测器;评分前后核对,漂移即拒绝,不能事后刷新冻结。准备阶段曾因旧 JSON 覆盖被安全工具拒绝而迁移新报告路径,准备产物及弃用原因由历史目录 `preparation_manifest.json` 保留;最终路径评测器另行预冻结,没有追认旧冻结。
|
||||
|
||||
## 可复算产物与验收
|
||||
|
||||
@@ -59,6 +75,19 @@ python scripts/research/reported_offset_sweep.py --json
|
||||
python -m pytest tests/test_reported_offset_research.py tests/test_rectification_validation_integrity_gate.py -q
|
||||
```
|
||||
|
||||
`reported_offset_2026_09_20.json` 保留逐例四项必要输出:真值是否在窗、真值排名、头名分钟误差、交付区间覆盖,以及每个格子的分母、全部参数和实现哈希。仅保存公开集序号,不保存出生分钟、坐标或事件正文。
|
||||
`reported_offset_cross_midnight_2026_09_20.json` 保留逐例四项必要输出:真值是否在窗、真值排名、头名分钟误差、交付区间覆盖,以及每个格子的分母、全部参数和实现哈希。仅保存公开集序号,不保存出生分钟、坐标或事件正文。
|
||||
|
||||
**最终全量 900/900 组合完成,用时 399.902 秒;排除 0 例。与旧日期分组版逐例全部字段相同,改变 0/900,45/45 汇总格一致。** JSON 的 `historical_comparison.trials` 保留所有 before/after 与 changed_fields(含未变化行)。这符合旧最终版已经做日期分组修正的事实,不表示生产旧 helper 没有 Bug。本轮没有调参、删失败样本、换半径或把已曝光集包装成新独立盲测。
|
||||
|
||||
### 既有测试断言变更
|
||||
|
||||
| 原值 / 原断言 | 新值 / 新断言 | 原因 |
|
||||
| --- | --- | --- |
|
||||
| 跨日实引擎 `grouped == expected` 且生产 `old_rows != expected` | 原生单矩阵 `native == expected` 且 `native_rows == expected`,并锁一次矩阵调用 | 原测试第二条锁住的是已知生产缺陷;修复后必须相等。独立逐候选真实引擎 oracle 原样保留并强化,不是删掉差异验证 |
|
||||
| `replay_revision=candidate_date_grouped_v2` | `native_candidate_date_v3` | 原生路径已修,不再离线分组 |
|
||||
| 记录测试读取旧 JSON | 读取脚本新 `REPORT` 路径,全部数据集/评测器/生产/研究哈希断言保留 | 旧产物原字节留档,新报告独立创建 |
|
||||
| 仅运行后记录参数身份 | 预冻结完整身份,评分前后核对;新增漂移提前拒绝、时间先后、完整旧新对比断言 | 不能事后伪造冻结或只比较有利样本 |
|
||||
|
||||
本轮定向验证:`test_reported_offset_research.py`、`test_sealed_holdout_contract_freshness.py`、`test_rectification_validation_integrity_gate.py`、`test_rectification_confirmation_and.py` 合计 **52/52 通过**(含桥接重复收集)。独立逐行比较新旧 900 trials、45 summary 全相同;当前报告 SHA-256 为 `26c3d4a6e5e96ffaacdd00c64b5f3ed08fb9eacb18ef6fffc4b01c15159207b4`。旧产物与准备产物 manifest 的全部字节校验通过。预检因既有 Windows 镜像路径断言失败(23 通过 / 1 失败),不是预检通过。完整环境/广域验收由 `../tasks/PROGRESS-rectification-cross-midnight-20260920.md` 记录。
|
||||
|
||||
排名不能与 T2 的竞争排名 top-1/top-3 混比:T2 按旧协议允许并列多个分钟同时算第一;本表把同分完全按哈希打破。最终测试数字与已知基线失败见 [进度](../tasks/PROGRESS-rectification-validation-20260920.md)。真实分布采集与新未曝光样本见 [v5 协议](sealed_holdout_v5_protocol_2026_09_20.md)。
|
||||
|
||||
@@ -0,0 +1,108 @@
|
||||
{
|
||||
"record_version": "reported-offset-native-candidate-date-v3",
|
||||
"frozen_at_utc": "2026-09-20T04:56:55.260470+00:00",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"radii_minutes": [
|
||||
15,
|
||||
30,
|
||||
60
|
||||
],
|
||||
"offsets_minutes": [
|
||||
-30,
|
||||
-20,
|
||||
-15,
|
||||
-10,
|
||||
-8,
|
||||
-5,
|
||||
-3,
|
||||
0,
|
||||
3,
|
||||
5,
|
||||
8,
|
||||
10,
|
||||
15,
|
||||
20,
|
||||
30
|
||||
],
|
||||
"minute_step": 1,
|
||||
"dataset": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"dataset_benchmark_id": "minute_rectification_fact_ranker_v4_holdout_v3",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"implementation_sha256_prefix": "b15d9ea15227cd58",
|
||||
"historical_artifacts_manifest_sha256": "6102a26a840be207b5858b3a4c0509274468ae9071e86d308cbbac7371d6864e",
|
||||
"production_scoring_files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/rectification/candidate_contrast.py",
|
||||
"scripts/rectification/case_holdout.py",
|
||||
"scripts/rectification/contracts.py",
|
||||
"scripts/rectification/dasha_transition_proximity.py",
|
||||
"scripts/rectification/event_probes.py",
|
||||
"scripts/rectification/scoring_service.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"production_scoring_sha256": "7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51",
|
||||
"research_files": [
|
||||
"scripts/research/reported_offset_sweep.py",
|
||||
"scripts/research/cluster_width_lib.py",
|
||||
"scripts/research/probe_supply_after_six.py",
|
||||
"scripts/research/sealed_holdout_rerun.py",
|
||||
"scripts/minute_rectification_blind_eval.py",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py",
|
||||
"scripts/minute_rectification_holdout_validator.py"
|
||||
],
|
||||
"research_implementation_sha256": "2931d4661bf53dde5c57f6492de65be4c00877293b7277353279a05eda5af215",
|
||||
"file_sha256": {
|
||||
"scripts/active_rectification_event_engine.py": "edbe93ef3c8bd65f986139dba6b95b96d9b0fa9920341fbe677cd6f53e4db84f",
|
||||
"scripts/active_rectification_events.py": "03e333ccf1103caf685dfe50fe890ab1793e945037085f50ce834a4a3b232222",
|
||||
"scripts/ashtakavarga.py": "cc32776674cf60713634a360a6f09774b3af0ec5d0f650687f52b79597b0dfbe",
|
||||
"scripts/dasha_analyzer.py": "5e53b1b4d4c7414c08dd0e7716f120a63c70ec3428ec93547f64d55c580da0e1",
|
||||
"scripts/divisional_charts_extended.py": "73bd50a83ac444830940184d0988cea621855f91db76a3f1a232014b992eacea",
|
||||
"scripts/domain_calculation_service.py": "9ee955874eea806c7b56f4246d90009c02ee8318b8f3dfe1d7132b1c0b4262e1",
|
||||
"scripts/jaimini.py": "b98bf965603dd2e3cc871fe40f3c8184731a701875ca50fef3ca68e9682eda81",
|
||||
"scripts/minute_rectification_blind_eval.py": "07aa30d8bad3278669495d5fea8be937aaf6367c8af980ff4f2fe9aa84ccc6ff",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py": "bdf460e3467208e8b87345c5beb2128f3e22de7eeffd68f72b281b3ad9affc92",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py": "f1a0d8d2482cd754d2e7f253a9606b0373519435a2bf58cfc3753e2b9ae0fb76",
|
||||
"scripts/minute_rectification_feature_facts_v4.py": "3d899a2cdd4825b8e32062443d487edcd2173a4ea55a36eb9f3abe587c0faf0f",
|
||||
"scripts/minute_rectification_holdout_validator.py": "b60b1c752e58672beaee8b36eb3b2348169726cda070ec969db0f13e7ed715af",
|
||||
"scripts/narayana_dasha.py": "7ff2c3238cd113b11b815e14b41967bd8e486ff62e36650e55613b9c79d1bfe3",
|
||||
"scripts/rectification/candidate_contrast.py": "10a360586359e72846593ef6d7c15a2cea79358245a34c288851406fe9e47ca8",
|
||||
"scripts/rectification/case_holdout.py": "fe0e699b617d9e216acdfeff87af4221c1d6f4c101bf33e1dbf14e2bcd75a8b5",
|
||||
"scripts/rectification/contracts.py": "fdd1a47b2e3dac8579e870a282ed74c68518b39f6c0576b27df766a4b8021143",
|
||||
"scripts/rectification/dasha_transition_proximity.py": "0afba494aa12899621173c3cc941c6c1d0be9cec0bdf4c4a1c9bcd004ab22de8",
|
||||
"scripts/rectification/event_probes.py": "45478ee3fdc36dc36e1a602bee5dd0627a5464275e370e64adb5826fe2aa58f5",
|
||||
"scripts/rectification/scoring_service.py": "0adf5700bb22dd8c4c3147027251210e21ec66a2c9f9a91217edba3ad9e5d1f5",
|
||||
"scripts/research/cluster_width_lib.py": "1d8f2ffa7795d62069deebb2de8e16749317037d4ed3f24a2c0092676e5cfcc6",
|
||||
"scripts/research/probe_supply_after_six.py": "098a5b4398d6f8d996917d809c2fc922ad64031d3b70660f8ef7c876aba92438",
|
||||
"scripts/research/reported_offset_sweep.py": "f81c681fd0fe257e90a761dfab370f623e5f8a881b62cc6b1c596cbd58800f8f",
|
||||
"scripts/research/sealed_holdout_rerun.py": "34a6724530fbecdecff9e5e07fa64980b8e8aff1fac2e000f4268debe1ee96e1",
|
||||
"scripts/shadbala.py": "912e0e6d169c2172aab85f71e4347e3e39193020777b43825b775d906400956b",
|
||||
"scripts/varga.py": "4331de5a25ea08729af91aae863943c4223183937f419d52dc63fd6fe7f27be6"
|
||||
},
|
||||
"hash_scope": "explicit_identity_file_sets_not_a_transitive_dependency_lock",
|
||||
"scorer": "native_event_contribution_matrix_not_shadow_fact_ranker",
|
||||
"delivery": "initial_signature_clusters_peak_gap_lt_8_envelope_no_answers_no_elimination",
|
||||
"rank": "score_desc_then_sha256(benchmark_id:case_id:candidate_time)",
|
||||
"grid": "every_minute_inclusive_not_production_two_minute_sampling",
|
||||
"cross_midnight": "date_aware_candidates_distance_and_native_single_matrix",
|
||||
"evaluator_sha256": "f81c681fd0fe257e90a761dfab370f623e5f8a881b62cc6b1c596cbd58800f8f",
|
||||
"replay_revision": "native_candidate_date_v3",
|
||||
"supersedes": "candidate_date_grouped_v2_identity_and_path_not_assumed_numerically_wrong",
|
||||
"is_blind_evaluation": false,
|
||||
"truth_hidden_from_ranker": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"official_blind_trial_count": 0,
|
||||
"results_previously_seen": true,
|
||||
"must_not_use_for_tuning": true
|
||||
}
|
||||
@@ -0,0 +1,108 @@
|
||||
{
|
||||
"record_version": "reported-offset-native-candidate-date-v3",
|
||||
"frozen_at_utc": "2026-09-20T04:50:20.160896+00:00",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"radii_minutes": [
|
||||
15,
|
||||
30,
|
||||
60
|
||||
],
|
||||
"offsets_minutes": [
|
||||
-30,
|
||||
-20,
|
||||
-15,
|
||||
-10,
|
||||
-8,
|
||||
-5,
|
||||
-3,
|
||||
0,
|
||||
3,
|
||||
5,
|
||||
8,
|
||||
10,
|
||||
15,
|
||||
20,
|
||||
30
|
||||
],
|
||||
"minute_step": 1,
|
||||
"dataset": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"dataset_benchmark_id": "minute_rectification_fact_ranker_v4_holdout_v3",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"implementation_sha256_prefix": "b15d9ea15227cd58",
|
||||
"historical_artifacts_manifest_sha256": "6102a26a840be207b5858b3a4c0509274468ae9071e86d308cbbac7371d6864e",
|
||||
"production_scoring_files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/rectification/candidate_contrast.py",
|
||||
"scripts/rectification/case_holdout.py",
|
||||
"scripts/rectification/contracts.py",
|
||||
"scripts/rectification/dasha_transition_proximity.py",
|
||||
"scripts/rectification/event_probes.py",
|
||||
"scripts/rectification/scoring_service.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"production_scoring_sha256": "7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51",
|
||||
"research_files": [
|
||||
"scripts/research/reported_offset_sweep.py",
|
||||
"scripts/research/cluster_width_lib.py",
|
||||
"scripts/research/probe_supply_after_six.py",
|
||||
"scripts/research/sealed_holdout_rerun.py",
|
||||
"scripts/minute_rectification_blind_eval.py",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py",
|
||||
"scripts/minute_rectification_holdout_validator.py"
|
||||
],
|
||||
"research_implementation_sha256": "61dbb47eb46f9a01c78096677ac6238e3a65e9674ce4f2514b18e1ca50c0b954",
|
||||
"file_sha256": {
|
||||
"scripts/active_rectification_event_engine.py": "edbe93ef3c8bd65f986139dba6b95b96d9b0fa9920341fbe677cd6f53e4db84f",
|
||||
"scripts/active_rectification_events.py": "03e333ccf1103caf685dfe50fe890ab1793e945037085f50ce834a4a3b232222",
|
||||
"scripts/ashtakavarga.py": "cc32776674cf60713634a360a6f09774b3af0ec5d0f650687f52b79597b0dfbe",
|
||||
"scripts/dasha_analyzer.py": "5e53b1b4d4c7414c08dd0e7716f120a63c70ec3428ec93547f64d55c580da0e1",
|
||||
"scripts/divisional_charts_extended.py": "73bd50a83ac444830940184d0988cea621855f91db76a3f1a232014b992eacea",
|
||||
"scripts/domain_calculation_service.py": "9ee955874eea806c7b56f4246d90009c02ee8318b8f3dfe1d7132b1c0b4262e1",
|
||||
"scripts/jaimini.py": "b98bf965603dd2e3cc871fe40f3c8184731a701875ca50fef3ca68e9682eda81",
|
||||
"scripts/minute_rectification_blind_eval.py": "07aa30d8bad3278669495d5fea8be937aaf6367c8af980ff4f2fe9aa84ccc6ff",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py": "bdf460e3467208e8b87345c5beb2128f3e22de7eeffd68f72b281b3ad9affc92",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py": "f1a0d8d2482cd754d2e7f253a9606b0373519435a2bf58cfc3753e2b9ae0fb76",
|
||||
"scripts/minute_rectification_feature_facts_v4.py": "3d899a2cdd4825b8e32062443d487edcd2173a4ea55a36eb9f3abe587c0faf0f",
|
||||
"scripts/minute_rectification_holdout_validator.py": "b60b1c752e58672beaee8b36eb3b2348169726cda070ec969db0f13e7ed715af",
|
||||
"scripts/narayana_dasha.py": "7ff2c3238cd113b11b815e14b41967bd8e486ff62e36650e55613b9c79d1bfe3",
|
||||
"scripts/rectification/candidate_contrast.py": "10a360586359e72846593ef6d7c15a2cea79358245a34c288851406fe9e47ca8",
|
||||
"scripts/rectification/case_holdout.py": "fe0e699b617d9e216acdfeff87af4221c1d6f4c101bf33e1dbf14e2bcd75a8b5",
|
||||
"scripts/rectification/contracts.py": "fdd1a47b2e3dac8579e870a282ed74c68518b39f6c0576b27df766a4b8021143",
|
||||
"scripts/rectification/dasha_transition_proximity.py": "0afba494aa12899621173c3cc941c6c1d0be9cec0bdf4c4a1c9bcd004ab22de8",
|
||||
"scripts/rectification/event_probes.py": "45478ee3fdc36dc36e1a602bee5dd0627a5464275e370e64adb5826fe2aa58f5",
|
||||
"scripts/rectification/scoring_service.py": "0adf5700bb22dd8c4c3147027251210e21ec66a2c9f9a91217edba3ad9e5d1f5",
|
||||
"scripts/research/cluster_width_lib.py": "1d8f2ffa7795d62069deebb2de8e16749317037d4ed3f24a2c0092676e5cfcc6",
|
||||
"scripts/research/probe_supply_after_six.py": "098a5b4398d6f8d996917d809c2fc922ad64031d3b70660f8ef7c876aba92438",
|
||||
"scripts/research/reported_offset_sweep.py": "cdcd7b583d9ecee8a313c25cac795923da2f8af45875d572f0a3efbc25f5d071",
|
||||
"scripts/research/sealed_holdout_rerun.py": "33c294b21db0631a68068dcf135b70db7aae6a1a3e66414bb72959a448883465",
|
||||
"scripts/shadbala.py": "912e0e6d169c2172aab85f71e4347e3e39193020777b43825b775d906400956b",
|
||||
"scripts/varga.py": "4331de5a25ea08729af91aae863943c4223183937f419d52dc63fd6fe7f27be6"
|
||||
},
|
||||
"hash_scope": "explicit_identity_file_sets_not_a_transitive_dependency_lock",
|
||||
"scorer": "native_event_contribution_matrix_not_shadow_fact_ranker",
|
||||
"delivery": "initial_signature_clusters_peak_gap_lt_8_envelope_no_answers_no_elimination",
|
||||
"rank": "score_desc_then_sha256(benchmark_id:case_id:candidate_time)",
|
||||
"grid": "every_minute_inclusive_not_production_two_minute_sampling",
|
||||
"cross_midnight": "date_aware_candidates_distance_and_native_single_matrix",
|
||||
"evaluator_sha256": "cdcd7b583d9ecee8a313c25cac795923da2f8af45875d572f0a3efbc25f5d071",
|
||||
"replay_revision": "native_candidate_date_v3",
|
||||
"supersedes": "candidate_date_grouped_v2_identity_and_path_not_assumed_numerically_wrong",
|
||||
"is_blind_evaluation": false,
|
||||
"truth_hidden_from_ranker": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"official_blind_trial_count": 0,
|
||||
"results_previously_seen": true,
|
||||
"must_not_use_for_tuning": true
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -1,4 +1,6 @@
|
||||
# 当前冻结实现的 v3 固定口径重跑(2026-09-20)
|
||||
# 当前冻结实现的 v3 固定口径重跑(2026-09-20,BUG-981 后重新冻结)
|
||||
|
||||
> 当前正式结果已迁到 `sealed_holdout_rerun_cross_midnight_2026_09_20.json`;旧同名 JSON 与旧 `.freeze.json` 原字节保留,不代表本轮身份。本文修订前的 Markdown、JSON、冻结记录、契约和评测器均逐字节归档在 `history/rectification_pre_cross_midnight_2026_09_20/`,由 `manifest.json` 校验。历史冻结不刷新、不追认。
|
||||
|
||||
## 结论
|
||||
|
||||
@@ -21,9 +23,18 @@
|
||||
| 排名标签隔离 | 排序和稀疏证据判定完成后才读取真值标签 |
|
||||
| 独立官方盲测 | `official_valid_independent_blind=false`;有效官方盲测试次仍为 0 |
|
||||
|
||||
冻结记录:`sealed_holdout_rerun_2026_09_20.freeze.json`。先以独占创建模式记录 12 个文件、数据集与评测器字节哈希,再执行评分;评分前验证全部身份,任何漂移即拒绝。历史 v3 `frozen_scoring` 原值 `f41c298dd6cdcebe…` 未改。
|
||||
本轮冻结记录:`sealed_holdout_rerun_cross_midnight_2026_09_20.final.freeze.json`。在 **2026-09-20T04:56:54.914584+00:00** 以独占创建模式预冻结,然后重跑。12 文件身份继续为 `b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18`;实际修复在这 12 文件之外,故增加以下身份,不能拿旧 12 文件哈希不变冒充生产未变:
|
||||
|
||||
脱敏逐例结果:`sealed_holdout_rerun_2026_09_20.json`,只含公开集序号、排名、误差、门禁原因,不含出生分钟、坐标或事件正文。复算:
|
||||
| 身份 | SHA-256 |
|
||||
| --- | --- |
|
||||
| 扩展生产评分(18 文件,含 `scoring_service.py` 与 `dasha_transition_proximity.py`) | `7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51` |
|
||||
| 研究/校验(7 文件) | `2931d4661bf53dde5c57f6492de65be4c00877293b7277353279a05eda5af215` |
|
||||
| 本轮冻结文件 | `b74f70d21011f98b78f141535b1bfa49b20701418812a3ad97f9350cc793c67d` |
|
||||
| 旧产物归档 manifest | `6102a26a840be207b5858b3a4c0509274468ae9071e86d308cbbac7371d6864e` |
|
||||
|
||||
每个文件的字节哈希与有序文件集身份均在冻结记录内;评分前后都核对数据集、评测器与扩展实现。文件集为显式范围,不冒称完整传递依赖或环境锁。历史 v3 `frozen_scoring` 原值 `f41c298dd6cdcebe…` 未改。
|
||||
|
||||
脱敏逐例结果:`sealed_holdout_rerun_cross_midnight_2026_09_20.json`,只含公开集序号、排名、误差、门禁原因,不含出生分钟、坐标或事件正文。复算:
|
||||
|
||||
```bash
|
||||
python scripts/research/sealed_holdout_rerun.py --json
|
||||
@@ -58,7 +69,15 @@ python scripts/research/sealed_holdout_rerun.py --json
|
||||
- `source_audit_status` 对齐数据集;`evaluated_on` 表示最新固定口径重跑日期,并新增作用域说明。历史顶层指标仍有 `metrics_produced_by` 的旧日期与来源,未伪装成本次结果。
|
||||
- `status`、`valid_public_aa_cases`、`required_cases`、`top_1_rate`、`confirmation_coverage_rate`、`sealed_benchmark_id` 六个 runtime 键的值与类型全不变;尤其 `status=not_ready`、确认覆盖率为零。
|
||||
- `official_eval_implementation_hash_matches=false`、`official_eval_trial_count=0` 保留;新增重跑记录不冒充已重新获得独立性。
|
||||
- 未改任何冻结打分文件、生产判定、现有 `_candidate_moments()` 或 `request_from_case()`。
|
||||
- 本轮生产代理修复 transition-proximity 候选日期并将矩阵算法身份升为 `rectification-v5-matrix-scoring-8`,落在历史 12 文件之外。本研究未改权重、候选/事件/半径,未改 `_candidate_moments()` 或 `request_from_case()`。
|
||||
|
||||
## BUG-981 前后逐例对比
|
||||
|
||||
**20/20 例全部字段相同,改变 0 例;五项指标和附加 opaque 指标均不变。** 当前 JSON 的 `historical_comparison.trials` 保留全部 20 例 before/after 与 changed_fields,包括未变化行,并校验旧报告 SHA-256。不是只保留变化样本。
|
||||
|
||||
原因边界:本协议是 shadow fact ranker,`build_feature_fact_rows()` 原先已使用 `candidate_at` 日期,不调用被修复的 transition-proximity helper。扩展生产身份用于标识同一工作树上下文,不得写成该 shadow 路径执行了生产 helper。旧成绩因身份更新不再充当当前报告,但不能把数值本身说成错误;重新冻结也不恢复独立盲测资格。
|
||||
|
||||
准备阶段曾生成两份 `*.freeze.json` 和一个 `*.cross_midnight.rerun.json`。因安全工具拒绝覆盖旧 JSON,改用新的正式报告路径,停止未完成的 900 组合准备运行,评测器最终路径修改后重新冻结再完整重跑。所有准备产物保留并由历史目录的 `preparation_manifest.json` 标为 `superseded_preparation_not_final_report`;不得当最终结果。
|
||||
|
||||
## 防复发与既有断言变更
|
||||
|
||||
@@ -69,4 +88,16 @@ python scripts/research/sealed_holdout_rerun.py --json
|
||||
| Python / 前端:`current_tree_scorer.implementation_sha256` 等于旧 post-audit sidecar 的哈希 | 等于新固定口径报告 `frozen_record.implementation_sha256`,另验证当前树与新报告 | 契约当前树块的含义就是当前实现;继续绑定旧报告会强制过期。新增身份断言更严格,不弱化门禁 |
|
||||
| 官方标志 false、官方试次 0、六键等于历史报告、确认门全套行为断言 | 原值保留,另加固定口径重跑身份/试次和独立盲测 false | 固定口径重跑不是新的独立盲测,不借元数据迁移打开确认门 |
|
||||
|
||||
本机 Python 3.11.7 可跑;`python3` Windows 别名退出 49。该任务定向命令的最终测试清单与基线失败对照由 `PROGRESS-rectification-validation-20260920.md` 统一记录。特别是既有 v2 冻结身份测试不属本次刷新范围,不得为了通过而改其历史封存或断言。
|
||||
本轮既有测试调整三栏(此前 BUG-978 的三栏保留在上表):
|
||||
|
||||
| 原值 / 原断言 | 新值 / 新断言 | 原因 |
|
||||
| --- | --- | --- |
|
||||
| `FREEZE` 指向旧 `sealed_holdout_rerun_2026_09_20.freeze.json`,当前 `REPORT` 指向旧同名 JSON | 指向新 `sealed_holdout_rerun_cross_midnight_2026_09_20.final.freeze.json` 与新 JSON;保留原完整身份相等断言 | 历史冻结不可刷新,新实现/评测器必须重新预冻结;旧文件继续逐字节校验 |
|
||||
| 当前树身份只核对 12 文件 | 保留 12 文件相等与数量断言,另核对生产 18 文件、研究 7 文件、每文件哈希、归档 manifest、契约及报告 | 修复位于旧 12 文件之外;新增身份漂移拒绝,不弱化原断言 |
|
||||
| 当前报告与冻结身份一致 | 原断言全部保留,另核对冻结时间早于评分开始、评分前后身份一致、全部旧新逐例对比、六键类型与值不变 | 防止事后追认、选样与误开确认门 |
|
||||
| Python 确认合同仅比较 legacy 实现身份 | 原断言保留,新增扩展身份相等与 helper 在文件集内 | 原 12 文件不变不足以证明当前修复已被绑定 |
|
||||
| 前端确认测试读取旧固定口径 JSON | 由 T4 协调改为新正式 JSON,原门禁与身份断言保留 | 报告地址迁移,不改变确认行为 |
|
||||
|
||||
本轮定向四文件合计 **52/52 通过**(含 quick 桥接重复收集),身份/历史字节/全部 before-after 独立复核通过。20 例重跑用时 **3.158 秒**;当前报告 SHA-256:`0122d3cd1545416e1de9029b912d665d130049e9ea768a8972d18bb917331d38`。
|
||||
|
||||
本机 Python 3.11.7 可跑;`python3` Windows 别名退出 49。开工预检 remote verified,但 23 通过 / 1 既有 Windows 镜像路径断言失败,不冒写预检通过。本轮测试清单与环境/广域基线对照由 `../tasks/PROGRESS-rectification-cross-midnight-20260920.md` 统一记录;前轮记录仍在 `../tasks/PROGRESS-rectification-validation-20260920.md`。既有 v2 冻结身份测试不属刷新范围,不得为了通过而改历史封存或断言。
|
||||
|
||||
@@ -0,0 +1,106 @@
|
||||
{
|
||||
"record_version": "exposed-v3-fixed-protocol-cross-midnight-rerun-v2",
|
||||
"extended_identity": {
|
||||
"historical_artifacts_manifest_sha256": "6102a26a840be207b5858b3a4c0509274468ae9071e86d308cbbac7371d6864e",
|
||||
"production_scoring_files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/rectification/candidate_contrast.py",
|
||||
"scripts/rectification/case_holdout.py",
|
||||
"scripts/rectification/contracts.py",
|
||||
"scripts/rectification/dasha_transition_proximity.py",
|
||||
"scripts/rectification/event_probes.py",
|
||||
"scripts/rectification/scoring_service.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"production_scoring_sha256": "7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51",
|
||||
"research_files": [
|
||||
"scripts/research/reported_offset_sweep.py",
|
||||
"scripts/research/cluster_width_lib.py",
|
||||
"scripts/research/probe_supply_after_six.py",
|
||||
"scripts/research/sealed_holdout_rerun.py",
|
||||
"scripts/minute_rectification_blind_eval.py",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py",
|
||||
"scripts/minute_rectification_holdout_validator.py"
|
||||
],
|
||||
"research_implementation_sha256": "2931d4661bf53dde5c57f6492de65be4c00877293b7277353279a05eda5af215",
|
||||
"file_sha256": {
|
||||
"scripts/active_rectification_event_engine.py": "edbe93ef3c8bd65f986139dba6b95b96d9b0fa9920341fbe677cd6f53e4db84f",
|
||||
"scripts/active_rectification_events.py": "03e333ccf1103caf685dfe50fe890ab1793e945037085f50ce834a4a3b232222",
|
||||
"scripts/ashtakavarga.py": "cc32776674cf60713634a360a6f09774b3af0ec5d0f650687f52b79597b0dfbe",
|
||||
"scripts/dasha_analyzer.py": "5e53b1b4d4c7414c08dd0e7716f120a63c70ec3428ec93547f64d55c580da0e1",
|
||||
"scripts/divisional_charts_extended.py": "73bd50a83ac444830940184d0988cea621855f91db76a3f1a232014b992eacea",
|
||||
"scripts/domain_calculation_service.py": "9ee955874eea806c7b56f4246d90009c02ee8318b8f3dfe1d7132b1c0b4262e1",
|
||||
"scripts/jaimini.py": "b98bf965603dd2e3cc871fe40f3c8184731a701875ca50fef3ca68e9682eda81",
|
||||
"scripts/minute_rectification_blind_eval.py": "07aa30d8bad3278669495d5fea8be937aaf6367c8af980ff4f2fe9aa84ccc6ff",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py": "bdf460e3467208e8b87345c5beb2128f3e22de7eeffd68f72b281b3ad9affc92",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py": "f1a0d8d2482cd754d2e7f253a9606b0373519435a2bf58cfc3753e2b9ae0fb76",
|
||||
"scripts/minute_rectification_feature_facts_v4.py": "3d899a2cdd4825b8e32062443d487edcd2173a4ea55a36eb9f3abe587c0faf0f",
|
||||
"scripts/minute_rectification_holdout_validator.py": "b60b1c752e58672beaee8b36eb3b2348169726cda070ec969db0f13e7ed715af",
|
||||
"scripts/narayana_dasha.py": "7ff2c3238cd113b11b815e14b41967bd8e486ff62e36650e55613b9c79d1bfe3",
|
||||
"scripts/rectification/candidate_contrast.py": "10a360586359e72846593ef6d7c15a2cea79358245a34c288851406fe9e47ca8",
|
||||
"scripts/rectification/case_holdout.py": "fe0e699b617d9e216acdfeff87af4221c1d6f4c101bf33e1dbf14e2bcd75a8b5",
|
||||
"scripts/rectification/contracts.py": "fdd1a47b2e3dac8579e870a282ed74c68518b39f6c0576b27df766a4b8021143",
|
||||
"scripts/rectification/dasha_transition_proximity.py": "0afba494aa12899621173c3cc941c6c1d0be9cec0bdf4c4a1c9bcd004ab22de8",
|
||||
"scripts/rectification/event_probes.py": "45478ee3fdc36dc36e1a602bee5dd0627a5464275e370e64adb5826fe2aa58f5",
|
||||
"scripts/rectification/scoring_service.py": "0adf5700bb22dd8c4c3147027251210e21ec66a2c9f9a91217edba3ad9e5d1f5",
|
||||
"scripts/research/cluster_width_lib.py": "1d8f2ffa7795d62069deebb2de8e16749317037d4ed3f24a2c0092676e5cfcc6",
|
||||
"scripts/research/probe_supply_after_six.py": "098a5b4398d6f8d996917d809c2fc922ad64031d3b70660f8ef7c876aba92438",
|
||||
"scripts/research/reported_offset_sweep.py": "f81c681fd0fe257e90a761dfab370f623e5f8a881b62cc6b1c596cbd58800f8f",
|
||||
"scripts/research/sealed_holdout_rerun.py": "34a6724530fbecdecff9e5e07fa64980b8e8aff1fac2e000f4268debe1ee96e1",
|
||||
"scripts/shadbala.py": "912e0e6d169c2172aab85f71e4347e3e39193020777b43825b775d906400956b",
|
||||
"scripts/varga.py": "4331de5a25ea08729af91aae863943c4223183937f419d52dc63fd6fe7f27be6"
|
||||
},
|
||||
"hash_scope": "explicit_identity_file_sets_not_a_transitive_dependency_lock"
|
||||
},
|
||||
"production_identity_scope": "context_only_shadow_scorer_does_not_call_transition_proximity",
|
||||
"frozen_at_utc": "2026-09-20T04:56:54.914584+00:00",
|
||||
"dataset_path": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"algorithm_version": "birth-time-event-fact-ranker-v4-shadow",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"historical_frozen_sha256": "f41c298dd6cdcebe7a93e632f7191954be7987f6012ca2a34cecb6e447fbf196",
|
||||
"evaluator_sha256": "34a6724530fbecdecff9e5e07fa64980b8e8aff1fac2e000f4268debe1ee96e1",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"candidate_radius_minutes": [
|
||||
10
|
||||
],
|
||||
"minute_step": 1,
|
||||
"release_metrics": {
|
||||
"top_1_rate_minimum": 0.6,
|
||||
"top_3_rate_minimum": 0.85,
|
||||
"mean_absolute_minute_error_maximum": 2.0,
|
||||
"false_confirmation_rate_maximum": 0.05,
|
||||
"correct_insufficient_evidence_rejection_rate_minimum": 0.9
|
||||
},
|
||||
"results_previously_seen": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"must_not_use_for_tuning": true,
|
||||
"tie_breaker": "sha256(benchmark_id:case_id:candidate_time); published minute is never passed to the ranker",
|
||||
"metric_rank_definition": "competition_rank_1_plus_strictly_higher_scores_legacy_protocol",
|
||||
"extra_metric_rank_definition": "score_desc_then_existing_opaque_sha256_total_order"
|
||||
}
|
||||
@@ -0,0 +1,106 @@
|
||||
{
|
||||
"record_version": "exposed-v3-fixed-protocol-cross-midnight-rerun-v2",
|
||||
"extended_identity": {
|
||||
"historical_artifacts_manifest_sha256": "6102a26a840be207b5858b3a4c0509274468ae9071e86d308cbbac7371d6864e",
|
||||
"production_scoring_files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/rectification/candidate_contrast.py",
|
||||
"scripts/rectification/case_holdout.py",
|
||||
"scripts/rectification/contracts.py",
|
||||
"scripts/rectification/dasha_transition_proximity.py",
|
||||
"scripts/rectification/event_probes.py",
|
||||
"scripts/rectification/scoring_service.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"production_scoring_sha256": "7fffd1db612af1f244a8b5d5eac2f1a3f5a8e9e138c09ac6b3743b98f9a47a51",
|
||||
"research_files": [
|
||||
"scripts/research/reported_offset_sweep.py",
|
||||
"scripts/research/cluster_width_lib.py",
|
||||
"scripts/research/probe_supply_after_six.py",
|
||||
"scripts/research/sealed_holdout_rerun.py",
|
||||
"scripts/minute_rectification_blind_eval.py",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py",
|
||||
"scripts/minute_rectification_holdout_validator.py"
|
||||
],
|
||||
"research_implementation_sha256": "61dbb47eb46f9a01c78096677ac6238e3a65e9674ce4f2514b18e1ca50c0b954",
|
||||
"file_sha256": {
|
||||
"scripts/active_rectification_event_engine.py": "edbe93ef3c8bd65f986139dba6b95b96d9b0fa9920341fbe677cd6f53e4db84f",
|
||||
"scripts/active_rectification_events.py": "03e333ccf1103caf685dfe50fe890ab1793e945037085f50ce834a4a3b232222",
|
||||
"scripts/ashtakavarga.py": "cc32776674cf60713634a360a6f09774b3af0ec5d0f650687f52b79597b0dfbe",
|
||||
"scripts/dasha_analyzer.py": "5e53b1b4d4c7414c08dd0e7716f120a63c70ec3428ec93547f64d55c580da0e1",
|
||||
"scripts/divisional_charts_extended.py": "73bd50a83ac444830940184d0988cea621855f91db76a3f1a232014b992eacea",
|
||||
"scripts/domain_calculation_service.py": "9ee955874eea806c7b56f4246d90009c02ee8318b8f3dfe1d7132b1c0b4262e1",
|
||||
"scripts/jaimini.py": "b98bf965603dd2e3cc871fe40f3c8184731a701875ca50fef3ca68e9682eda81",
|
||||
"scripts/minute_rectification_blind_eval.py": "07aa30d8bad3278669495d5fea8be937aaf6367c8af980ff4f2fe9aa84ccc6ff",
|
||||
"scripts/minute_rectification_fact_blind_eval_v4.py": "bdf460e3467208e8b87345c5beb2128f3e22de7eeffd68f72b281b3ad9affc92",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py": "f1a0d8d2482cd754d2e7f253a9606b0373519435a2bf58cfc3753e2b9ae0fb76",
|
||||
"scripts/minute_rectification_feature_facts_v4.py": "3d899a2cdd4825b8e32062443d487edcd2173a4ea55a36eb9f3abe587c0faf0f",
|
||||
"scripts/minute_rectification_holdout_validator.py": "b60b1c752e58672beaee8b36eb3b2348169726cda070ec969db0f13e7ed715af",
|
||||
"scripts/narayana_dasha.py": "7ff2c3238cd113b11b815e14b41967bd8e486ff62e36650e55613b9c79d1bfe3",
|
||||
"scripts/rectification/candidate_contrast.py": "10a360586359e72846593ef6d7c15a2cea79358245a34c288851406fe9e47ca8",
|
||||
"scripts/rectification/case_holdout.py": "fe0e699b617d9e216acdfeff87af4221c1d6f4c101bf33e1dbf14e2bcd75a8b5",
|
||||
"scripts/rectification/contracts.py": "fdd1a47b2e3dac8579e870a282ed74c68518b39f6c0576b27df766a4b8021143",
|
||||
"scripts/rectification/dasha_transition_proximity.py": "0afba494aa12899621173c3cc941c6c1d0be9cec0bdf4c4a1c9bcd004ab22de8",
|
||||
"scripts/rectification/event_probes.py": "45478ee3fdc36dc36e1a602bee5dd0627a5464275e370e64adb5826fe2aa58f5",
|
||||
"scripts/rectification/scoring_service.py": "0adf5700bb22dd8c4c3147027251210e21ec66a2c9f9a91217edba3ad9e5d1f5",
|
||||
"scripts/research/cluster_width_lib.py": "1d8f2ffa7795d62069deebb2de8e16749317037d4ed3f24a2c0092676e5cfcc6",
|
||||
"scripts/research/probe_supply_after_six.py": "098a5b4398d6f8d996917d809c2fc922ad64031d3b70660f8ef7c876aba92438",
|
||||
"scripts/research/reported_offset_sweep.py": "cdcd7b583d9ecee8a313c25cac795923da2f8af45875d572f0a3efbc25f5d071",
|
||||
"scripts/research/sealed_holdout_rerun.py": "33c294b21db0631a68068dcf135b70db7aae6a1a3e66414bb72959a448883465",
|
||||
"scripts/shadbala.py": "912e0e6d169c2172aab85f71e4347e3e39193020777b43825b775d906400956b",
|
||||
"scripts/varga.py": "4331de5a25ea08729af91aae863943c4223183937f419d52dc63fd6fe7f27be6"
|
||||
},
|
||||
"hash_scope": "explicit_identity_file_sets_not_a_transitive_dependency_lock"
|
||||
},
|
||||
"production_identity_scope": "context_only_shadow_scorer_does_not_call_transition_proximity",
|
||||
"frozen_at_utc": "2026-09-20T04:50:19.815357+00:00",
|
||||
"dataset_path": "references/real_case_calibration/minute_rectification_holdout_v3.json",
|
||||
"dataset_sha256": "45aa76ab79dfc6c6404b6252083e1aa5c7e2e2b7de811fc11042a8ad2da3f0ea",
|
||||
"algorithm_version": "birth-time-event-fact-ranker-v4-shadow",
|
||||
"implementation_sha256": "b15d9ea15227cd58b3097555537fa7c37fb5a993905ada4f5aa04f52de1f8f18",
|
||||
"files": [
|
||||
"scripts/active_rectification_event_engine.py",
|
||||
"scripts/active_rectification_events.py",
|
||||
"scripts/ashtakavarga.py",
|
||||
"scripts/dasha_analyzer.py",
|
||||
"scripts/divisional_charts_extended.py",
|
||||
"scripts/domain_calculation_service.py",
|
||||
"scripts/jaimini.py",
|
||||
"scripts/minute_rectification_fact_ranker_v4.py",
|
||||
"scripts/minute_rectification_feature_facts_v4.py",
|
||||
"scripts/narayana_dasha.py",
|
||||
"scripts/shadbala.py",
|
||||
"scripts/varga.py"
|
||||
],
|
||||
"historical_frozen_sha256": "f41c298dd6cdcebe7a93e632f7191954be7987f6012ca2a34cecb6e447fbf196",
|
||||
"evaluator_sha256": "33c294b21db0631a68068dcf135b70db7aae6a1a3e66414bb72959a448883465",
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"candidate_radius_minutes": [
|
||||
10
|
||||
],
|
||||
"minute_step": 1,
|
||||
"release_metrics": {
|
||||
"top_1_rate_minimum": 0.6,
|
||||
"top_3_rate_minimum": 0.85,
|
||||
"mean_absolute_minute_error_maximum": 2.0,
|
||||
"false_confirmation_rate_maximum": 0.05,
|
||||
"correct_insufficient_evidence_rejection_rate_minimum": 0.9
|
||||
},
|
||||
"results_previously_seen": true,
|
||||
"official_valid_independent_blind": false,
|
||||
"must_not_use_for_tuning": true,
|
||||
"tie_breaker": "sha256(benchmark_id:case_id:candidate_time); published minute is never passed to the ranker",
|
||||
"metric_rank_definition": "competition_rank_1_plus_strictly_higher_scores_legacy_protocol",
|
||||
"extra_metric_rank_definition": "score_desc_then_existing_opaque_sha256_total_order"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user