research(rectification): archive partial varga-resolution study (BUG-1105)
Archive research scripts, regression tests, M1 results and safe M0 smoke. Keep the incomplete study and failing quick gate explicit. Exclude full M0 JSON, raw logs and unrelated oracle newline changes. Co-Authored-By: Claude Code <noreply@anthropic.com>
This commit is contained in:
+11
@@ -1,5 +1,16 @@
|
||||
# BLOCKED
|
||||
|
||||
## BUG-1105:离线研究验收的本机环境缺口(2026-09-30)
|
||||
|
||||
- 推送范围更新:用户要求阶段性归档到 staging;M0 全量 JSON 与原始日志保留本地不提交。明确暂存的研究代码、测试、M1 JSON、M0 smoke 及文档已随 tracked 隐私门复验(与研究定向合计 140 passed)。只解除提交子集的隐私覆盖缺口,不解除下述未提交产物、完整 quick 或研究验收阻塞。
|
||||
|
||||
- 当前 worktree 无可用项目 `.venv`,使用本机 Python 3.11.7,不改依赖锁或生产代码。
|
||||
- 开工 `artifacts/varga-resolution/baseline-quick.log` 的 Python 测试子段通过,随后 npm 缺 `tsx`,quick 整体 exit 1;`pre-work.log` 是另一项失败,focused 测试要求不存在的历史 `.workbuddy/skills/jyotish-vedic-astrology` 镜像路径。没有伪造镜像或弱化测试。
|
||||
- 本轮最终 quick **exit 1**:Python 子段 1000 passed、1 failed、1 skipped、3 subtests passed,失败为 `tests/test_consultation_native_layers.py::test_the_new_layer_leaves_every_existing_output_unchanged` 的原输出不变性断言。日志差异片段涉及 VedAstro 请求清单,输出被截断,根因未确定;不能归因为开工时的 tsx 缺失,也不能声称与基线失败逐项一致。详见 `artifacts/varga-resolution/closure-quick.log` 与 PROGRESS。恢复条件:独立定位新增失败后,在依赖齐全的验证环境复跑完整 quick;不修改生产代码或弱化断言消红。历史镜像断言由仓库环境合同维护方核实,不让研究脚本创建假镜像。
|
||||
- quick 还重写了 `references/oracle/artifacts/pending_packets/` 下 5 个已跟踪 JSON。独立只读比对确认 JSON 语义与 HEAD 相同,字节差异仅 LF→CRLF;来源是 oracle dashboard 链调用 `tajika_annual_oracle_queue.py` 的文本模式写入。文件现场保留、未恢复、未暂存,不纳入 BUG-1105 研究候选交付;此副作用不是上条不变性失败的已证根因。
|
||||
- 产物隐私补扫未通过:tracked 守卫 63 passed 不覆盖本轮未跟踪文件。显式复用隐私规则扫描新增产物时,M0 baseline JSON 的时刻数组 R003 命中 3 次;旧 baseline quick 日志混合编码无法以 UTF-8 或 GB18030 严格解码。未打印标记原值、未改既有隐私例外、未删除命中数据;需独立核查精确字段碰撞及完成日志保真扫描后,才可考虑纳入提交。证据 `artifacts/varga-resolution/closure-explicit-privacy.log`。
|
||||
- M0 两次字节复跑未完成、M2/M3 not_started 是研究欠项,**不是上述环境失败的推断后果**。不得据此关单或宣称已交付。
|
||||
|
||||
## TASK-chart-surface-polish:受控登录、实体手机与基线全量失败(2026-09-29)
|
||||
|
||||
- 本轮 Chrome、Edge、Docker 均可用,不套用历史“无 Chrome / 无 Docker”。真实 Chrome + golden 的本地组件验收及骨架高度修复后复验已完成(70+8 项通过);没有受控线上登录账号、实体手机和读屏实测;不能代替完整账户/人物/请求链路。清单见 `docs/testing/chart-surface-polish-20260929.md`。
|
||||
|
||||
@@ -0,0 +1,242 @@
|
||||
{
|
||||
"aggregates": [
|
||||
{
|
||||
"case_count": 1,
|
||||
"delivery_envelope": {
|
||||
"by_varga": {
|
||||
"D1": {
|
||||
"denominator": 1,
|
||||
"envelope_at_most_two_signs": 1,
|
||||
"envelope_at_most_two_signs_rate": 1.0,
|
||||
"envelope_majority_ties": 0,
|
||||
"envelope_majority_truth": 1,
|
||||
"envelope_majority_truth_rate": 1.0,
|
||||
"envelope_only_one_sign": 1,
|
||||
"envelope_only_one_sign_rate": 1.0,
|
||||
"real_at_most_two_segments": 1,
|
||||
"real_at_most_two_segments_rate": 1.0,
|
||||
"real_truth_segment_retained": 1,
|
||||
"real_truth_segment_retained_rate": 1.0
|
||||
},
|
||||
"D10": {
|
||||
"denominator": 1,
|
||||
"envelope_at_most_two_signs": 0,
|
||||
"envelope_at_most_two_signs_rate": 0.0,
|
||||
"envelope_majority_ties": 0,
|
||||
"envelope_majority_truth": 1,
|
||||
"envelope_majority_truth_rate": 1.0,
|
||||
"envelope_only_one_sign": 0,
|
||||
"envelope_only_one_sign_rate": 0.0,
|
||||
"real_at_most_two_segments": 0,
|
||||
"real_at_most_two_segments_rate": 0.0,
|
||||
"real_truth_segment_retained": 1,
|
||||
"real_truth_segment_retained_rate": 1.0
|
||||
},
|
||||
"D9": {
|
||||
"denominator": 1,
|
||||
"envelope_at_most_two_signs": 0,
|
||||
"envelope_at_most_two_signs_rate": 0.0,
|
||||
"envelope_majority_ties": 0,
|
||||
"envelope_majority_truth": 1,
|
||||
"envelope_majority_truth_rate": 1.0,
|
||||
"envelope_only_one_sign": 0,
|
||||
"envelope_only_one_sign_rate": 0.0,
|
||||
"real_at_most_two_segments": 0,
|
||||
"real_at_most_two_segments_rate": 0.0,
|
||||
"real_truth_segment_retained": 1,
|
||||
"real_truth_segment_retained_rate": 1.0
|
||||
}
|
||||
},
|
||||
"combination_at_most_two": 0,
|
||||
"combination_majority_ties": 0,
|
||||
"combination_majority_truth": 1,
|
||||
"combination_only_one": 0,
|
||||
"denominator": 1
|
||||
},
|
||||
"radius": 10,
|
||||
"reconciliation": {
|
||||
"all_zero_diff": true,
|
||||
"case_count": 1,
|
||||
"changed_case_count": 0,
|
||||
"denominator": 1
|
||||
},
|
||||
"window": {
|
||||
"case_count": 1,
|
||||
"d1_single_sign": 1,
|
||||
"mean_combination_count": 5.0,
|
||||
"mean_segment_counts": {
|
||||
"D1": 1.0,
|
||||
"D10": 3.0,
|
||||
"D9": 3.0
|
||||
},
|
||||
"mean_sign_counts": {
|
||||
"D1": 1.0,
|
||||
"D10": 3.0,
|
||||
"D9": 3.0
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"errors": [],
|
||||
"metadata": {
|
||||
"aggregation_denominator": "77 cases per radius when full run completes; empty coverage and ties are separate counts",
|
||||
"ayanamsa": "raman",
|
||||
"case_count_completed": 1,
|
||||
"case_count_requested": 1,
|
||||
"deterministic_json": true,
|
||||
"holdout": "references/real_case_calibration/minute_rectification_holdout_v5.json",
|
||||
"interval_envelope": "unionStillValidRange equivalent: min/max clock edges over the real valid candidate clusters; envelope is not the real set",
|
||||
"node_mode": "mean",
|
||||
"questions_answered_is_recorded_per_case": true,
|
||||
"questions_requested": 6,
|
||||
"radii": [
|
||||
10
|
||||
],
|
||||
"real_valid_candidate_set": "union of cluster_times from still_valid_public after replay; scoring-grid candidates only",
|
||||
"reconciliation_scope": "every case and radius, scoring grid only; single-case zero-diff is smoke evidence, not full-set evidence",
|
||||
"refresh_probes": false,
|
||||
"replay_probe_source": "existing event_probes.discriminating_event_probes; no refresh",
|
||||
"scoring_candidate_step_minutes": 2,
|
||||
"segment_definition": "maximal contiguous one-minute scan run of equal D1/D9/D10 sign; repeated non-contiguous signs retain separate IDs; 23:59->00:00 is contiguous",
|
||||
"segment_scan_step_minutes": 1,
|
||||
"vargas": [
|
||||
"D1",
|
||||
"D9",
|
||||
"D10"
|
||||
]
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"answered_count": 0,
|
||||
"case_id": "albert_brooks_1947_aa_v4_holdout",
|
||||
"probe_count": 0,
|
||||
"radius": 10,
|
||||
"reconciliation": {
|
||||
"candidates": 11,
|
||||
"changed": 0,
|
||||
"details": []
|
||||
},
|
||||
"scan_point_count": 21,
|
||||
"scoring_candidate_count": 11,
|
||||
"six_question_delivery": {
|
||||
"by_varga": {
|
||||
"D1": {
|
||||
"envelope_end": "03:10",
|
||||
"envelope_majority_tie": false,
|
||||
"envelope_majority_truth": true,
|
||||
"envelope_scan_point_count": 21,
|
||||
"envelope_segment_count": 1,
|
||||
"envelope_sign_count": 1,
|
||||
"envelope_start": "02:50",
|
||||
"real_valid_candidate_count": 11,
|
||||
"real_valid_majority_tie": false,
|
||||
"real_valid_majority_truth": true,
|
||||
"real_valid_segment_count": 1,
|
||||
"real_valid_sign_count": 1,
|
||||
"truth_segment_id": 0,
|
||||
"truth_segment_retained_in_real_set": true,
|
||||
"truth_sign": 2
|
||||
},
|
||||
"D10": {
|
||||
"envelope_end": "03:10",
|
||||
"envelope_majority_tie": false,
|
||||
"envelope_majority_truth": true,
|
||||
"envelope_scan_point_count": 21,
|
||||
"envelope_segment_count": 3,
|
||||
"envelope_sign_count": 3,
|
||||
"envelope_start": "02:50",
|
||||
"real_valid_candidate_count": 11,
|
||||
"real_valid_majority_tie": true,
|
||||
"real_valid_majority_truth": null,
|
||||
"real_valid_segment_count": 3,
|
||||
"real_valid_sign_count": 3,
|
||||
"truth_segment_id": 1,
|
||||
"truth_segment_retained_in_real_set": true,
|
||||
"truth_sign": 5
|
||||
},
|
||||
"D9": {
|
||||
"envelope_end": "03:10",
|
||||
"envelope_majority_tie": false,
|
||||
"envelope_majority_truth": true,
|
||||
"envelope_scan_point_count": 21,
|
||||
"envelope_segment_count": 3,
|
||||
"envelope_sign_count": 3,
|
||||
"envelope_start": "02:50",
|
||||
"real_valid_candidate_count": 11,
|
||||
"real_valid_majority_tie": true,
|
||||
"real_valid_majority_truth": null,
|
||||
"real_valid_segment_count": 3,
|
||||
"real_valid_sign_count": 3,
|
||||
"truth_segment_id": 1,
|
||||
"truth_segment_retained_in_real_set": true,
|
||||
"truth_sign": 9
|
||||
}
|
||||
},
|
||||
"combination": {
|
||||
"envelope_count": 5,
|
||||
"envelope_majority_tie": false,
|
||||
"envelope_majority_truth": true,
|
||||
"real_valid_count": 5
|
||||
},
|
||||
"interval_envelope": {
|
||||
"end": "03:10",
|
||||
"scan_point_count": 21,
|
||||
"start": "02:50",
|
||||
"times": [
|
||||
"02:50",
|
||||
"02:51",
|
||||
"02:52",
|
||||
"02:53",
|
||||
"02:54",
|
||||
"02:55",
|
||||
"02:56",
|
||||
"02:57",
|
||||
"02:58",
|
||||
"02:59",
|
||||
"03:00",
|
||||
"03:01",
|
||||
"03:02",
|
||||
"03:03",
|
||||
"03:04",
|
||||
"03:05",
|
||||
"03:06",
|
||||
"03:07",
|
||||
"03:08",
|
||||
"03:09",
|
||||
"03:10"
|
||||
]
|
||||
},
|
||||
"scoring_candidate_times": [
|
||||
"02:56",
|
||||
"02:58",
|
||||
"03:00",
|
||||
"03:02",
|
||||
"02:52",
|
||||
"02:54",
|
||||
"03:06",
|
||||
"03:08",
|
||||
"03:04",
|
||||
"03:10",
|
||||
"02:50"
|
||||
]
|
||||
},
|
||||
"true_time": "03:00",
|
||||
"window": {
|
||||
"combination_count": 5,
|
||||
"d1_single_sign": true,
|
||||
"scan_point_count": 21,
|
||||
"segment_counts": {
|
||||
"D1": 1,
|
||||
"D10": 3,
|
||||
"D9": 3
|
||||
},
|
||||
"sign_counts": {
|
||||
"D1": 1,
|
||||
"D10": 3,
|
||||
"D9": 3
|
||||
}
|
||||
}
|
||||
}
|
||||
],
|
||||
"schema": "bug-1105-varga-resolution-v1"
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -14906,6 +14906,23 @@
|
||||
- 复发自:无(BUG-1040 的覆盖缺口)
|
||||
- 修复版本:分支 `codex/new-chat-prewarm-20260929`
|
||||
|
||||
## BUG-1105 | 校正目标是分钟而非所需分盘,缺少分盘不可判出口
|
||||
|
||||
- 状态:investigating(离线研究部分完成;生产未修改,未验收)
|
||||
- 首次发现 / 最近更新:2026-09-30 / 2026-09-30
|
||||
- 影响面:生时校正目标、分盘段汇总与按问题域选题;本轮仅 `scripts/research/varga_resolution_*.py`、研究测试与文档。
|
||||
- 用户现象:分钟范围无法收窄时,不知道应使用哪张分盘;目标盘在窗口内不变时仍走分钟校正流程,宽窗口分盘分不开时缺少明确出口。
|
||||
- 触发条件:候选窗口包含多个分钟,但目标盘未变或 D9/D10 存在多个连续上升段。
|
||||
- 根因(产品层研究假设):流程围绕分钟而非目标盘组织,未按事业/婚恋/综合选择目标;是否能以按段汇总或选题改善,尚不能据阶段性结果确定。
|
||||
- 修复:尚未修复;不改生产引擎、评分权重、采用/确认门、前端及冻结文件,不重开 BUG-1091 的打分路线。
|
||||
- 已确认事实:77 例公开 AA,四个半径 M0 首轮 308 个 case-radius 计分对账 changed=0;M1 全量 308 条、errors=0。±60 D9/D10 真值段保留各 74/77,低于任务书既有区间基线 75/77,当前方案不过门。±30 D9/D10 的 LOO 1.000 仅覆盖 11/77、16/77;top 命中包含并列,不能写成唯一段已确定。
|
||||
- 验证:既有 M0 定向 4 passed;M1 smoke/full exit 0;本轮补充测试与门禁证据见 `docs/tasks/PROGRESS-rectification-varga-resolution-research-20260930.md`。M0 起点表逐格一致性及二次字节复跑未闭环,M2/M3 not_started,M4 仅阶段性报告。
|
||||
- 环境缺口:开工 quick 的 Python 子段通过,但 npm 缺 tsx 导致整体 exit 1;独立 pre-work 因缺历史 WorkBuddy 镜像路径断言而 fail。没有伪造目录或弱化测试。
|
||||
- 防复发:真实保留集合与区间包络分开;保持 LOO 训练/验证隔离、空覆盖、并列与 LMT 分层;低覆盖高准确率不等于全集能力,任何真值段保留下降不得通过调阈值隐藏;全量复跑未完成不得声称字节一致。
|
||||
- 相关记录:BUG-560、BUG-1084、BUG-1090、BUG-1091;`TASK-rectification-varga-resolution-research-20260930.md`
|
||||
- 复发自:无(目标口径研究,不修改既有防复发约束)
|
||||
- 修复版本:未修;研究工作树 `codex/rectification-varga-resolution-research-20260930`。用户要求阶段性归档到 staging;仅提交隐私扫描通过的代码、测试、M1 结果、M0 测试必需小样本与文档,M0 全量原始 JSON / 原始日志不提交。提交不等于研究验收、生产修复或部署;交付范围与门禁失败见 PROGRESS。
|
||||
|
||||
## BUG-1106 | 打开报告详情先闪一个卡在顶部、没有顶栏的转圈,再换成居中的另一个转圈
|
||||
|
||||
- 状态:resolved(代码 + 回归测试 + 本机 Chromium 手机视口截图;iPhone 真机欠)
|
||||
|
||||
@@ -116,6 +116,12 @@ Do not open new product surfaces before at least one of these four lanes is clos
|
||||
- Vimsopaka semantic mapping for `NEECHA_BHANGA / GREAT_FRIEND / GREAT_ENEMY`
|
||||
- functional role now enters strict evidence; follow-up is Technique Audit Table rendering.
|
||||
|
||||
## 生时校正:分盘上升段研究(2026-09-30,BUG-1105)
|
||||
|
||||
- [任务书](../tasks/TASK-rectification-varga-resolution-research-20260930.md)、[进度](../tasks/PROGRESS-rectification-varga-resolution-research-20260930.md)、[阶段性报告](rectification_varga_resolution_2026_09_30.md)。仅离线研究,无生产实现授权。
|
||||
- M0 首轮与 M1 全量各 308 个 case-radius 已在本地落盘;本次 staging 归档仅含研究代码/测试、M1 JSON、测试必需的 M0 smoke 和文档,M0 全量原始 JSON 与运行日志因交付检查缺口不提交。M0 起点表逐格一致性、两次字节复跑尚未闭环,不能称整单验收通过。
|
||||
- M2/M3 `not_started`,M4 仅阶段性建议;±60 D9/D10 真值段保留各 74/77,低于既有区间基线 75/77,不通过研究红线。筛选子集的高准确率不能冒充全集或唯一段准确率。
|
||||
|
||||
## 生时校正:候选分不开(2026-09-14 收口)
|
||||
|
||||
- 定论页(先读这个):`<repo>/docs/research/rectification_minute_resolution_closure_2026_09_14.md`
|
||||
|
||||
@@ -2,6 +2,18 @@
|
||||
|
||||
Purpose: read this file before substantial project work. It exists to stop repeat mistakes caused by multiple Codex windows, WorkBuddy mirrors, local drafts, backup folders, and partial cloud-git visibility.
|
||||
|
||||
## 2026-09-30 · BUG-1105 获准阶段性归档后的推送准备
|
||||
|
||||
- 用户获知未验收、quick 和产物隐私缺口后要求 push staging。fetch 到 `5208f19b7`,研究基线之后 4 个提交不改 Python 研究依赖。仅提交隐私守卫覆盖并通过的 13 个精确路径;M0 全量 JSON、原始日志和 5 个 oracle 换行副作用留在原工作树,不编码规避、不改守卫。
|
||||
- pre-work 再次 exit 1:remote verified,focused 两项失败为历史镜像缺失以及 candidate_count=6 小于 workspace_residue_count=29;后者与已有同类机器目录数量假设一致。不得以允许推送冒充预检/研究通过。加入 index 后研究与隐私合跑 140 passed,完整 quick 的新增失败仍未解除。
|
||||
|
||||
## 2026-09-30 · BUG-1105 离线研究收尾验证
|
||||
|
||||
- 研究工作树 `codex/rectification-varga-resolution-research-20260930` 使用本机 Python 3.11.7。开工 quick 的 Python 子段通过后因 npm 缺 tsx 失败;pre-work 的历史 WorkBuddy 镜像断言是另一项失败,不混写。
|
||||
- 最终 `artifacts/varga-resolution/closure-quick.log` 为 exit 1:1000 passed、1 failed、1 skipped、3 subtests passed。新增失败是 `test_the_new_layer_leaves_every_existing_output_unchanged` 原输出不变性断言;日志仅给出被截断的请求清单差异,根因未确定,不得套用开工失败豁免,也不为消红修改生产代码。
|
||||
- quick 经 external oracle sanity → master dashboard → annual closure status → annual oracle queue 重写 5 个 tracked pending packet JSON;独立核查 JSON 全等、换行归一后的字节与 HEAD 相同,仅 LF→CRLF。Windows 文本模式写入造成工作树副作用,未证明与上述断言失败有关;保留现场、不暂存、不顺带提交。后续验证应将生成型检查隔离到临时输出,不默认门禁只读。
|
||||
- tracked 隐私 63 passed 不覆盖新研究产物;显式补扫有 3 处时刻字段 R003 命中及旧日志严格解码失败。保持未通过,不删除证据、不改例外;详见本轮 PROGRESS / BLOCKED。未提交、未推送、未部署。
|
||||
|
||||
## 2026-09-29 · 星盘展示获准推送后的预检
|
||||
|
||||
- 产品在获知本轮全量 29 项既有失败后明确授权推送 staging。重新 fetch,远端仍为已验基线 `78b9c5b65cda44650e8c7ca019931ce33a85a95d`;19 个业务/测试变动文件与 final-r2 快照哈希一致,隐私检查 63/63 通过。
|
||||
|
||||
@@ -0,0 +1,99 @@
|
||||
# 生时校正「分盘上升段」研究结论(BUG-1105,2026-09-30)
|
||||
|
||||
## 1. 范围与诚实边界
|
||||
|
||||
这是基于公开 Rodden-AA v5 开放集的离线研究,不是盲测,也不是生产准确率承诺。评价集为 77 例公开资料;所有数字均保留分母、空 coverage、并列和 LMT(1900 年前)分层。参数固定为 Raman ayanamsa、mean node、2 分钟评分候选网格、1 分钟分盘扫描,既有探针 `refresh_probes=false`。生产代码、权重、门槛和前端均未修改。
|
||||
|
||||
M0 的首轮完整结果为 308 个 case-radius、逐案计分对账 changed=0;第二次 M0 字节复跑因耗时中止,不能声称逐字节复现。M1 全量 308 个 case-radius、0 errors 已完成。M2、M3 尚未执行,M4 只有本页的研究边界与阶段性结论,不构成实现授权。
|
||||
|
||||
## 2. 一句话结论表
|
||||
|
||||
| 阶段 | 结论 | 是否建议实现 |
|
||||
|---|---|---|
|
||||
| M0 窗内分盘段 | D1 在 ±10/±15 几乎稳定;D9/D10 在 ±10 可较常收到 ≤2 段;±30 起分盘候选明显增多,±60 更不稳定。真值段首轮真实保留集合仍为 D1 76/77、D9 76/77、D10 76/77(±10);±60 为 D1 76/77、D9 74/77、D10 74/77。 | 仅支持继续研究,不直接实现 |
|
||||
| M1 段级汇总 | LOO 选择阈值后,D9/D10 的高 accuracy 只覆盖少量高占比案例;±60 D9 仅 17/77、D10 9/77 eligible,且 D9 accuracy 0.588。uniform 并列更多。 | 不建议单独以阈值交付分盘结论 |
|
||||
| M2 按段选题 | 未执行;没有证据证明按段信息增益优于线上选题。 | 不建议实现 |
|
||||
| M3 按问题域取盘 | 未执行;事业/婚恋/综合三策略没有实测比较。 | 不建议实现 |
|
||||
|
||||
## 3. M0 结果:盘型比分钟更有信息,但存在窗口边界
|
||||
|
||||
| 半径 | D1 窗内单一上升 | D9 平均段数 | D10 平均段数 | D1×D9×D10 平均组合 |
|
||||
|---|---:|---:|---:|---:|
|
||||
| ±10 | 64/77 | 2.43 | 2.51 | 3.84 |
|
||||
| ±15 | 51/77 | 3.22 | 3.26 | 5.27 |
|
||||
| ±30 | 37/77 | 5.34 | 5.60 | 9.62 |
|
||||
| ±60 | 10/77 | 9.64 | 10.29 | 18.36 |
|
||||
|
||||
六题后真实保留集合中的真值段保留:
|
||||
|
||||
| 半径 | D1 | D9 | D10 |
|
||||
|---|---:|---:|---:|
|
||||
| ±10 | 76/77 | 76/77 | 76/77 |
|
||||
| ±15 | 75/77 | 75/77 | 76/77 |
|
||||
| ±30 | 76/77 | 76/77 | 76/77 |
|
||||
| ±60 | 76/77 | 74/77 | 74/77 |
|
||||
|
||||
这些数字只能说明“真值段通常没有被排除”,不能说明系统已经正确选中唯一分盘。特别是 ±30/±60 的 D9/D10,段数和并列快速增加。**按任务书硬红线 1,±60 的 D9/D10 各 74/77,低于既有区间基线 75/77,当前口径不过门**;不能用阈值筛掉失败案例后重报通过。±15 未在原三档区间基线中给出参考数,不能自行套用另一半径。M0 首轮运行完成不等于 M0 已验收:任务书起点表与新表仍有种数/连续段、区间包络/真实集合差异需要逐格解释,第二次字节复跑也未完成。
|
||||
|
||||
## 4. M1 结果:阈值可以筛出高置信子集,但不能冒充全集能力
|
||||
|
||||
下面是 `raw` 模式全量 top 段正确 / 真值段保留 / 有效候选 ≤2 段,以及逐案例 LOO 选择阈值后的 eligible、accuracy、truth-retained。`top_segment_correct` 是**真值段属于最高质量段集合**(并列也算命中),不是唯一段选对率;并列另报,不能据此宣称唯一盘型已确定。
|
||||
|
||||
“六题后”表示最多六题预算,并非每例实际答足六题:308 个 case-radius 中 36 个无探针、38 个零回答,实际回答 0/1/2/3/4/5/6 题分别为 38/6/5/9/11/16/223 个;全部保留在分母中。
|
||||
|
||||
| 半径 | 盘 | 全量 top 正确 | 真值保留 | ≤2 段 | LOO eligible / accuracy / retained |
|
||||
|---|---|---:|---:|---:|---:|
|
||||
| ±10 | D1 | 76/77 | 76/77 | 77/77 | 77 / 0.987 / 0.987 |
|
||||
| ±10 | D9 | 71/77 | 76/77 | 68/77 | 23 / 0.913 / 0.957 |
|
||||
| ±10 | D10 | 69/77 | 76/77 | 63/77 | 29 / 0.931 / 0.966 |
|
||||
| ±15 | D1 | 74/77 | 75/77 | 77/77 | 77 / 0.961 / 0.974 |
|
||||
| ±15 | D9 | 61/77 | 75/77 | 50/77 | 14 / 0.929 / 0.929 |
|
||||
| ±15 | D10 | 59/77 | 76/77 | 52/77 | 8 / 0.875 / 0.875 |
|
||||
| ±30 | D1 | 75/77 | 76/77 | 77/77 | 75 / 0.973 / 0.987 |
|
||||
| ±30 | D9 | 50/77 | 76/77 | 25/77 | 11 / 1.000 / 1.000 |
|
||||
| ±30 | D10 | 52/77 | 76/77 | 25/77 | 16 / 1.000 / 1.000 |
|
||||
| ±60 | D1 | 68/77 | 76/77 | 76/77 | 54 / 0.944 / 0.981 |
|
||||
| ±60 | D9 | 41/77 | 74/77 | 11/77 | 17 / 0.588 / 0.941 |
|
||||
| ±60 | D10 | 34/77 | 74/77 | 10/77 | 9 / 0.778 / 1.000 |
|
||||
|
||||
±30 的 D9/D10 “1.000”只是在 11/77 和 16/77 的筛选子集上成立;不能用于对所有用户交付“盘型已确定”。±60 的 D9 结果直接显示阈值筛选也不稳定。
|
||||
|
||||
三种段质量均已计算:`raw`、`percent`、`uniform`。先将代表候选的后验分复制到其 `cluster_times` 中的评分网格候选,再只对真实保留候选求段内和;不是给独立扫描网格的每分钟插值。`percent` 在当前实现中只做非负总量比例化,非负分数下与 raw 的 top-share 数学等价,不是另一种概率校准或线上百分比先验回放;`uniform` 不使用分数,依赖段内保留候选数量,且产生更多并列。
|
||||
|
||||
LOO 每次只以另外 76 例选择阈值,优先顺序为训练准确率、真值保留率、eligible 数、较低阈值;训练 eligible 至少 5 例。固定阈值表不拟合参数,其汇总可与逐例留出评估相同,但不能当作已择阈值的无偏验证。当前名为 `full_fit` 的字段只是全集固定阈值对照,不代表另一轮学习器拟合。M1 JSON 尚缺独立完整的运行参数 metadata,复现还须结合本页与 M0 metadata;`deterministic_json=true` 只声明确定性序列化意图,不证明两次全量已逐字节一致。完整逐案数组、固定阈值表、LOO 选择阈值表、空 coverage、并列及 LMT 分层见:
|
||||
|
||||
- `artifacts/varga-resolution/varga_resolution_m1.json`
|
||||
|
||||
## 5. 可信度口径(研究建议,不是生产实现)
|
||||
|
||||
- **可直接说“无需校正”**:只有目标盘在当前窗口扫描中只有一个连续段,且该盘的输入窗口与资料精度满足产品既有门槛;不能由高 top-share 单独推出。
|
||||
- **较可信**:目标盘剩余不超过两段、真值保留门没有被触发、且不是 ±30/±60 的高并列情况;仍只能说“目前更像/可按该盘继续看”,不能写 confirmed。
|
||||
- **倾向**:存在多个段但 top-share 达到训练折选择的阈值;必须显示仍有多少段、是否并列和覆盖分母。
|
||||
- **不可判**:±30 以上 D9/D10 剩余段多、并列或 LOO eligible 很低;诚实出口应只交付 D1 或明确“分盘暂时分不开”,不要继续给伪精确分钟。
|
||||
|
||||
以上档位是研究报告用语,不是批准生产阈值;M2/M3 尚未提供足够证据支持交互或停题策略。
|
||||
|
||||
## 6. 产物与复跑
|
||||
|
||||
**归档范围说明**:用户要求将阶段性研究推到 staging;本次仅纳入隐私扫描通过的研究代码、测试、M1 JSON、测试必需的 M0 smoke JSON 和文档。下面的 M0 baseline JSON 及所有原始日志仍只在本地研究树,未提交;不将路径列出冒充远端证据已齐。完整提交/排除清单及 M0、M1 哈希见 PROGRESS 顶部。研究验收、完整 quick 和未提交产物的隐私缺口保持未通过。
|
||||
|
||||
- `scripts/research/varga_resolution_lib.py`
|
||||
- `scripts/research/varga_resolution_probe.py`
|
||||
- `scripts/research/varga_resolution_m1.py`
|
||||
- `docs/research/varga_resolution_baseline_2026_09_30.json`
|
||||
- `artifacts/varga-resolution/varga_resolution_m1.json`
|
||||
- `tests/test_varga_resolution_research.py`
|
||||
|
||||
既有验证:M0 定向测试 4 passed;M1 烟测、M1 全量均 exit 0;M1 全量为 308 条、0 errors。开工快速门 `baseline-quick.log` 的 Python 子段为 1001 passed、1 skipped、3 subtests passed,但随后 `npm test` 因 `tsx` 不可用失败,整体 exit 1。另一个独立命令的 `pre-work.log` 为 fail:focused preflight fragment-scan 测试要求本机不存在的 `.workbuddy/skills/jyotish-vedic-astrology` 镜像路径。两项不能混为一谈,均不能写成通过;本轮补验见 PROGRESS。
|
||||
|
||||
本轮补充的研究测试独立复验 **77 passed(exit 0)**,包括上述表格与机器结果一致性、36 组 M1 汇总及 LOO 重算;这不是全引擎复跑。tracked 隐私测试 63 passed,但显式补扫未跟踪产物未通过:M0 JSON 有 3 处 R003 时刻字段命中、旧 quick 日志严格解码失败,详见 PROGRESS / BLOCKED。没有修改隐私例外,当前不可声明隐私全绿或可交付。
|
||||
|
||||
最终完整 quick 补验 **exit 1**:Python 子段为 **1000 passed、1 failed、1 skipped、3 subtests passed**。失败是 `test_the_new_layer_leaves_every_existing_output_unchanged` 的原输出不变性断言,不是开工 quick 的 tsx 缺失;根因尚未确定,不能宣称失败清单与基线一致。快速门另产生 5 个 oracle JSON 的 LF→CRLF 副作用,独立比对确认无语义变化;保留现场且不纳入本研究候选交付。证据与范围见 PROGRESS / BLOCKED。
|
||||
|
||||
## 7. 后续实现单必须先补的内容
|
||||
|
||||
若产品仍要推进实现,另立实现单并明确串行依赖:
|
||||
|
||||
1. 先补齐 M0/M1 的验收缺口;按任务书让步顺序优先 M3:事业 D1+D10、婚恋 D1+D9、综合 D1+D9+D10 的取盘策略和“不用校正”比例。
|
||||
2. 再完成 M2:同一题池、同样六题预算下,按段信息增益与线上按分钟选题逐格比较;必须包含 wrong1/wrong2、事件日期 ±7 天、LMT 分层。
|
||||
3. 只有两者通过真值段保留红线,才能讨论 UI 交付卡、段扫描时机、段级停止和阈值文案;不得先把本研究脚本接入生产。
|
||||
@@ -0,0 +1,124 @@
|
||||
# PROGRESS · 生时校正分盘上升段研究(BUG-1105)
|
||||
|
||||
更新时间:2026-09-30
|
||||
|
||||
## 本次 staging 归档范围(用户要求推送后)
|
||||
|
||||
用户在收到未验收、快速门失败及隐私补扫缺口说明后要求 push 到 staging。本次仅归档可通过隐私扫描的阶段性研究代码、测试、M1 机器结果和文档;**不是研究验收通过,不批准生产实现,不关闭 BUG-1105**。下文“未提交/未推送”均是此前收尾时的历史状态;本次远端是否写入以 push 与远端 SHA 核对为准。
|
||||
|
||||
- 提交范围:3 个研究脚本、研究测试、`artifacts/varga-resolution/varga_resolution_m1.json`、现有测试必需的 `artifacts/varga-resolution/m0-smoke.json`,以及本轮报告、进度、索引、BUG 历史、BLOCKED 和错误台账。M0 smoke 与 M1 JSON 均经显式隐私扫描,0 findings;smoke 不代替 M0 全量证据。
|
||||
- **明确不提交** M0 全量原始 JSON、所有原始运行日志/其余烟测产物,以及 5 个 oracle 换行副作用文件。它们在原研究工作树保留,不删除、不改编码规避扫描、不改隐私例外;文中这些路径指本地证据,不代表远端已包含。
|
||||
- M0 原始 JSON 的 SHA-256:`c19faccb95076855420b9e5fbbbff9de13b4e611e2fe4f699af39474bd82506c`;M1 JSON 的 SHA-256:`a0f481e85e926fa3ca400c6cae06b608407540418f73fd6b0960067627b2ed24`。哈希仅用于身份追溯,不证明二次复跑。
|
||||
- 推送前 fetch 成功,远端从研究基线 `82db6373d` 前进到 `5208f19b7`(4 个提交);新增远端提交未改 `scripts/`、`tests/`、公开评价集及本任务书。先整合远端,禁止强推覆盖。
|
||||
- 推送前 `pre_work_check.py` 再次 exit 1:remote verified;focused 除历史镜像路径断言外,还出现 `candidate_count=6 < workspace_residue_count=29` 的数量断言。保持失败,不造目录、不清其他会话文件;本地日志 `artifacts/varga-resolution/push-pre-work.log`。
|
||||
- 全量 quick 仍采用本页所列实际失败结果;已知失败和 M0/M1 验收欠项不因允许归档而转绿。隐私要求只对明确选择的提交子集重新验收,未提交产物仍保留未通过状态。
|
||||
- 将上述 13 个精确路径加入 index 后,研究与隐私测试合跑 **140 passed,exit 0(6.73s)**。此次 tracked 隐私门已覆盖拟提交的新脚本、测试和两个 JSON,不再仅依赖未跟踪文件的补扫。两个机器 JSON 保留原始 CRLF 字节及哈希;默认 whitespace check 报 CRLF,采用单次命令 `core.whitespace=blank-at-eol,blank-at-eof,space-before-tab,cr-at-eol` 检查通过,不改仓库配置,不改机器证据。
|
||||
|
||||
## 当前结论(M0/M1 产物已落盘,整单未验收)
|
||||
|
||||
本轮已完成 **M0 首轮运行与 M1 全量分析**,未修改生产代码、未提交、未 push。M0 起点表逐格一致性及第二次字节复跑未闭环;M1 专项测试与记录补验见下文。M2/M3 为 `not_started`,M4 仅有阶段性报告,不能写成整单结论。±60 D9/D10 真值段保留各 74/77,低于任务书区间基线 75/77,明确不过硬红线 1,不建议接入生产。
|
||||
|
||||
## M0 · 首轮运行完成,验收未闭环
|
||||
|
||||
- 评价集:`references/real_case_calibration/minute_rectification_holdout_v5.json`,77 例。
|
||||
- 参数:`ayanamsa=raman`、`node_mode=mean`。
|
||||
- 评分候选网格:固定 **2 分钟**;每档评分候选数为 ±10=11、±15=16、±30=31、±60=61。
|
||||
- 分盘段扫描网格:独立固定 **1 分钟**;每档扫描点数为 ±10=21、±15=31、±30=61、±60=121。
|
||||
- 探针口径:既有 `event_probes.discriminating_event_probes`,固定 `refresh=False`;首轮空探针如实记录,不切换 `refresh=True` 制造题目。
|
||||
- 连续段:按 D1/D9/D10 上升星座在 1 分钟扫描序列中的最大连续运行编号;非连续再次出现的同一星座不合并;`23:59 -> 00:00` 视为连续,缺样本不跨接。
|
||||
- 线上结果口径:机器结果同时保存评分网格候选、`still_valid_public` 展开的真实保留候选集合和 `unionStillValidRange` 等价的首尾区间包络;包络不等于真实集合。
|
||||
- 完整逐案对账:四个半径均为 77/77 例,`changed_case_count=0`,错误数 0。单例零差只作 smoke;本结论使用了 77 例×每档半径完整对账。
|
||||
|
||||
### M0 实际汇总(分母均为 77)
|
||||
|
||||
| 半径 | D1 窗内单一上升星 | D9 平均段数 | D10 平均段数 | D1×D9×D10 平均组合 | 真值段仍在真实保留集合 |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| ±10 | 64/77 | 2.43 | 2.51 | 3.84 | D1 76/77;D9 76/77;D10 76/77 |
|
||||
| ±15 | 51/77 | 3.22 | 3.26 | 5.27 | D1 75/77;D9 75/77;D10 76/77 |
|
||||
| ±30 | 37/77 | 5.34 | 5.60 | 9.62 | D1 76/77;D9 76/77;D10 76/77 |
|
||||
| ±60 | 10/77 | 9.64 | 10.29 | 18.36 | D1 76/77;D9 74/77;D10 74/77 |
|
||||
|
||||
注意:以上是真实保留集合上的段保留统计,不是只对区间包络统计;区间包络中的并列、空 coverage 另列于机器 JSON。
|
||||
|
||||
## 机器结果与脚本
|
||||
|
||||
- `scripts/research/varga_resolution_lib.py`
|
||||
- `scripts/research/varga_resolution_probe.py`
|
||||
- `docs/research/varga_resolution_baseline_2026_09_30.json`
|
||||
- `artifacts/varga-resolution/m0-full.log`
|
||||
- `tests/test_varga_resolution_research.py`
|
||||
|
||||
机器 JSON metadata 明确记录:评分 2 分钟、扫描 1 分钟、`refresh_probes=false`、真实集合/区间包络区别、77 例逐档对账分母和跨午夜段定义;不含生成时间或耗时字段,便于逐字节复跑。
|
||||
|
||||
## M1 · 已完成(研究侧)
|
||||
|
||||
- 新增 `scripts/research/varga_resolution_m1.py`,沿用 M0 的 2 分钟评分 / 1 分钟扫描 / `refresh_probes=false` 链路,不改生产代码。
|
||||
- 全量完成 `77 × 4 = 308` 个 case-radius,`errors=0`;结果见 `artifacts/varga-resolution/varga_resolution_m1.json`,日志见 `m1-full-v2.log`。
|
||||
- 每盘分别计算 `raw`、`percent`、`uniform` 段质量;输出固定阈值 `0.5/0.6/0.7/0.8/0.9`、全样本拟合对照、逐案例留一法阈值选择与验证结果。
|
||||
- 留一法严格排除验证案例;输出 `eligible / coverage / accuracy / truth_retained`,并单列空 coverage、并列、真值排除和 `1900` 年前 LMT 分层。
|
||||
- 关键结果(raw,全量正确率 / 真值段保留;LOO 选阈值后的 eligible、accuracy、retained):
|
||||
|
||||
| 半径 | 盘 | 全量 top 段正确 | 全量真值段保留 | 全量 ≤2 段 | LOO eligible / accuracy / retained |
|
||||
|---|---|---:|---:|---:|---:|
|
||||
| ±10 | D1 | 76/77 | 76/77 | 77/77 | 77 / 0.987 / 0.987 |
|
||||
| ±10 | D9 | 71/77 | 76/77 | 68/77 | 23 / 0.913 / 0.957 |
|
||||
| ±10 | D10 | 69/77 | 76/77 | 63/77 | 29 / 0.931 / 0.966 |
|
||||
| ±15 | D1 | 74/77 | 75/77 | 77/77 | 77 / 0.961 / 0.974 |
|
||||
| ±15 | D9 | 61/77 | 75/77 | 50/77 | 14 / 0.929 / 0.929 |
|
||||
| ±15 | D10 | 59/77 | 76/77 | 52/77 | 8 / 0.875 / 0.875 |
|
||||
| ±30 | D1 | 75/77 | 76/77 | 77/77 | 75 / 0.973 / 0.987 |
|
||||
| ±30 | D9 | 50/77 | 76/77 | 25/77 | 11 / 1.000 / 1.000 |
|
||||
| ±30 | D10 | 52/77 | 76/77 | 25/77 | 16 / 1.000 / 1.000 |
|
||||
| ±60 | D1 | 68/77 | 76/77 | 76/77 | 54 / 0.944 / 0.981 |
|
||||
| ±60 | D9 | 41/77 | 74/77 | 11/77 | 17 / 0.588 / 0.941 |
|
||||
| ±60 | D10 | 34/77 | 74/77 | 10/77 | 9 / 0.778 / 1.000 |
|
||||
|
||||
LOO 的高准确率部分来自只交付少量 eligible 案例,不能当全集准确率;±30 以上分盘的覆盖和并列退化明显。三种模式完整逐案例结果在机器 JSON,uniform 的并列数显著更高;LMT 单列,不能剔除。
|
||||
|
||||
- M0 第二次全量字节复跑未完成:`m0-full-rerun.log` 仅约 103 行,目标 rerun JSON 不存在,不能宣称逐字节一致;首轮 M0 全量仍是唯一完整复现证据。
|
||||
|
||||
|
||||
## M2 · not_started
|
||||
|
||||
尚未执行段级信息增益排序、六题与早停分开结果、候选保留集合逐格比较、wrong1/wrong2 和日期 ±7 稳健性。`refresh=False` 首轮探针为空的案例必须保留为空,不能用刷新探针补齐。
|
||||
|
||||
## M3 · not_started
|
||||
|
||||
尚未执行事业(D1+D10)、婚恋(D1+D9)、综合(D1+D9+D10)三种取盘策略,也未统计各半径目标盘只有 1 段的“不用校正”比例。
|
||||
|
||||
## M4 · pending
|
||||
|
||||
已创建阶段性报告 `docs/research/rectification_varga_resolution_2026_09_30.md`,含候选档位、诚实出口与实现前置条件;M2/M3 无实测,尚未形成最终研究结论或实现单。不得把研究候选档位当成已校准的产品阈值。
|
||||
|
||||
## 运行命令与证据
|
||||
|
||||
- M0 全量:
|
||||
`PYTHONUTF8=1 PYTHONHASHSEED=0 python scripts/research/varga_resolution_probe.py --radii 10,15,30,60 --json-out docs/research/varga_resolution_baseline_2026_09_30.json`
|
||||
- M0 日志:`artifacts/varga-resolution/m0-full.log`
|
||||
- 首轮 M0 定向测试:`PYTHONUTF8=1 python -m pytest -q tests/test_varga_resolution_research.py`,当时 4 passed;当前收尾复验为 77 passed,见下文。
|
||||
- 既有基线定向测试:`artifacts/varga-resolution/baseline-targeted.log`,exit 0。
|
||||
- 快速门:`artifacts/varga-resolution/baseline-quick.log`;需以命令实际 exit 与日志内容报告,不能把开工预检当作通过。
|
||||
- 开工预检:`artifacts/varga-resolution/pre-work.log`,status=`fail`;失败原因是仓库环境的 preflight fragment-scan 合同缺少预期 `.workbuddy/skills/jyotish-vedic-astrology` 镜像路径。该环境缺口不归因于本研究,也不能伪称通过。
|
||||
|
||||
## 本轮收尾复验(2026-09-30)
|
||||
|
||||
- 报告 ±30 D1 的 LOO 数字已从误写的 `75 / 1.000 / 1.000` 修正为 `75 / 0.973 / 0.987`,精确值为 0.97333333 / 0.98666667;新增测试逐行对比 M1 报表全部 12 行与机器 JSON。
|
||||
- `PYTHONUTF8=1 PYTHONHASHSEED=0 python -m pytest -o addopts='' -q tests/test_varga_resolution_research.py`:**77 passed,exit 0**(主会话独立复验 1.19s),日志 `artifacts/varga-resolution/closure-targeted.log`。原 4 个测试断言不变,新增合成算术、空训练、LOO held-out 身份隔离、并列、LMT、308 唯一 case-radius、36 组汇总重算测试;没有重跑引擎,不证明全量字节复现。
|
||||
- M0/M1 `--help` 均 exit 0,日志 `closure-m0-help.log`、`closure-m1-help.log`。
|
||||
- `PYTHONUTF8=1 PYTHONHASHSEED=0 python -m pytest -o addopts='' -q tests/test_repo_privacy_markers.py`:**63 passed,exit 0**,日志 `closure-privacy.log`。只覆盖当前 tracked repository;新增研究文件仍未跟踪,不能据此宣称产物隐私全绿。
|
||||
- 对全部本轮脚本、文档、测试及研究 JSON/log 复用现有 `build_rules` / `scan_text` 显式补扫:**未通过**。29 个成功解码文件中 M0 baseline JSON 有 R003 共 3 次命中(扫描时刻字符串字段,未输出标记值);旧 `baseline-quick.log` UTF-8 与 GB18030 严格解码均失败。初次 UTF-8-only 扫描 exit 1,带严格回退的补扫也 exit 1,证据 `closure-explicit-privacy.log`。未删改研究证据或放宽隐私例外;提交前需独立核查这些命中并完成原日志的保真扫描。
|
||||
- 回放不是每例答满六题:308 个 case-radius 中无探针 36、零回答 38,实际 0–6 题分布为 38/6/5/9/11/16/223。全部保留分母。top 命中包含并列;percent 为后验质量归一化,不是独立概率校准。报告已补真实定义与局限。
|
||||
- M1 metadata 的确定性标志不构成复跑证据;M0 任务书起点表逐格对齐仍欠。M2/M3 未执行,BUG-1105 保持 investigating,索引已同步为部分完成。
|
||||
- 最终完整 quick 补验使用本机 Python 3.11.7,日志 `artifacts/varga-resolution/closure-quick.log`,**exit 1**。Python 子段为 **1000 passed、1 failed、1 skipped、3 subtests passed(473.06s)**;唯一列出的 pytest 失败是 `tests/test_consultation_native_layers.py::test_the_new_layer_leaves_every_existing_output_unchanged`,在去除既有 volatile 字段后比较两次输出的 JSON 等价性断言失败。日志差异片段涉及 VedAstro 请求清单,但详情截断,根因未确定。本次在 Python 步骤停止,不以开工时的 tsx 失败解释本次失败,不能声称与基线逐项一致。
|
||||
- quick 运行期间 5 个 tracked oracle pending packet JSON 变脏。独立只读核查确认:全部 JSON 与 HEAD 语义相同,归一换行后字节相同,只有 LF→CRLF。写入链为 `run_quality_gate.py` → `test_external_oracle_sanity_closure.py` → `external_oracle_sanity_closure.py` → `oracle_closure_master_dashboard.py` → `tajika_annual_closure_status.py` → `tajika_annual_oracle_queue.py:272–284` 的文本模式写入。未恢复或暂存这些非计划文件,排除在本研究候选交付之外;不能把这一副作用当作上条测试失败的已证根因。
|
||||
- 独立验收复核全部 M1 表格和回答分布与 JSON 一致,另跑研究定向 **77 passed,exit 0(1.34s)**;只发现本页一处残留“当前 4 passed”,已改为首轮历史值。复核不覆盖完整引擎二次复跑,也不推翻 quick / 隐私补扫未通过的结论。
|
||||
|
||||
### 最终文档收口复验
|
||||
|
||||
- 最终文档更新后合跑研究定向与 tracked 隐私守卫:**140 passed,exit 0(6.63s)**,即 77 项研究测试 + 63 项隐私测试;日志 `artifacts/varga-resolution/closure-final-checks.log`。
|
||||
- 对本轮 7 份修改文档(含未跟踪报告/进度及错误台账)复用现有隐私规则显式扫描:**0 findings**。此结论仅覆盖这 7 份文档,不覆盖全部研究 JSON/log,不解除此前产物补扫的 3 处命中及旧日志解码缺口。
|
||||
- 本轮 tracked 文档 `git diff --check` exit 0;全工作树另有已确认的 5 个 oracle 换行副作用,未恢复、未暂存。生产代码、前端、线上评分权重和门槛未改,未提交、未推送、未部署。完整 quick 仍为未通过,整单仍未验收。
|
||||
|
||||
## 未完成原因与恢复点
|
||||
|
||||
第二次全量 M0 复跑已于本轮中止,**未完成**:`artifacts/varga-resolution/m0-full-rerun.log` 仅有 103 行(已输出 约 26 例×4 档中的前三档,远少于完整 77×4=308 个 case-radius 行),`artifacts/varga-resolution/varga_resolution_baseline_rerun.json` 不存在,因此没有第二次复跑 exit/hash 可报告,也没有逐字节一致性证据。当前没有残留 Python 进程。不得把该日志算作全量通过;首跑的 `m0-full.log`/baseline JSON 才是 M0 全量证据。后续若恢复,须从头完成 77×4 并再比较 hash;本轮不再等待。M1 后续已完成 308 条全量分析,见上文,不再沿用“M1 未开始”的旧状态。M2/M3 为 not_started、M4 为部分完成;后续按原任务书让步顺序 M0 > M1 > M3 > M2 推进,不能以当前文档收尾替代研究验收。
|
||||
@@ -257,7 +257,7 @@
|
||||
| `TASK-rectification-futile-collect-stop-20260929.md` | `PROGRESS-rectification-futile-collect-stop-20260929.md` | **生时校正停掉无效补经历循环(止血)**:打字经历不收窄(BUG-560 后果),流程却一路索要,真机 22 件整窗不动;采集只为开闸、门开后只问点选卡问完即出卡;交付正文去吻合率;多段时旁白误报「范围没变」;交付轮带采集题 | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1084~1087 |
|
||||
| `TASK-rectification-holdout-expansion-20260929.md` | `PROGRESS-rectification-holdout-expansion-20260929.md` | **开放评价集扩到 ≥60 例(v5)**:所有打分判定都在 20 例上做,1 例 = 5pp,KP / 精度闸 / V1n 的 no_benefit 都只差 1–3 例。只用 Rodden AA 公开名人,事件人工核对出处,分层(年代 / 纬度 / 南半球 / 跨午夜 / UTC+8),v4 不动;基线成绩单 v4 子集须与已发布数字逐格一致。研究单的判定以 v5 为准 | 已验收(Claude 09-29:77 例、v4 子集数字逐格一致、抽样 6 例来源一致、冻结文件零改动;UTC+8 5/6 为数据可得性缺口,接受;已合入 staging)→ 研究单判定可开工 | 分支 `codex/rectification-holdout-expansion-20260929`(BUG-1090) |
|
||||
| `TASK-rectification-scoring-research-20260929.md` | `PROGRESS-rectification-scoring-research-20260929.md` | **打分方法研究(离线)**:R-A 似然比校准权重(leave-one-case-out,KP / Pranapada / D60 作为特征由数据定权重、允许负权重)、R-B 缺席证据(口述时间线覆盖时段内的空白年扣分,必测漏说稳健性)、R-C 精度追问可行性(对已说事件追问月份,算在定向追问 2 条额度内)。判定必须在 v5 上做;有收益才立实现单并按 ERR-110 重冻结 | 已验收关单(Claude 09-30:v5 77 例三项全部不过硬红线 1,稳定特征 0 个;BUG-1091 closed_by_design;打分路线关闭,后续走产品口径「盘型选择」研究) | 分支 `codex/rectification-scoring-research-20260929`(BUG-1091 closed_by_design) |
|
||||
| `TASK-rectification-varga-resolution-research-20260930.md` | — | **「盘型口径」研究(离线)**:打分层关闭后,把校正目标从分钟改为分盘上升段。v5 实测:±10 六题后 D1 单一 75/77,D9/D10 ≤2 种 68/63,多数=真值 88%/84%;±30 起分盘分不开。M0 入库段扫描与区间脚本、M1 段级汇总与阈值-准确率(留一法)、M2 按段选题、M3 按问题域取盘 + 「不用校正」比例、M4 实现单要点。不改引擎 / 常数 / 冻结文件;红线:真值段不被排除 ≥ 现在真值在区间比例 | 待领取 | — |
|
||||
| `TASK-rectification-varga-resolution-research-20260930.md` | [PROGRESS](PROGRESS-rectification-varga-resolution-research-20260930.md) | **「盘型口径」研究(BUG-1105,离线)**:M0 首轮、M1 全量各 308 个 case-radius;±60 D9/D10 真值段保留各 74/77 < 基线 75/77,不通过红线;LOO 高准确率仅在小覆盖子集成立。生产零改动 | 执行中(部分完成;M0 逐格一致性/二次复跑未闭环,M2/M3 not_started,M4 阶段性) | `codex/rectification-varga-resolution-research-20260930`;用户要求归档到 staging,提交子集与排除项见 PROGRESS(不代表验收);[研究报告](../research/rectification_varga_resolution_2026_09_30.md) |
|
||||
| `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 已验收(Claude 2026-09-29 合并验收:全量前端 4302/24 与基线同 24 条环境失败、tsc 0、lint 0 error、`/` Static、gzip +0.16%、快速门 pytest 1001 passed、两份回放复跑一致),待部署核对 | BUG-1088、1089 |
|
||||
| `TASK-report-reader-polish-20260929.md` | `PROGRESS-report-reader-polish-20260929.md` | **报告页打磨**:生成入口挪进页面主体(删标题栏按钮)、详情页到底部按钮(懒渲染一次到底)、导出按钮带文字、目录一级/二级分层、去掉「字段」与 RL/NL 缩写表头、状态列同义重复去重、报告表格淡底色、分块导出显示真实文件大小 | Claude 直接执行并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +7 B、快速门 Python 1000 passed),已部署 staging `d7772011`,真机欠 | BUG-1092~1094 |
|
||||
| `TASK-report-english-edition-20260929.md` | `PROGRESS-report-english-edition-20260929.md` | **报告中英两版**:同一次引擎计算渲染 zh/en 两遍(不用模型翻译),瑜伽库 477 条补英文、模板与前端表头英文化;中文逐字节不变、英文零汉字、两版数字序列一致;阅读页 `?lang=en` 切换,导出当前语言 + 「问 AI 建议导出英文版」提示;旧报告不补英文 | Claude 直接执行(含两个 fork 子代理)并自验(tsc 0、lint 0 error、全量 fail 与基线同 24 条、`/` Static、gzip +0.06%、Python 定向全绿),已部署 staging `d7772011`,真机欠 | — |
|
||||
|
||||
@@ -0,0 +1,302 @@
|
||||
"""Offline segment-oriented rectification research for BUG-1105.
|
||||
|
||||
The production candidate scorer is only used as an observation source. This
|
||||
module never changes production defaults. A segment is a maximal contiguous
|
||||
run of equal divisional ascendant sign; equal signs separated by another sign
|
||||
remain different segments.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from collections import defaultdict
|
||||
from dataclasses import dataclass
|
||||
from datetime import date
|
||||
from typing import Any, Iterable, Sequence
|
||||
|
||||
from scripts.active_rectification_event_engine import compute_candidate_static_contexts
|
||||
from scripts.rectification.refinement_packet import window_scan
|
||||
from scripts.rectification.scoring_service import build_event_contribution_matrix, score_from_matrix
|
||||
from scripts.research.minute_resolution_sweep import scoring_request_for
|
||||
from scripts.research.probe_supply_after_six import ASK_COUNT, apply_answer, optimal_answer
|
||||
from scripts.research.scoring_research_lib import (
|
||||
public_for,
|
||||
replay,
|
||||
reconcile_rows,
|
||||
score_map,
|
||||
truth_cluster_times,
|
||||
)
|
||||
from scripts.research.cluster_width_lib import SEPARATION_LEAD, delivery_from_public, still_valid_public
|
||||
from scripts.rectification.event_probes import discriminating_event_probes
|
||||
|
||||
VARGA_PREFIXES = ("D1", "D9", "D10")
|
||||
RADII = (10, 15, 30, 60)
|
||||
THRESHOLDS = (0.5, 0.6, 0.7, 0.8, 0.9)
|
||||
TODAY = date(2026, 9, 16)
|
||||
|
||||
|
||||
def clock(stamp: str) -> int:
|
||||
hour, minute = str(stamp)[:5].split(":")
|
||||
return int(hour) * 60 + int(minute)
|
||||
|
||||
|
||||
def sign_value(row: dict[str, Any], prefix: str) -> Any:
|
||||
value = row.get(prefix)
|
||||
if isinstance(value, dict):
|
||||
return value.get("sign_idx", value.get("sign"))
|
||||
return value
|
||||
|
||||
|
||||
def segment_rows(rows: Sequence[dict[str, Any]], prefix: str) -> list[dict[str, Any]]:
|
||||
"""Return maximal sampled runs without merging a non-contiguous sign.
|
||||
|
||||
Clock time is cyclic here: 23:59 followed by 00:00 is one minute apart.
|
||||
A missing sample (for example 23:59 followed by 00:01) still starts a new
|
||||
run. This deliberately uses the observation order rather than a set of
|
||||
signs, so A-B-A produces three separately numbered segments.
|
||||
"""
|
||||
segments: list[dict[str, Any]] = []
|
||||
current: dict[str, Any] | None = None
|
||||
for row in rows:
|
||||
stamp = str(row.get("time") or row.get("stamp") or "")[:5]
|
||||
value = sign_value(row, prefix)
|
||||
if not stamp or value is None:
|
||||
current = None
|
||||
continue
|
||||
minute = clock(stamp)
|
||||
contiguous = (
|
||||
current is not None
|
||||
and (minute - int(current["last_minute"])) % 1440 == 1
|
||||
)
|
||||
if current is None or current["value"] != value or not contiguous:
|
||||
current = {
|
||||
"segment_id": len(segments),
|
||||
"varga": prefix,
|
||||
"value": value,
|
||||
"start": stamp,
|
||||
"end": stamp,
|
||||
"times": [stamp],
|
||||
"last_minute": minute,
|
||||
}
|
||||
segments.append(current)
|
||||
else:
|
||||
current["end"] = stamp
|
||||
current["times"].append(stamp)
|
||||
current["last_minute"] = minute
|
||||
for segment in segments:
|
||||
segment.pop("last_minute", None)
|
||||
return segments
|
||||
|
||||
|
||||
def segment_members(rows: Sequence[dict[str, Any]], prefix: str) -> list[list[str]]:
|
||||
return [list(item["times"]) for item in segment_rows(rows, prefix)]
|
||||
|
||||
|
||||
def row_signature(context: dict[str, Any], prefix: str) -> int | str | None:
|
||||
if prefix == "D1":
|
||||
value = context.get("ascendant_index")
|
||||
else:
|
||||
charts = context.get("varga_charts") or {}
|
||||
value = ((charts.get(prefix) or {}).get("Ascendant") or {}).get("sign_idx")
|
||||
return value if isinstance(value, (int, str)) else None
|
||||
|
||||
|
||||
def contexts_to_rows(contexts: Sequence[dict[str, Any]], prefixes: Iterable[str] = VARGA_PREFIXES) -> list[dict[str, Any]]:
|
||||
rows: list[dict[str, Any]] = []
|
||||
for context in contexts:
|
||||
feature = context.get("feature")
|
||||
stamp = feature.get("time") if isinstance(feature, dict) else None
|
||||
if not stamp:
|
||||
stamp = context["candidate_at"].strftime("%H:%M")
|
||||
rows.append({"time": str(stamp)[:5], **{p: row_signature(context, p) for p in prefixes}})
|
||||
return rows
|
||||
|
||||
|
||||
def truth_segment(rows: Sequence[dict[str, Any]], prefix: str, truth_time: str) -> dict[str, Any] | None:
|
||||
for segment in segment_rows(rows, prefix):
|
||||
if truth_time[:5] in segment["times"]:
|
||||
return segment
|
||||
return None
|
||||
|
||||
|
||||
def unique_sign_count(rows: Sequence[dict[str, Any]], prefix: str) -> int:
|
||||
return len({sign_value(row, prefix) for row in rows if sign_value(row, prefix) is not None})
|
||||
|
||||
|
||||
def unique_segment_count(rows: Sequence[dict[str, Any]], prefix: str) -> int:
|
||||
return len(segment_rows(rows, prefix))
|
||||
|
||||
|
||||
def window_payload(case: dict[str, Any], radius: int) -> tuple[dict[str, Any], list[dict[str, Any]], list[dict[str, Any]]]:
|
||||
"""Build the two-minute scoring grid and an independent one-minute scan grid."""
|
||||
scoring_request = scoring_request_for({**case, "candidate_radius_minutes": radius}, radius)
|
||||
scoring_request["minute_step"] = 2
|
||||
scoring_contexts = compute_candidate_static_contexts(scoring_request)
|
||||
scan_request = scoring_request_for({**case, "candidate_radius_minutes": radius}, radius)
|
||||
scan_request["minute_step"] = 1
|
||||
scan_contexts = compute_candidate_static_contexts(scan_request)
|
||||
return scoring_request, scoring_contexts, contexts_to_rows(scan_contexts)
|
||||
|
||||
|
||||
def _probe_payload(request: dict[str, Any], built: dict[str, Any], times: Sequence[str], true_time: str) -> list[dict[str, Any]]:
|
||||
return discriminating_event_probes(
|
||||
{**request, "refresh_probes": False, "asked_probe_keys": []},
|
||||
built,
|
||||
scan=window_scan(built),
|
||||
candidate_times=list(times),
|
||||
representative_time=true_time,
|
||||
today=TODAY,
|
||||
)
|
||||
|
||||
|
||||
def replay_state(rows: Sequence[dict[str, Any]], contexts: Sequence[dict[str, Any]], request: dict[str, Any], true_time: str, *, probes: Sequence[dict[str, Any]] | None = None, answers: Sequence[str | None] | None = None) -> dict[str, Any]:
|
||||
times = [str(row["time"])[:5] for row in rows]
|
||||
public = public_for(rows, contexts)
|
||||
reps = [str(row["time"])[:5] for row in public]
|
||||
scores = {stamp: float(row.get("score") or 0) for stamp, row in ((str(item["time"])[:5], item) for item in public)}
|
||||
conflicts = {stamp: 0 for stamp in reps}
|
||||
eliminated: set[str] = set()
|
||||
actual_probes = list(probes) if probes is not None else _probe_payload(request, {"static_contexts": list(contexts), "rows": list(rows)}, times, true_time)
|
||||
given = list(answers) if answers is not None else [optimal_answer(p, true_time) for p in actual_probes[:ASK_COUNT]]
|
||||
for probe, answer in zip(actual_probes[:ASK_COUNT], given):
|
||||
if answer in {"yes", "no", "weak_yes"}:
|
||||
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, answer, reps)
|
||||
posterior = [{**row, "score": scores.get(str(row["time"])[:5], row.get("score") or 0)} for row in public]
|
||||
valid = still_valid_public(posterior, scores, eliminated, lead=SEPARATION_LEAD)
|
||||
return {
|
||||
"result": {"questions": sum(answer is not None for answer in given)},
|
||||
"public": public,
|
||||
"posterior": posterior,
|
||||
"valid": valid,
|
||||
"scores": scores,
|
||||
"eliminated": eliminated,
|
||||
"probes": actual_probes,
|
||||
"delivery": delivery_from_public(valid),
|
||||
}
|
||||
|
||||
|
||||
def native_case(case: dict[str, Any], radius: int, *, do_reconcile: bool = False) -> dict[str, Any]:
|
||||
request, contexts, chart_rows = window_payload(case, radius)
|
||||
true_time = str(case["birth"]["time"])[:5]
|
||||
built = build_event_contribution_matrix(request, static_contexts=contexts)
|
||||
rows = score_from_matrix(request, built)
|
||||
times = [str(row["time"])[:5] for row in rows]
|
||||
probes = _probe_payload(request, built, times, true_time)
|
||||
state = replay_state(rows, contexts, request, true_time, probes=probes)
|
||||
reconciliation = {"status": "not_run"}
|
||||
if do_reconcile:
|
||||
from scripts.research.scoring_research_lib import FeatureStore, d60_charts, make_recording_provider
|
||||
store = FeatureStore()
|
||||
recording = build_event_contribution_matrix(request, row_provider=make_recording_provider(contexts, store, d60_charts(contexts)), static_contexts=contexts)
|
||||
recording_rows = score_from_matrix(request, recording)
|
||||
reconciliation = reconcile_rows(rows, recording_rows)
|
||||
return {"case_id": str(case["case_id"]), "radius": radius, "true_time": true_time, "request": request, "contexts": contexts, "chart_rows": chart_rows, "rows": rows, "probes": probes, "state": state, "reconciliation": reconciliation}
|
||||
|
||||
|
||||
def valid_minute_scores(
|
||||
state: dict[str, Any],
|
||||
chart_rows: Sequence[dict[str, Any]],
|
||||
scores: dict[str, float] | None = None,
|
||||
) -> dict[str, float]:
|
||||
"""Project representative posterior scores onto each minute in its cluster."""
|
||||
effective_scores = scores if scores is not None else state["scores"]
|
||||
output: dict[str, float] = {}
|
||||
for row in state["posterior"]:
|
||||
representative = str(row["time"])[:5]
|
||||
value = float(effective_scores.get(representative, row.get("score") or 0))
|
||||
for stamp in row.get("cluster_times") or [representative]:
|
||||
output[str(stamp)[:5]] = value
|
||||
return output
|
||||
|
||||
|
||||
def segment_metrics(state: dict[str, Any], chart_rows: Sequence[dict[str, Any]], prefix: str, true_time: str, mode: str) -> dict[str, Any]:
|
||||
segments = segment_rows(chart_rows, prefix)
|
||||
scores = {str(key)[:5]: float(value) for key, value in state["scores"].items()}
|
||||
if mode == "percent":
|
||||
total = sum(max(value, 0.0) for value in scores.values())
|
||||
scores = (
|
||||
{key: max(value, 0.0) / total * 100.0 for key, value in scores.items()}
|
||||
if total > 0
|
||||
else {key: 0.0 for key in scores}
|
||||
)
|
||||
minute_scores = valid_minute_scores(state, chart_rows, scores)
|
||||
valid_times = {str(t)[:5] for row in state["valid"] for t in (row.get("cluster_times") or [row.get("time")])}
|
||||
truth = truth_segment(chart_rows, prefix, true_time)
|
||||
qualities: list[float] = []
|
||||
for segment in segments:
|
||||
values = [minute_scores.get(t, 0.0) for t in segment["times"] if t in valid_times]
|
||||
if mode == "uniform":
|
||||
quality = float(len(values))
|
||||
elif mode == "percent":
|
||||
quality = sum(max(v, 0.0) for v in values)
|
||||
else:
|
||||
quality = sum(values)
|
||||
qualities.append(quality)
|
||||
total = sum(qualities)
|
||||
share = max(qualities) / total if qualities and total > 0 else None
|
||||
leaders = [i for i, value in enumerate(qualities) if share is not None and abs(value - max(qualities)) <= 1e-9]
|
||||
truth_id = truth["segment_id"] if truth else None
|
||||
retained = truth_id is not None and truth_id in {segment["segment_id"] for segment in segments if any(t in valid_times for t in segment["times"])}
|
||||
correct = truth_id is not None and truth_id in leaders
|
||||
valid_segment_count = sum(1 for segment in segments if any(t in valid_times for t in segment["times"]))
|
||||
return {
|
||||
"prefix": prefix,
|
||||
"mode": mode,
|
||||
"segment_count_window": len(segments),
|
||||
"valid_segment_count": valid_segment_count,
|
||||
"truth_segment_id": truth_id,
|
||||
"truth_retained": bool(retained),
|
||||
"top_segment_correct": bool(correct),
|
||||
"top_segment_tie": len(leaders) > 1,
|
||||
"top_segment_ids": leaders,
|
||||
"top_share": None if share is None else round(share, 8),
|
||||
"segment_qualities": [round(v, 8) for v in qualities],
|
||||
}
|
||||
|
||||
|
||||
def threshold_scan(rows: Sequence[dict[str, Any]], thresholds: Sequence[float] = THRESHOLDS) -> dict[str, Any]:
|
||||
out: dict[str, Any] = {}
|
||||
for threshold in thresholds:
|
||||
eligible = [row for row in rows if row.get("top_share") is not None and float(row["top_share"]) >= threshold]
|
||||
out[str(threshold)] = {"n": len(eligible), "denominator": len(rows), "coverage": round(len(eligible) / len(rows), 8) if rows else None, "accuracy": round(sum(bool(row.get("top_segment_correct")) for row in eligible) / len(eligible), 8) if eligible else None, "truth_retained": round(sum(bool(row.get("truth_retained")) for row in eligible) / len(eligible), 8) if eligible else None}
|
||||
return out
|
||||
|
||||
|
||||
def choose_loo_threshold(training: Sequence[dict[str, Any]], thresholds: Sequence[float] = THRESHOLDS, minimum: int = 5) -> float | None:
|
||||
if not training:
|
||||
return None
|
||||
candidates = []
|
||||
required = min(minimum, len(training))
|
||||
for threshold in thresholds:
|
||||
eligible = [row for row in training if row.get("top_share") is not None and float(row["top_share"]) >= threshold]
|
||||
if len(eligible) < required or not eligible:
|
||||
continue
|
||||
accuracy = sum(bool(row.get("top_segment_correct")) for row in eligible) / len(eligible)
|
||||
retained = sum(bool(row.get("truth_retained")) for row in eligible) / len(eligible)
|
||||
candidates.append((accuracy, retained, len(eligible), -float(threshold), float(threshold)))
|
||||
if not candidates:
|
||||
return None
|
||||
return max(candidates)[-1]
|
||||
|
||||
|
||||
def segment_probe_score(probe: dict[str, Any], chart_rows: Sequence[dict[str, Any]], prefix: str, minute_weights: dict[str, float]) -> float:
|
||||
segments = segment_rows(chart_rows, prefix)
|
||||
segment_by_time = {t: segment["segment_id"] for segment in segments for t in segment["times"]}
|
||||
yes = {str(t)[:5] for item in probe.get("expected_outcomes") or [] if item.get("answer_class") in {"yes", "weak_yes"} for t in item.get("supports") or []}
|
||||
no = {str(t)[:5] for item in probe.get("expected_outcomes") or [] if item.get("answer_class") == "no" for t in item.get("supports") or []}
|
||||
totals: dict[int, float] = defaultdict(float)
|
||||
for stamp in yes | no:
|
||||
if stamp in segment_by_time:
|
||||
totals[segment_by_time[stamp]] += max(minute_weights.get(stamp, 0.0), 0.0)
|
||||
if len(totals) < 2:
|
||||
return 0.0
|
||||
values = sorted(totals.values(), reverse=True)
|
||||
return round(values[0] - values[1], 8)
|
||||
|
||||
|
||||
def reorder_probes_by_segments(probes: Sequence[dict[str, Any]], chart_rows: Sequence[dict[str, Any]], prefix: str, minute_weights: dict[str, float]) -> list[dict[str, Any]]:
|
||||
return sorted(enumerate(probes), key=lambda item: (-segment_probe_score(item[1], chart_rows, prefix, minute_weights), item[0])) and [item[1] for item in sorted(enumerate(probes), key=lambda item: (-segment_probe_score(item[1], chart_rows, prefix, minute_weights), item[0]))]
|
||||
|
||||
|
||||
def strategy_prefixes(domain: str) -> tuple[str, ...]:
|
||||
if domain == "career": return ("D1", "D10")
|
||||
if domain == "relationship": return ("D1", "D9")
|
||||
return VARGA_PREFIXES
|
||||
@@ -0,0 +1,221 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Offline M1 segment-quality and leave-one-out analysis for BUG-1105.
|
||||
|
||||
This deliberately reuses ``native_case`` and does not modify production
|
||||
rectification code. The output contains both full-sample fixed-threshold
|
||||
figures and case-level leave-one-out validation; no threshold is selected from
|
||||
the validation case itself.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any, Sequence
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.research.varga_resolution_lib import ( # noqa: E402
|
||||
RADII,
|
||||
THRESHOLDS,
|
||||
VARGA_PREFIXES,
|
||||
native_case,
|
||||
segment_metrics,
|
||||
choose_loo_threshold,
|
||||
threshold_scan,
|
||||
)
|
||||
|
||||
HOLDOUT = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v5.json"
|
||||
SCHEMA = "bug-1105-varga-resolution-m1-v1"
|
||||
MODES = ("raw", "percent", "uniform")
|
||||
|
||||
|
||||
def load_cases() -> list[dict[str, Any]]:
|
||||
return list(json.loads(HOLDOUT.read_text(encoding="utf-8")).get("cases") or [])
|
||||
|
||||
|
||||
def is_lmt(case: dict[str, Any]) -> bool:
|
||||
return int(str(case.get("birth", {}).get("date", "9999"))[:4]) < 1900
|
||||
|
||||
|
||||
def metrics_for_case(case: dict[str, Any], radius: int) -> dict[str, Any]:
|
||||
result = native_case(case, radius, do_reconcile=False)
|
||||
state = result["state"]
|
||||
rows = result["chart_rows"]
|
||||
true_time = result["true_time"]
|
||||
return {
|
||||
"case_id": str(case["case_id"]),
|
||||
"radius": radius,
|
||||
"lmt_before_1900": is_lmt(case),
|
||||
"answered_count": int((state.get("result") or {}).get("questions") or 0),
|
||||
"probe_count": len(result.get("probes") or []),
|
||||
"by_varga": {
|
||||
prefix: {
|
||||
mode: segment_metrics(state, rows, prefix, true_time, mode)
|
||||
for mode in MODES
|
||||
}
|
||||
for prefix in VARGA_PREFIXES
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def aggregate_rows(items: Sequence[dict[str, Any]], prefix: str, mode: str) -> dict[str, Any]:
|
||||
rows = [item["by_varga"][prefix][mode] for item in items]
|
||||
denominator = len(rows)
|
||||
return {
|
||||
"denominator": denominator,
|
||||
"top_share_mean": round(sum(float(r["top_share"] or 0) for r in rows) / denominator, 8) if denominator else None,
|
||||
"truth_retained": sum(bool(r["truth_retained"]) for r in rows),
|
||||
"truth_retained_rate": round(sum(bool(r["truth_retained"]) for r in rows) / denominator, 8) if denominator else None,
|
||||
"truth_excluded": sum(not bool(r["truth_retained"]) for r in rows),
|
||||
"top_segment_correct": sum(bool(r["top_segment_correct"]) for r in rows),
|
||||
"top_segment_correct_rate": round(sum(bool(r["top_segment_correct"]) for r in rows) / denominator, 8) if denominator else None,
|
||||
"top_segment_ties": sum(bool(r.get("top_segment_tie")) for r in rows),
|
||||
"top_segment_tie_rate": round(sum(bool(r.get("top_segment_tie")) for r in rows) / denominator, 8) if denominator else None,
|
||||
"valid_segment_count_le_2": sum(int(r["valid_segment_count"]) <= 2 for r in rows),
|
||||
"valid_segment_count_le_2_rate": round(sum(int(r["valid_segment_count"]) <= 2 for r in rows) / denominator, 8) if denominator else None,
|
||||
"thresholds_full_fit": threshold_scan(rows),
|
||||
}
|
||||
|
||||
|
||||
def stratified_rows(items: Sequence[dict[str, Any]], prefix: str, mode: str) -> dict[str, Any]:
|
||||
"""Keep the pre-1900 LMT stratum auditable without changing denominators."""
|
||||
return {
|
||||
"all": aggregate_rows(items, prefix, mode),
|
||||
"lmt_before_1900": aggregate_rows(
|
||||
[item for item in items if item["lmt_before_1900"]], prefix, mode
|
||||
),
|
||||
"post_1900": aggregate_rows(
|
||||
[item for item in items if not item["lmt_before_1900"]], prefix, mode
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def summarize_loo_rows(rows: Sequence[dict[str, Any]]) -> dict[str, Any]:
|
||||
eligible = [row for row in rows if row["eligible"]]
|
||||
return {
|
||||
"eligible": len(eligible),
|
||||
"denominator": len(rows),
|
||||
"coverage": round(len(eligible) / len(rows), 8) if rows else None,
|
||||
"accuracy": round(sum(bool(row["top_segment_correct"]) for row in eligible) / len(eligible), 8) if eligible else None,
|
||||
"truth_retained": round(sum(bool(row["truth_retained"]) for row in eligible) / len(eligible), 8) if eligible else None,
|
||||
}
|
||||
|
||||
|
||||
def loo_fixed(items: Sequence[dict[str, Any]], prefix: str, mode: str) -> dict[str, Any]:
|
||||
out: dict[str, Any] = {}
|
||||
for threshold in THRESHOLDS:
|
||||
selected = [item["by_varga"][prefix][mode] for item in items if float(item["by_varga"][prefix][mode]["top_share"] or 0) >= threshold]
|
||||
out[str(threshold)] = {
|
||||
"eligible": len(selected),
|
||||
"denominator": len(items),
|
||||
"coverage": round(len(selected) / len(items), 8) if items else None,
|
||||
"accuracy": round(sum(bool(row["top_segment_correct"]) for row in selected) / len(selected), 8) if selected else None,
|
||||
"truth_retained": round(sum(bool(row["truth_retained"]) for row in selected) / len(selected), 8) if selected else None,
|
||||
}
|
||||
return out
|
||||
|
||||
|
||||
def loo_selected(items: Sequence[dict[str, Any]], prefix: str, mode: str) -> dict[str, Any]:
|
||||
validations: list[dict[str, Any]] = []
|
||||
for index, item in enumerate(items):
|
||||
training = [other["by_varga"][prefix][mode] for j, other in enumerate(items) if j != index]
|
||||
threshold = choose_loo_threshold(training)
|
||||
row = item["by_varga"][prefix][mode]
|
||||
eligible = threshold is not None and float(row["top_share"] or 0) >= threshold
|
||||
validations.append({
|
||||
"case_id": item["case_id"],
|
||||
"threshold": threshold,
|
||||
"eligible": bool(eligible),
|
||||
"top_segment_correct": bool(row["top_segment_correct"]) if eligible else None,
|
||||
"truth_retained": bool(row["truth_retained"]) if eligible else None,
|
||||
"lmt_before_1900": item["lmt_before_1900"],
|
||||
})
|
||||
eligible = [row for row in validations if row["eligible"]]
|
||||
return {
|
||||
"eligible": len(eligible),
|
||||
"denominator": len(validations),
|
||||
"coverage": round(len(eligible) / len(validations), 8) if validations else None,
|
||||
"accuracy": round(sum(bool(row["top_segment_correct"]) for row in eligible) / len(eligible), 8) if eligible else None,
|
||||
"truth_retained": round(sum(bool(row["truth_retained"]) for row in eligible) / len(eligible), 8) if eligible else None,
|
||||
"selected_threshold_counts": {str(t): sum(row["threshold"] == t for row in validations) for t in THRESHOLDS},
|
||||
"validation_rows": validations,
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--radii", default=",".join(str(value) for value in RADII))
|
||||
parser.add_argument("--limit", type=int, default=0)
|
||||
parser.add_argument("--json-out", required=True)
|
||||
args = parser.parse_args()
|
||||
radii = tuple(int(value) for value in str(args.radii).split(",") if value.strip())
|
||||
cases = load_cases()
|
||||
if args.limit:
|
||||
cases = cases[: args.limit]
|
||||
items: list[dict[str, Any]] = []
|
||||
errors: list[dict[str, Any]] = []
|
||||
for case in cases:
|
||||
for radius in radii:
|
||||
label = f"{case.get('case_id')} ±{radius}"
|
||||
try:
|
||||
items.append(metrics_for_case(case, radius))
|
||||
print(label, flush=True)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
errors.append({"case_id": str(case.get("case_id")), "radius": radius, "error": f"{type(exc).__name__}: {exc}"})
|
||||
print(label, "ERROR", type(exc).__name__, exc, flush=True)
|
||||
aggregates: list[dict[str, Any]] = []
|
||||
for radius in radii:
|
||||
subset = [item for item in items if item["radius"] == radius]
|
||||
by_varga: dict[str, Any] = {}
|
||||
for prefix in VARGA_PREFIXES:
|
||||
by_varga[prefix] = {}
|
||||
for mode in MODES:
|
||||
loo = loo_selected(subset, prefix, mode)
|
||||
by_varga[prefix][mode] = {
|
||||
"full_fit": stratified_rows(subset, prefix, mode),
|
||||
"loo_fixed_thresholds": loo_fixed(subset, prefix, mode),
|
||||
"loo_selected_threshold": {
|
||||
**summarize_loo_rows(loo["validation_rows"]),
|
||||
"selected_threshold_counts": loo["selected_threshold_counts"],
|
||||
"validation_rows": loo["validation_rows"],
|
||||
"lmt_before_1900": summarize_loo_rows([row for row in loo["validation_rows"] if row["lmt_before_1900"]]),
|
||||
"post_1900": summarize_loo_rows([row for row in loo["validation_rows"] if not row["lmt_before_1900"]]),
|
||||
},
|
||||
}
|
||||
aggregates.append({
|
||||
"radius": radius,
|
||||
"case_count": len(subset),
|
||||
"lmt_case_count": sum(item["lmt_before_1900"] for item in subset),
|
||||
"by_varga": by_varga,
|
||||
})
|
||||
payload = {
|
||||
"schema": SCHEMA,
|
||||
"metadata": {
|
||||
"holdout": str(HOLDOUT.relative_to(ROOT)).replace("\\", "/"),
|
||||
"case_count_requested": len(cases),
|
||||
"case_count_completed": len({item["case_id"] for item in items}),
|
||||
"radii": list(radii),
|
||||
"vargas": list(VARGA_PREFIXES),
|
||||
"modes": list(MODES),
|
||||
"thresholds": list(THRESHOLDS),
|
||||
"loo": "one case held out; threshold selected only from the other cases; validation case never used for selection",
|
||||
"lmt_stratum": "birth year < 1900, reported separately",
|
||||
"production_code_modified": False,
|
||||
"deterministic_json": True,
|
||||
},
|
||||
"aggregates": aggregates,
|
||||
"items": items,
|
||||
"errors": errors,
|
||||
}
|
||||
out = Path(args.json_out)
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True) + "\n", encoding="utf-8")
|
||||
return 0 if not errors and len(items) == len(cases) * len(radii) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,349 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run the deterministic BUG-1105 varga-resolution M0 replay.
|
||||
|
||||
The production scorer remains untouched. Scoring candidates are sampled at a
|
||||
fixed two-minute step, while an independent one-minute context scan supplies
|
||||
varga rising-sign segments. ``refresh_probes=False`` is intentional and is
|
||||
recorded in the machine result; an empty probe pool is reported rather than
|
||||
silently replaced with refreshed dasha-boundary probes.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
from typing import Any, Iterable, Sequence
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
from scripts.research.cluster_width_lib import delivery_from_public # noqa: E402
|
||||
from scripts.research.varga_resolution_lib import ( # noqa: E402
|
||||
RADII,
|
||||
VARGA_PREFIXES,
|
||||
native_case,
|
||||
segment_rows,
|
||||
sign_value,
|
||||
)
|
||||
|
||||
HOLDOUT = ROOT / "references" / "real_case_calibration" / "minute_rectification_holdout_v5.json"
|
||||
SCHEMA = "bug-1105-varga-resolution-v1"
|
||||
SCORING_STEP = 2
|
||||
SCAN_STEP = 1
|
||||
REFRESH_PROBES = False
|
||||
|
||||
|
||||
def load_cases() -> list[dict[str, Any]]:
|
||||
payload = json.loads(HOLDOUT.read_text(encoding="utf-8"))
|
||||
return list(payload.get("cases") or [])
|
||||
|
||||
|
||||
def hhmm(value: object) -> str:
|
||||
return str(value or "")[:5]
|
||||
|
||||
|
||||
def clock(stamp: str) -> int:
|
||||
hour, minute = stamp[:5].split(":")
|
||||
return int(hour) * 60 + int(minute)
|
||||
|
||||
|
||||
def in_envelope(stamp: str, start: str | None, end: str | None) -> bool:
|
||||
if not start or not end:
|
||||
return False
|
||||
value, lower, upper = clock(stamp), clock(start), clock(end)
|
||||
return lower <= value <= upper if lower <= upper else value >= lower or value <= upper
|
||||
|
||||
|
||||
def stable_unique(values: Iterable[str]) -> list[str]:
|
||||
return list(dict.fromkeys(str(value)[:5] for value in values if value))
|
||||
|
||||
|
||||
def interval_times(chart_rows: Sequence[dict[str, Any]], delivery: dict[str, Any]) -> list[str]:
|
||||
return [
|
||||
hhmm(row.get("time"))
|
||||
for row in chart_rows
|
||||
if in_envelope(hhmm(row.get("time")), delivery.get("start"), delivery.get("end"))
|
||||
]
|
||||
|
||||
|
||||
def valid_candidate_times(state: dict[str, Any]) -> list[str]:
|
||||
values: list[str] = []
|
||||
for row in state.get("valid") or []:
|
||||
values.extend(hhmm(value) for value in row.get("cluster_times") or [row.get("time")])
|
||||
return stable_unique(values)
|
||||
|
||||
|
||||
def represented_segments(
|
||||
chart_rows: Sequence[dict[str, Any]],
|
||||
prefix: str,
|
||||
times: Iterable[str],
|
||||
) -> list[dict[str, Any]]:
|
||||
wanted = set(stable_unique(times))
|
||||
return [
|
||||
segment for segment in segment_rows(chart_rows, prefix)
|
||||
if wanted.intersection(segment.get("times") or [])
|
||||
]
|
||||
|
||||
|
||||
def represented_values(rows: Sequence[dict[str, Any]], prefix: str, times: Iterable[str]) -> list[Any]:
|
||||
wanted = set(stable_unique(times))
|
||||
return stable_values(sign_value(row, prefix) for row in rows if hhmm(row.get("time")) in wanted)
|
||||
|
||||
|
||||
def stable_values(values: Iterable[Any]) -> list[Any]:
|
||||
result: list[Any] = []
|
||||
for value in values:
|
||||
if value is None or value in result:
|
||||
continue
|
||||
result.append(value)
|
||||
return result
|
||||
|
||||
|
||||
def majority(values: Sequence[Any], truth: Any) -> tuple[bool | None, bool]:
|
||||
if not values:
|
||||
return None, False
|
||||
counts = Counter(values)
|
||||
peak = max(counts.values())
|
||||
leaders = {value for value, count in counts.items() if count == peak}
|
||||
return (truth in leaders if len(leaders) == 1 else None), len(leaders) > 1
|
||||
|
||||
|
||||
def window_summary(chart_rows: Sequence[dict[str, Any]]) -> dict[str, Any]:
|
||||
counts: dict[str, int] = {}
|
||||
segments: dict[str, int] = {}
|
||||
for prefix in VARGA_PREFIXES:
|
||||
counts[prefix] = len(stable_values(sign_value(row, prefix) for row in chart_rows))
|
||||
segments[prefix] = len(segment_rows(chart_rows, prefix))
|
||||
combinations = len({tuple(sign_value(row, prefix) for prefix in VARGA_PREFIXES) for row in chart_rows})
|
||||
return {
|
||||
"scan_point_count": len(chart_rows),
|
||||
"sign_counts": counts,
|
||||
"segment_counts": segments,
|
||||
"combination_count": combinations,
|
||||
"d1_single_sign": counts["D1"] == 1,
|
||||
}
|
||||
|
||||
|
||||
def interval_summary(
|
||||
chart_rows: Sequence[dict[str, Any]],
|
||||
state: dict[str, Any],
|
||||
true_time: str,
|
||||
) -> dict[str, Any]:
|
||||
delivery = state.get("delivery") or {}
|
||||
envelope = interval_times(chart_rows, delivery)
|
||||
real = valid_candidate_times(state)
|
||||
values: dict[str, dict[str, Any]] = {}
|
||||
for prefix in VARGA_PREFIXES:
|
||||
truth_row = next((row for row in chart_rows if hhmm(row.get("time")) == true_time), None)
|
||||
truth = sign_value(truth_row or {}, prefix)
|
||||
envelope_segments = represented_segments(chart_rows, prefix, envelope)
|
||||
real_segments = represented_segments(chart_rows, prefix, real)
|
||||
envelope_values = [sign_value(row, prefix) for row in chart_rows if hhmm(row.get("time")) in set(envelope)]
|
||||
exact_values = represented_values(chart_rows, prefix, real)
|
||||
envelope_majority, envelope_tie = majority(envelope_values, truth)
|
||||
exact_majority, exact_tie = majority(exact_values, truth)
|
||||
truth_segment = next(
|
||||
(segment for segment in segment_rows(chart_rows, prefix) if true_time in (segment.get("times") or [])),
|
||||
None,
|
||||
)
|
||||
truth_segment_id = truth_segment.get("segment_id") if truth_segment else None
|
||||
values[prefix] = {
|
||||
"truth_sign": truth,
|
||||
"truth_segment_id": truth_segment_id,
|
||||
"envelope_start": delivery.get("start"),
|
||||
"envelope_end": delivery.get("end"),
|
||||
"envelope_scan_point_count": len(envelope),
|
||||
"envelope_sign_count": len(stable_values(envelope_values)),
|
||||
"envelope_segment_count": len(envelope_segments),
|
||||
"envelope_majority_truth": envelope_majority,
|
||||
"envelope_majority_tie": envelope_tie,
|
||||
"real_valid_candidate_count": len(real),
|
||||
"real_valid_sign_count": len(exact_values),
|
||||
"real_valid_segment_count": len(real_segments),
|
||||
"real_valid_majority_truth": exact_majority,
|
||||
"real_valid_majority_tie": exact_tie,
|
||||
"truth_segment_retained_in_real_set": truth_segment_id is not None and any(
|
||||
segment.get("segment_id") == truth_segment_id for segment in real_segments
|
||||
),
|
||||
}
|
||||
combo_envelope = {
|
||||
tuple(sign_value(row, prefix) for prefix in VARGA_PREFIXES)
|
||||
for row in chart_rows
|
||||
if hhmm(row.get("time")) in set(envelope)
|
||||
}
|
||||
combo_real = {
|
||||
tuple(sign_value(row, prefix) for prefix in VARGA_PREFIXES)
|
||||
for row in chart_rows
|
||||
if hhmm(row.get("time")) in set(real)
|
||||
}
|
||||
truth_row = next((row for row in chart_rows if hhmm(row.get("time")) == true_time), None)
|
||||
truth_combo = tuple(sign_value(truth_row or {}, prefix) for prefix in VARGA_PREFIXES)
|
||||
envelope_combo_majority, envelope_combo_tie = majority(
|
||||
[tuple(sign_value(row, prefix) for prefix in VARGA_PREFIXES) for row in chart_rows if hhmm(row.get("time")) in set(envelope)],
|
||||
truth_combo,
|
||||
)
|
||||
return {
|
||||
"scoring_candidate_times": real,
|
||||
"interval_envelope": {
|
||||
"start": delivery.get("start"),
|
||||
"end": delivery.get("end"),
|
||||
"scan_point_count": len(envelope),
|
||||
"times": envelope,
|
||||
},
|
||||
"by_varga": values,
|
||||
"combination": {
|
||||
"envelope_count": len(combo_envelope),
|
||||
"real_valid_count": len(combo_real),
|
||||
"envelope_majority_truth": envelope_combo_majority,
|
||||
"envelope_majority_tie": envelope_combo_tie,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def case_result(case: dict[str, Any], radius: int) -> dict[str, Any]:
|
||||
result = native_case(case, radius, do_reconcile=True)
|
||||
state = result["state"]
|
||||
chart_rows = result["chart_rows"]
|
||||
return {
|
||||
"case_id": str(case.get("case_id") or ""),
|
||||
"radius": radius,
|
||||
"true_time": result["true_time"],
|
||||
"scoring_candidate_count": len(result["rows"]),
|
||||
"scan_point_count": len(chart_rows),
|
||||
"probe_count": len(result["probes"]),
|
||||
"answered_count": int((state.get("result") or {}).get("questions") or 0),
|
||||
"reconciliation": result["reconciliation"],
|
||||
"window": window_summary(chart_rows),
|
||||
"six_question_delivery": interval_summary(chart_rows, state, result["true_time"]),
|
||||
}
|
||||
|
||||
|
||||
def ratio(numerator: int, denominator: int) -> float | None:
|
||||
return round(numerator / denominator, 8) if denominator else None
|
||||
|
||||
|
||||
def aggregate(results: Sequence[dict[str, Any]], radius: int) -> dict[str, Any]:
|
||||
rows = [row for row in results if row["radius"] == radius]
|
||||
n = len(rows)
|
||||
window = {
|
||||
"case_count": n,
|
||||
"d1_single_sign": sum(row["window"]["d1_single_sign"] for row in rows),
|
||||
"mean_sign_counts": {
|
||||
prefix: round(sum(row["window"]["sign_counts"][prefix] for row in rows) / n, 8) if n else None
|
||||
for prefix in VARGA_PREFIXES
|
||||
},
|
||||
"mean_segment_counts": {
|
||||
prefix: round(sum(row["window"]["segment_counts"][prefix] for row in rows) / n, 8) if n else None
|
||||
for prefix in VARGA_PREFIXES
|
||||
},
|
||||
"mean_combination_count": round(sum(row["window"]["combination_count"] for row in rows) / n, 8) if n else None,
|
||||
}
|
||||
by_varga: dict[str, Any] = {}
|
||||
for prefix in VARGA_PREFIXES:
|
||||
items = [row["six_question_delivery"]["by_varga"][prefix] for row in rows]
|
||||
by_varga[prefix] = {
|
||||
"envelope_only_one_sign": sum(item["envelope_sign_count"] == 1 for item in items),
|
||||
"envelope_at_most_two_signs": sum(item["envelope_sign_count"] <= 2 for item in items),
|
||||
"envelope_majority_truth": sum(item["envelope_majority_truth"] is True for item in items),
|
||||
"envelope_majority_ties": sum(item["envelope_majority_tie"] for item in items),
|
||||
"real_truth_segment_retained": sum(item["truth_segment_retained_in_real_set"] for item in items),
|
||||
"real_at_most_two_segments": sum(item["real_valid_segment_count"] <= 2 for item in items),
|
||||
"denominator": n,
|
||||
}
|
||||
for key in (
|
||||
"envelope_only_one_sign",
|
||||
"envelope_at_most_two_signs",
|
||||
"envelope_majority_truth",
|
||||
"real_truth_segment_retained",
|
||||
"real_at_most_two_segments",
|
||||
):
|
||||
by_varga[prefix][f"{key}_rate"] = ratio(by_varga[prefix][key], n)
|
||||
combo = [row["six_question_delivery"]["combination"] for row in rows]
|
||||
return {
|
||||
"radius": radius,
|
||||
"case_count": n,
|
||||
"reconciliation": {
|
||||
"all_zero_diff": all(not row["reconciliation"].get("changed") for row in rows),
|
||||
"case_count": n,
|
||||
"changed_case_count": sum(bool(row["reconciliation"].get("changed")) for row in rows),
|
||||
"denominator": n,
|
||||
},
|
||||
"window": window,
|
||||
"delivery_envelope": {
|
||||
"by_varga": by_varga,
|
||||
"combination_only_one": sum(item["envelope_count"] == 1 for item in combo),
|
||||
"combination_at_most_two": sum(item["envelope_count"] <= 2 for item in combo),
|
||||
"combination_majority_truth": sum(item["envelope_majority_truth"] is True for item in combo),
|
||||
"combination_majority_ties": sum(item["envelope_majority_tie"] for item in combo),
|
||||
"denominator": n,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--radii", default=",".join(str(value) for value in RADII))
|
||||
parser.add_argument("--vargas", default=",".join(VARGA_PREFIXES))
|
||||
parser.add_argument("--limit", type=int, default=0)
|
||||
parser.add_argument("--json-out", required=True)
|
||||
args = parser.parse_args()
|
||||
radii = tuple(int(value) for value in str(args.radii).split(",") if value.strip())
|
||||
vargas = tuple(value.strip() for value in str(args.vargas).split(",") if value.strip())
|
||||
cases = load_cases()
|
||||
if args.limit:
|
||||
cases = cases[: args.limit]
|
||||
results: list[dict[str, Any]] = []
|
||||
errors: list[dict[str, str | int]] = []
|
||||
for case in cases:
|
||||
for radius in radii:
|
||||
label = f"{case.get('case_id')} ±{radius}"
|
||||
try:
|
||||
row = case_result(case, radius)
|
||||
results.append(row)
|
||||
print(
|
||||
f"{label} score={row['scoring_candidate_count']} scan={row['scan_point_count']} "
|
||||
f"probes={row['probe_count']} reconcile_changed={bool(row['reconciliation'].get('changed'))}",
|
||||
flush=True,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
errors.append({"case_id": str(case.get("case_id") or ""), "radius": radius, "error": f"{type(exc).__name__}: {exc}"})
|
||||
print(f"{label} ERROR {type(exc).__name__}: {exc}", flush=True)
|
||||
payload = {
|
||||
"schema": SCHEMA,
|
||||
"metadata": {
|
||||
"holdout": str(HOLDOUT.relative_to(ROOT)).replace("\\", "/"),
|
||||
"case_count_requested": len(cases),
|
||||
"case_count_completed": len({row["case_id"] for row in results}),
|
||||
"radii": list(radii),
|
||||
"vargas": list(vargas),
|
||||
"ayanamsa": "raman",
|
||||
"node_mode": "mean",
|
||||
"scoring_candidate_step_minutes": SCORING_STEP,
|
||||
"segment_scan_step_minutes": SCAN_STEP,
|
||||
"refresh_probes": REFRESH_PROBES,
|
||||
"replay_probe_source": "existing event_probes.discriminating_event_probes; no refresh",
|
||||
"questions_requested": 6,
|
||||
"questions_answered_is_recorded_per_case": True,
|
||||
"real_valid_candidate_set": "union of cluster_times from still_valid_public after replay; scoring-grid candidates only",
|
||||
"interval_envelope": "unionStillValidRange equivalent: min/max clock edges over the real valid candidate clusters; envelope is not the real set",
|
||||
"segment_definition": "maximal contiguous one-minute scan run of equal D1/D9/D10 sign; repeated non-contiguous signs retain separate IDs; 23:59->00:00 is contiguous",
|
||||
"aggregation_denominator": "77 cases per radius when full run completes; empty coverage and ties are separate counts",
|
||||
"reconciliation_scope": "every case and radius, scoring grid only; single-case zero-diff is smoke evidence, not full-set evidence",
|
||||
"deterministic_json": True,
|
||||
},
|
||||
"aggregates": [aggregate(results, radius) for radius in radii],
|
||||
"results": results,
|
||||
"errors": errors,
|
||||
}
|
||||
out = Path(args.json_out)
|
||||
out.parent.mkdir(parents=True, exist_ok=True)
|
||||
out.write_text(json.dumps(payload, ensure_ascii=False, indent=2, sort_keys=True) + "\n", encoding="utf-8")
|
||||
print(json.dumps(payload["aggregates"], ensure_ascii=False, indent=2, sort_keys=True), flush=True)
|
||||
return 0 if not errors and all(item["reconciliation"]["all_zero_diff"] for item in payload["aggregates"]) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,405 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from scripts.research import varga_resolution_m1 as m1
|
||||
from scripts.research.varga_resolution_lib import (
|
||||
RADII,
|
||||
THRESHOLDS,
|
||||
VARGA_PREFIXES,
|
||||
choose_loo_threshold,
|
||||
segment_metrics,
|
||||
segment_rows,
|
||||
threshold_scan,
|
||||
valid_minute_scores,
|
||||
)
|
||||
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
|
||||
|
||||
def test_segments_keep_non_contiguous_repeated_signs_separate() -> None:
|
||||
rows = [
|
||||
{"time": "12:00", "D9": 1},
|
||||
{"time": "12:01", "D9": 2},
|
||||
{"time": "12:02", "D9": 1},
|
||||
]
|
||||
segments = segment_rows(rows, "D9")
|
||||
assert [item["value"] for item in segments] == [1, 2, 1]
|
||||
assert [item["segment_id"] for item in segments] == [0, 1, 2]
|
||||
|
||||
|
||||
def test_segments_treat_midnight_as_contiguous() -> None:
|
||||
rows = [
|
||||
{"time": "23:59", "D10": 4},
|
||||
{"time": "00:00", "D10": 4},
|
||||
{"time": "00:01", "D10": 5},
|
||||
]
|
||||
segments = segment_rows(rows, "D10")
|
||||
assert len(segments) == 2
|
||||
assert segments[0]["times"] == ["23:59", "00:00"]
|
||||
assert segments[1]["times"] == ["00:01"]
|
||||
|
||||
|
||||
def test_segments_do_not_bridge_a_missing_minute() -> None:
|
||||
rows = [
|
||||
{"time": "23:59", "D1": 4},
|
||||
{"time": "00:01", "D1": 4},
|
||||
]
|
||||
segments = segment_rows(rows, "D1")
|
||||
assert len(segments) == 2
|
||||
|
||||
|
||||
def test_m0_metadata_declares_grid_and_set_envelope_distinction() -> None:
|
||||
payload = json.loads(
|
||||
(ROOT / "artifacts" / "varga-resolution" / "m0-smoke.json").read_text(encoding="utf-8")
|
||||
)
|
||||
metadata = payload["metadata"]
|
||||
assert metadata["scoring_candidate_step_minutes"] == 2
|
||||
assert metadata["segment_scan_step_minutes"] == 1
|
||||
assert metadata["refresh_probes"] is False
|
||||
assert "real_valid_candidate_set" in metadata
|
||||
assert "interval_envelope" in metadata
|
||||
|
||||
|
||||
# These are deliberately synthetic arithmetic fixtures, not engine-response
|
||||
# goldens or evidence of a deterministic full-engine repeat.
|
||||
def _metric(share, correct=True, retained=True, count=1, tie=False):
|
||||
return {
|
||||
"top_share": share,
|
||||
"top_segment_correct": correct,
|
||||
"truth_retained": retained,
|
||||
"valid_segment_count": count,
|
||||
"top_segment_tie": tie,
|
||||
}
|
||||
|
||||
|
||||
def _item(case_id, row, lmt=False):
|
||||
return {
|
||||
"case_id": case_id,
|
||||
"lmt_before_1900": lmt,
|
||||
"by_varga": {"D1": {"raw": row}},
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("training", "thresholds", "minimum", "expected"),
|
||||
[
|
||||
([], THRESHOLDS, 5, None),
|
||||
([_metric(None), _metric(0.49)], THRESHOLDS, 5, None),
|
||||
# A small fold requires all its rows, not an impossible five rows.
|
||||
([_metric(0.6), _metric(0.9)], THRESHOLDS, 5, 0.5),
|
||||
# The high-accuracy subset is too small at minimum=5, but not at 4.
|
||||
([_metric(0.9)] * 4 + [_metric(0.5, False)], THRESHOLDS, 5, 0.5),
|
||||
([_metric(0.9)] * 4 + [_metric(0.5, False)], THRESHOLDS, 4, 0.6),
|
||||
# Accuracy wins even when retention and coverage are worse.
|
||||
([_metric(0.9, True, False), _metric(0.5, False, True)], (0.5, 0.9), 1, 0.9),
|
||||
# Then retention, then coverage, then the lower threshold break ties.
|
||||
([_metric(0.9), _metric(0.5, True, False)], (0.5, 0.9), 1, 0.9),
|
||||
([_metric(0.9), _metric(0.5)], (0.9, 0.5), 1, 0.5),
|
||||
([_metric(0.9)], (0.9, 0.7, 0.5), 1, 0.5),
|
||||
([_metric(0.6)], (0.6,), 5, 0.6),
|
||||
],
|
||||
ids=["empty", "no-coverage", "small-fold", "minimum-five", "minimum-four",
|
||||
"accuracy-first", "retention-tiebreak", "coverage-tiebreak",
|
||||
"lower-threshold-tiebreak", "inclusive-boundary"],
|
||||
)
|
||||
def test_choose_loo_threshold(training, thresholds, minimum, expected) -> None:
|
||||
assert choose_loo_threshold(training, thresholds, minimum) == expected
|
||||
|
||||
|
||||
def test_loo_selected_excludes_each_held_out_identity(monkeypatch) -> None:
|
||||
rows = [_metric(0.5), _metric(0.8, False), _metric(0.9, True, False)]
|
||||
items = [_item(f"synthetic-{index}", row, index == 0) for index, row in enumerate(rows)]
|
||||
calls = []
|
||||
|
||||
def select(training):
|
||||
held_out = len(calls)
|
||||
expected = [row for index, row in enumerate(rows) if index != held_out]
|
||||
assert len(training) == len(expected)
|
||||
assert all(actual is wanted for actual, wanted in zip(training, expected))
|
||||
assert all(row is not rows[held_out] for row in training)
|
||||
calls.append(training)
|
||||
return (0.5, 0.9, 0.9)[held_out]
|
||||
|
||||
monkeypatch.setattr(m1, "choose_loo_threshold", select)
|
||||
actual = m1.loo_selected(items, "D1", "raw")
|
||||
assert len(calls) == 3
|
||||
assert actual == {
|
||||
"eligible": 2, "denominator": 3, "coverage": 0.66666667,
|
||||
"accuracy": 1.0, "truth_retained": 0.5,
|
||||
"selected_threshold_counts": {"0.5": 1, "0.6": 0, "0.7": 0, "0.8": 0, "0.9": 2},
|
||||
"validation_rows": [
|
||||
{"case_id": "synthetic-0", "threshold": 0.5, "eligible": True,
|
||||
"top_segment_correct": True, "truth_retained": True, "lmt_before_1900": True},
|
||||
{"case_id": "synthetic-1", "threshold": 0.9, "eligible": False,
|
||||
"top_segment_correct": None, "truth_retained": None, "lmt_before_1900": False},
|
||||
{"case_id": "synthetic-2", "threshold": 0.9, "eligible": True,
|
||||
"top_segment_correct": True, "truth_retained": False, "lmt_before_1900": False},
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def test_loo_selected_singleton_has_no_training_threshold() -> None:
|
||||
result = m1.loo_selected([_item("synthetic-only", _metric(1.0))], "D1", "raw")
|
||||
assert result == {
|
||||
"eligible": 0, "denominator": 1, "coverage": 0.0,
|
||||
"accuracy": None, "truth_retained": None,
|
||||
"selected_threshold_counts": {str(value): 0 for value in THRESHOLDS},
|
||||
"validation_rows": [{
|
||||
"case_id": "synthetic-only", "threshold": None, "eligible": False,
|
||||
"top_segment_correct": None, "truth_retained": None, "lmt_before_1900": False,
|
||||
}],
|
||||
}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def synthetic_cluster_state():
|
||||
# Two-minute clusters project their representative's score to BOTH minutes.
|
||||
# The high-scoring third cluster is posterior but not in the valid set.
|
||||
posterior = [
|
||||
{"time": "12:00", "cluster_times": ["12:00", "12:01"], "score": 999},
|
||||
{"time": "12:02", "score": 999},
|
||||
{"time": "12:03", "cluster_times": ["12:03", "12:04"], "score": 999},
|
||||
{"time": "12:05", "score": 999},
|
||||
]
|
||||
state = {
|
||||
"posterior": posterior,
|
||||
"valid": [posterior[0], posterior[1], posterior[3]],
|
||||
"scores": {"12:00": 2.0, "12:02": 3.0, "12:03": 100.0, "12:05": -1.0},
|
||||
}
|
||||
rows = [{"time": f"12:0{index}", "D1": sign} for index, sign in enumerate([1, 1, 2, 3, 3, 4])]
|
||||
return state, rows
|
||||
|
||||
|
||||
def test_minute_projection_includes_posterior_before_valid_filtering(synthetic_cluster_state) -> None:
|
||||
state, rows = synthetic_cluster_state
|
||||
assert valid_minute_scores(state, rows) == {
|
||||
"12:00": 2.0, "12:01": 2.0, "12:02": 3.0,
|
||||
"12:03": 100.0, "12:04": 100.0, "12:05": -1.0,
|
||||
}
|
||||
# Explicit scores override representatives, but missing keys use row.score.
|
||||
projected = valid_minute_scores(state, rows, {"12:00": 7.0})
|
||||
assert projected["12:00"] == projected["12:01"] == 7.0
|
||||
assert projected["12:02"] == 999.0
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("mode", "qualities", "share"),
|
||||
[("raw", [4.0, 3.0, 0.0, -1.0], 0.66666667),
|
||||
("percent", [3.80952381, 2.85714286, 0.0, 0.0], 0.57142857),
|
||||
("uniform", [2.0, 1.0, 0.0, 1.0], 0.5)],
|
||||
)
|
||||
def test_segment_metrics_cluster_projection_and_valid_filtering(
|
||||
synthetic_cluster_state, mode, qualities, share
|
||||
) -> None:
|
||||
state, rows = synthetic_cluster_state
|
||||
# Percent normalizes nonnegative representative scores (total=105), then
|
||||
# projects clusters; only segment_metrics filters the invalid 100-point row.
|
||||
assert segment_metrics(state, rows, "D1", "12:01", mode) == {
|
||||
"prefix": "D1", "mode": mode, "segment_count_window": 4,
|
||||
"valid_segment_count": 3, "truth_segment_id": 0, "truth_retained": True,
|
||||
"top_segment_correct": True, "top_segment_tie": False,
|
||||
"top_segment_ids": [0], "top_share": share, "segment_qualities": qualities,
|
||||
}
|
||||
excluded = segment_metrics(state, rows, "D1", "12:04", mode)
|
||||
assert excluded["truth_segment_id"] == 2
|
||||
assert excluded["truth_retained"] is False
|
||||
assert excluded["top_segment_correct"] is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", m1.MODES)
|
||||
def test_segment_metrics_counts_truth_in_any_tied_top_segment(mode) -> None:
|
||||
rows = [{"time": "12:00", "D1": 1}, {"time": "12:01", "D1": 2}]
|
||||
state = {"posterior": rows, "valid": rows, "scores": {"12:00": 2.0, "12:01": 2.0}}
|
||||
result = segment_metrics(state, rows, "D1", "12:01", mode)
|
||||
assert result["top_segment_ids"] == [0, 1]
|
||||
assert result["top_segment_tie"] is True
|
||||
assert result["top_segment_correct"] is True
|
||||
assert result["truth_retained"] is True
|
||||
assert result["top_share"] == 0.5
|
||||
absent = segment_metrics(state, rows, "D1", "12:02", mode)
|
||||
assert absent["truth_segment_id"] is None
|
||||
assert absent["truth_retained"] is False
|
||||
assert absent["top_segment_correct"] is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", m1.MODES)
|
||||
def test_segment_metrics_empty_window_and_empty_valid_set(mode) -> None:
|
||||
empty = {"posterior": [], "valid": [], "scores": {}}
|
||||
result = segment_metrics(empty, [], "D1", "12:00", mode)
|
||||
assert result == {
|
||||
"prefix": "D1", "mode": mode, "segment_count_window": 0,
|
||||
"valid_segment_count": 0, "truth_segment_id": None, "truth_retained": False,
|
||||
"top_segment_correct": False, "top_segment_tie": False,
|
||||
"top_segment_ids": [], "top_share": None, "segment_qualities": [],
|
||||
}
|
||||
rows = [{"time": "12:00", "D1": 1}]
|
||||
result = segment_metrics({**empty, "posterior": rows, "scores": {"12:00": 2}}, rows, "D1", "12:00", mode)
|
||||
assert result == {
|
||||
"prefix": "D1", "mode": mode, "segment_count_window": 1,
|
||||
"valid_segment_count": 0, "truth_segment_id": 0, "truth_retained": False,
|
||||
"top_segment_correct": False, "top_segment_tie": False,
|
||||
"top_segment_ids": [], "top_share": None, "segment_qualities": [0.0],
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("mode", m1.MODES)
|
||||
@pytest.mark.parametrize("score", [0.0, -2.0])
|
||||
def test_segment_metrics_nonpositive_scores_do_not_create_a_scored_leader(mode, score) -> None:
|
||||
rows = [{"time": "12:00", "D1": 1}]
|
||||
state = {"posterior": rows, "valid": rows, "scores": {"12:00": score}}
|
||||
result = segment_metrics(state, rows, "D1", "12:00", mode)
|
||||
assert result["valid_segment_count"] == 1
|
||||
assert result["truth_retained"] is True
|
||||
assert result["top_segment_tie"] is False
|
||||
assert result["segment_qualities"] == ([1.0] if mode == "uniform" else [0.0] if mode == "percent" else [score])
|
||||
assert result["top_share"] == (1.0 if mode == "uniform" else None)
|
||||
assert result["top_segment_ids"] == ([0] if mode == "uniform" else [])
|
||||
assert result["top_segment_correct"] is (mode == "uniform")
|
||||
|
||||
|
||||
def test_threshold_scan_empty_no_coverage_and_inclusive_boundary() -> None:
|
||||
assert threshold_scan([], [0.5]) == {
|
||||
"0.5": {"n": 0, "denominator": 0, "coverage": None, "accuracy": None, "truth_retained": None}
|
||||
}
|
||||
rows = [_metric(0.5), _metric(0.9, False, False), _metric(0.49999999), _metric(None)]
|
||||
assert threshold_scan(rows, [0.5, 0.9, 1.0]) == {
|
||||
"0.5": {"n": 2, "denominator": 4, "coverage": 0.5, "accuracy": 0.5, "truth_retained": 0.5},
|
||||
"0.9": {"n": 1, "denominator": 4, "coverage": 0.25, "accuracy": 0.0, "truth_retained": 0.0},
|
||||
"1.0": {"n": 0, "denominator": 4, "coverage": 0.0, "accuracy": None, "truth_retained": None},
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("case", "expected"),
|
||||
[({"birth": {"date": "1899-12-31"}}, True),
|
||||
({"birth": {"date": "1900-01-01"}}, False),
|
||||
({"birth": {"date": "2000-01-01"}}, False),
|
||||
({"birth": {}}, False), ({}, False)],
|
||||
)
|
||||
def test_is_lmt_synthetic_year_boundary(case, expected) -> None:
|
||||
assert m1.is_lmt(case) is expected
|
||||
|
||||
|
||||
def test_stratified_rows_preserves_denominators_and_empty_strata() -> None:
|
||||
items = [
|
||||
_item("synthetic-old", _metric(0.9, True, True, 1, True), True),
|
||||
_item("synthetic-new", _metric(0.5, False, True, 3)),
|
||||
_item("synthetic-empty", _metric(None, False, False, 0)),
|
||||
]
|
||||
result = m1.stratified_rows(items, "D1", "raw")
|
||||
for label, expected in {
|
||||
"all": (3, 0.46666667, 1, 2, 1, 2),
|
||||
"lmt_before_1900": (1, 0.9, 1, 1, 1, 1),
|
||||
"post_1900": (2, 0.25, 0, 1, 0, 1),
|
||||
}.items():
|
||||
row = result[label]
|
||||
assert tuple(row[key] for key in (
|
||||
"denominator", "top_share_mean", "top_segment_correct", "truth_retained",
|
||||
"top_segment_ties", "valid_segment_count_le_2",
|
||||
)) == expected
|
||||
denominator = expected[0]
|
||||
for count in ("top_segment_correct", "truth_retained", "valid_segment_count_le_2"):
|
||||
assert row[f"{count}_rate"] == round(row[count] / denominator, 8)
|
||||
assert row["top_segment_tie_rate"] == round(row["top_segment_ties"] / denominator, 8)
|
||||
assert row["truth_excluded"] == denominator - row["truth_retained"]
|
||||
empty = m1.stratified_rows(items[1:], "D1", "raw")["lmt_before_1900"]
|
||||
assert empty["denominator"] == 0
|
||||
assert empty["top_share_mean"] is None
|
||||
assert empty["truth_retained_rate"] is None
|
||||
assert empty["top_segment_correct_rate"] is None
|
||||
assert empty["top_segment_tie_rate"] is None
|
||||
assert empty["valid_segment_count_le_2_rate"] is None
|
||||
assert all(row["coverage"] is None for row in empty["thresholds_full_fit"].values())
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def m1_artifact():
|
||||
# Read the completed engine artifact; never regenerate it in this suite.
|
||||
return json.loads(
|
||||
(ROOT / "artifacts" / "varga-resolution" / "varga_resolution_m1.json").read_text(encoding="utf-8")
|
||||
)
|
||||
|
||||
|
||||
def test_m1_artifact_has_all_unique_case_radius_mode_combinations(m1_artifact) -> None:
|
||||
payload = m1_artifact
|
||||
assert payload["schema"] == m1.SCHEMA
|
||||
assert payload["errors"] == []
|
||||
metadata = payload["metadata"]
|
||||
assert metadata["case_count_requested"] == metadata["case_count_completed"] == 77
|
||||
assert metadata["radii"] == list(RADII) == [10, 15, 30, 60]
|
||||
assert metadata["vargas"] == list(VARGA_PREFIXES) == ["D1", "D9", "D10"]
|
||||
assert metadata["modes"] == list(m1.MODES) == ["raw", "percent", "uniform"]
|
||||
assert metadata["thresholds"] == list(THRESHOLDS)
|
||||
cases = m1.load_cases()
|
||||
case_ids = {str(case["case_id"]) for case in cases}
|
||||
assert len(cases) == len(case_ids) == 77
|
||||
items = payload["items"]
|
||||
pairs = [(item["case_id"], item["radius"]) for item in items]
|
||||
assert len(pairs) == len(set(pairs)) == 308
|
||||
assert set(pairs) == {(case_id, radius) for case_id in case_ids for radius in RADII}
|
||||
lmt_by_case = {str(case["case_id"]): m1.is_lmt(case) for case in cases}
|
||||
for item in items:
|
||||
assert item["lmt_before_1900"] is lmt_by_case[item["case_id"]]
|
||||
assert set(item["by_varga"]) == set(VARGA_PREFIXES)
|
||||
for prefix, modes in item["by_varga"].items():
|
||||
assert set(modes) == set(m1.MODES)
|
||||
for mode, row in modes.items():
|
||||
assert (row["prefix"], row["mode"]) == (prefix, mode)
|
||||
aggregates = payload["aggregates"]
|
||||
assert len(aggregates) == len(RADII)
|
||||
assert {row["radius"] for row in aggregates} == set(RADII)
|
||||
for aggregate in aggregates:
|
||||
assert aggregate["case_count"] == 77
|
||||
assert aggregate["lmt_case_count"] == sum(lmt_by_case.values())
|
||||
assert set(aggregate["by_varga"]) == set(VARGA_PREFIXES)
|
||||
assert all(set(modes) == set(m1.MODES) for modes in aggregate["by_varga"].values())
|
||||
|
||||
|
||||
@pytest.mark.parametrize("radius", RADII)
|
||||
@pytest.mark.parametrize("prefix", VARGA_PREFIXES)
|
||||
@pytest.mark.parametrize("mode", m1.MODES)
|
||||
def test_m1_stored_aggregates_and_loo_rows_match_offline_recomputation(m1_artifact, radius, prefix, mode) -> None:
|
||||
# Recompute only small summary arithmetic, not native_case/engine scoring.
|
||||
items = [item for item in m1_artifact["items"] if item["radius"] == radius]
|
||||
aggregate = next(row for row in m1_artifact["aggregates"] if row["radius"] == radius)
|
||||
stored = aggregate["by_varga"][prefix][mode]
|
||||
assert stored["full_fit"] == m1.stratified_rows(items, prefix, mode)
|
||||
assert stored["loo_fixed_thresholds"] == m1.loo_fixed(items, prefix, mode)
|
||||
loo = m1.loo_selected(items, prefix, mode)
|
||||
validation = loo["validation_rows"]
|
||||
assert stored["loo_selected_threshold"] == {
|
||||
**m1.summarize_loo_rows(validation),
|
||||
"selected_threshold_counts": loo["selected_threshold_counts"],
|
||||
"validation_rows": validation,
|
||||
"lmt_before_1900": m1.summarize_loo_rows([row for row in validation if row["lmt_before_1900"]]),
|
||||
"post_1900": m1.summarize_loo_rows([row for row in validation if not row["lmt_before_1900"]]),
|
||||
}
|
||||
|
||||
|
||||
def test_m1_report_raw_table_matches_stored_artifact(m1_artifact) -> None:
|
||||
report = (ROOT / "docs" / "research" / "rectification_varga_resolution_2026_09_30.md").read_text(encoding="utf-8")
|
||||
section = report.split("## 4. M1", 1)[1].split("## 5.", 1)[0]
|
||||
actual = [
|
||||
[cell.strip() for cell in line.strip().strip("|").split("|")]
|
||||
for line in section.splitlines() if line.startswith("| ±")
|
||||
]
|
||||
expected = []
|
||||
for aggregate in m1_artifact["aggregates"]:
|
||||
for prefix in VARGA_PREFIXES:
|
||||
raw = aggregate["by_varga"][prefix]["raw"]
|
||||
full = raw["full_fit"]["all"]
|
||||
loo = raw["loo_selected_threshold"]
|
||||
denominator = full["denominator"]
|
||||
expected.append([
|
||||
f"±{aggregate['radius']}", prefix,
|
||||
f"{full['top_segment_correct']}/{denominator}",
|
||||
f"{full['truth_retained']}/{denominator}",
|
||||
f"{full['valid_segment_count_le_2']}/{denominator}",
|
||||
f"{loo['eligible']} / {loo['accuracy']:.3f} / {loo['truth_retained']:.3f}",
|
||||
])
|
||||
assert actual == expected
|
||||
Reference in New Issue
Block a user