research(rectification): typed events scored as answered probes — no benefit (BUG-1089)
v4 open holdout, ±10/±30/±60, raw and percent priors: only 4/9/14 of 57 day-precision training events split candidates; widths unchanged, top-1 drops. Narrowed year blocking (R3) mixed. All arms no_benefit; no implementation brief recommended. Two runs byte-identical. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
56b51e2170
commit
bcf0c86feb
+11
-7
@@ -14647,25 +14647,29 @@
|
||||
|
||||
## BUG-1088 | 学业质量题一律问「那次上大学」,只记得年份的经历显示成「YYYY 年 1 月」
|
||||
|
||||
- 状态:investigating
|
||||
- 状态:resolved(代码 + 回归测试 + 重新冻结;未部署,真机待合入 staging 后复核)
|
||||
- 首次发现 / 最近更新:2026-09-29 / 2026-09-29
|
||||
- 影响面:`scripts/rectification/event_probes.py::_quality_user_meaning`、`_display_date_label` / `_event_month`(冻结计分文件,ERR-110)。
|
||||
- 影响面:`scripts/rectification/event_probes.py::_quality_user_meaning`、`_display_date_label`(冻结计分文件,ERR-110)。
|
||||
- 用户现象:用户说的是某年一次艺考,卡片问「YYYY 年 1 月那次上大学,更接近如愿、将就调剂…」。
|
||||
- 根因:学业模板写死「上大学」,不看 `event_kind`;年精度经历 `date_start = YYYY-01-01`,取月不看 `precision`。
|
||||
- 修复:待执行(研究单 R0,改文字后按 ERR-110 用新 freeze 路径重新冻结)。
|
||||
- 相关记录:ERR-110
|
||||
- 触发条件:学业经历类型为 education_start / change / interruption,且出现质量题;年精度经历存为 `YYYY-01-01`。
|
||||
- 根因:学业模板写死「上大学」,不看 `event_kind`;`_display_date_label` 用 `_event_month` 取月,不看 `precision`。
|
||||
- 修复:学业分支按类型称呼(升学 / 学业变动 / 学业中断,兜底「学业经历」);`precision == "year"` 的标签只写「YYYY 年」。`_event_month` 与 split hash 不变,计分身份不变。
|
||||
- 验证:`tests/test_rectification_event_probes.py::QualityWordingTests`(4 条);同机 A/B(`scripts/research/quality_wording_ab.py`,`PYTHONHASHSEED=0`)只有 `user_meaning` / `display_date_label` 文本变化,分数、候选决策、题目 id / split / outcomes 逐字节相同;按 ERR-110 新建冻结 `docs/research/sealed_holdout_rerun_quality_wording_2026_09_29.freeze.json`、`docs/research/reported_offset_quality_wording_2026_09_29.freeze.json`,重放 20 / 900 次结果与 09-21 完全一致;验证完整性门禁 83 passed。
|
||||
- 防复发:年精度经历不得在任何用户文案里补月份;新增学业经历类型时同步 `EDUCATION_QUALITY_EVENT_NAME`。质量题四个选项仍含「录取」字样,未改。
|
||||
- 相关记录:ERR-110、BUG-1089
|
||||
- 复发自:无
|
||||
- 修复版本:未修
|
||||
- 修复版本:分支 `codex/rectification-typed-event-research-20260929`,未合入
|
||||
|
||||
## BUG-1089 | 带年月的打字经历不按选择题规则计分,还挡掉同领域相邻年份的选择题
|
||||
|
||||
- 状态:investigating(离线研究单待执行;不上线)
|
||||
- 状态:investigating(离线研究已完成,结论不改代码;等止血单 BUG-1084 落地后由 Claude 决定是否关闭)
|
||||
- 首次发现 / 最近更新:2026-09-29 / 2026-09-29
|
||||
- 影响面:`core/build-state.ts::buildInferenceState`(evidence「是」跳过)、`event_probes.py::_existence_blocked_years`。
|
||||
- 用户现象:同一件「某年某月工作变动」,点选卡答「是」动 ±2 分,自己打字说约等于 0;说了之后那一年附近的选择题也不再出。
|
||||
- 已确认事实(虚构 12 例):12 件经历里去掉 4 件记得到月的,选择题 5/12 例变多,最终宽度中位 17 → 13 分钟。
|
||||
- 根因:计分通道不对称,见 `docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`。
|
||||
- 修复:未修。研究单量「打字经历按选择题边界规则计分」在 v4 开放集三档是否过门;过门才立实现单,合入前产品确认。
|
||||
- 研究结果(2026-09-29,`docs/research/rectification_typed_event_scoring_2026_09_29.md`,v4 开放集 20 例,开放集回放不是准确率):按选择题同一判定函数给打字经历计分(R1),57 件日精度训练经历在 ±10 / ±30 / ±60 下只有 4 / 9 / 14 件能把候选分到两边,宽度中位三档不变,头名在 raw 口径下三档都降,判 `no_benefit`;挡题缩到同领域同年(R3)宽窗头名 +0.10~0.15 但 ±10 头名 −0.15,判 `no_benefit`;年月记错 ±1–3 月挤出真值 0–2.9%。结论:计分通道不对称在代码上属实,但打字经历本身多半分不开候选,改计分不划算;不立实现单。
|
||||
- 防复发:不得在研究未过门时改先验尺度或淘汰线(BUG-560)。
|
||||
- 相关记录:BUG-560、BUG-1048、BUG-1084
|
||||
- 复发自:无
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,104 @@
|
||||
# 打字经历按选择题规则计分:离线研究(2026-09-29)
|
||||
|
||||
- 任务书:`docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`;进度:`docs/tasks/PROGRESS-rectification-typed-event-research-20260929.md`
|
||||
- 分支:`codex/rectification-typed-event-research-20260929`,基线 `origin/staging` @ `56722063`。
|
||||
- 性质:**离线测量,不改线上默认值。** 研究补丁(R3 的 `_existence_blocked_years`)只在一次调用内替换,调用后与脚本末尾都断言已恢复原函数。
|
||||
- 数据:`references/real_case_calibration/minute_rectification_holdout_v4.json`,20 例公开 Rodden-AA。v4 的事件是 59 件日精度、84 件年精度,**没有月精度**;每例留一件作 holdout,训练集为日精度 57 件、年精度 67 件。**这是开放集回放,不是盲测,下面的数字不是准确率。**
|
||||
- 口径:ayanamsa `raman`,node mode `mean`,步长 2 分钟,半径 ±10 / ±30 / ±60;六题回放(`ASK_COUNT = 6`,按真值作答);交付区间 = 未淘汰且落后头名不足 8 分的簇的并集。两种先验都报:`raw`(引擎行原始分,09-14 / 09-26 研究口径)与 `percent`(公开代表按正比例换成整数百分比,线上 `relative_support` 口径)。
|
||||
- 复跑:`PYTHONHASHSEED=0 python3 scripts/research/typed_event_as_probe.py`(约 1.5 分钟)→ `docs/research/rectification_typed_event_scoring_2026_09_29.json`。同机两次运行 JSON 逐字节一致。
|
||||
- 判门(任务书):三档**同时**满足「头名命中不降、真值在区间不降、宽度中位下降」才算有收益。
|
||||
|
||||
## 结论
|
||||
|
||||
| 项 | 一句话结论 | 是否建议立实现单 |
|
||||
| --- | --- | --- |
|
||||
| R0 学业质量题措辞(BUG-1088) | 已修:按经历类型称呼(升学 / 学业变动 / 学业中断),年精度经历不再显示「1 月」;计分输出逐字节不变,已按 ERR-110 用新路径重新冻结 | 已在本分支实现 |
|
||||
| R1 打字经历当作已答「是」的选择题 | **不过门。** 打字经历绝大多数分不开候选:±10 / ±30 / ±60 下,57 件日精度训练经历只有 4 / 9 / 14 件(7% / 16% / 25%)能把代表候选分到两边;宽度中位三档都不变,头名在 raw 口径下三档都降 | **不建议** |
|
||||
| R2 经历年月记错 ±1–3 个月 | 挤出真值的比例 0–2.9%,与 09-26 答错两题(±30 约 7%)相比不高;但 R1 本身没有收益,R2 不改变结论 | 不适用 |
|
||||
| R3 挡题范围缩到「同领域同年」 | **不过门,但宽窗有信号。** raw 口径下 ±30 头名 0.50 → 0.65、±60 0.35 → 0.45;±10 头名 0.75 → 0.60(宽度 15 → 14)。宽度三档基本不变。60 格中有 45 格实际问到的六题因此换了 | **不建议直接实现**;若要继续,需另立研究,先解释 ±10 的头名下降 |
|
||||
|
||||
**给产品的话:** 「打字说的经历不参与收窄」在代码里确实是两条不同的通道,但把打字经历改成按选择题规则计分**几乎没有用处**:一件你记得清清楚楚的事,多半恰好落在所有候选时刻都一样的那段大运里,本来就分不开候选。能分开候选的,是系统专门挑出来的「边界月份」选择题。这也解释了 09-26 研究 R3c 的发现:边界月份里只有 3–4.5% 能把代表候选分开。所以止血单「门开后不再索要经历、选择题问完就出卡」的方向是对的,本研究不支持再往打字经历上加计分。
|
||||
|
||||
## R1 · 打字经历当作已答选择题
|
||||
|
||||
**方法。** 对每件训练经历,用出题器的同一个判定函数 `event_probes._evaluate_contexts`(代表候选、签名聚类、同一套集合版本),以经历的领域和年月(日精度取所在月;R1y 的年精度取 `month=None`,即现有「整年」判定)生成一道「题」,按「是」作答(±2,强冲突计数同线上),然后照常回放六道选择题。函数返回 `None` 的经历(候选在那个月没有分边)不计分。
|
||||
|
||||
**能分开候选的经历(每档 20 例合计):**
|
||||
|
||||
| 半径 | 日精度训练经历 | 能分边的 | 年精度训练经历 | 能分边的 |
|
||||
| --- | ---: | ---: | ---: | ---: |
|
||||
| ±10 | 57 | 4 | 67 | 5 |
|
||||
| ±30 | 57 | 9 | 67 | 10 |
|
||||
| ±60 | 57 | 14 | 67 | 14 |
|
||||
|
||||
**raw 口径:**
|
||||
|
||||
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 并列率 | 每例提问 | 每例计分经历 |
|
||||
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||
| ±10 | B0 | 0.75 | 1.00 | 15 | 0.05 | 4.25 | 0 |
|
||||
| ±10 | R1d | 0.70 | 1.00 | 15 | 0.05 | 4.25 | 0.20 |
|
||||
| ±10 | R1y | 0.60 | 1.00 | 14 | 0.05 | 4.25 | 0.45 |
|
||||
| ±30 | B0 | 0.50 | 1.00 | 33 | 0.05 | 5.40 | 0 |
|
||||
| ±30 | R1d | 0.45 | 1.00 | 33 | 0.05 | 5.40 | 0.45 |
|
||||
| ±30 | R1y | 0.45 | 1.00 | 33 | 0.05 | 5.40 | 0.95 |
|
||||
| ±60 | B0 | 0.35 | 1.00 | 56 | 0.05 | 5.40 | 0 |
|
||||
| ±60 | R1d | 0.30 | 1.00 | 56 | 0.05 | 5.40 | 0.70 |
|
||||
| ±60 | R1y | 0.30 | 1.00 | 56 | 0.05 | 5.40 | 1.40 |
|
||||
|
||||
**percent 口径(线上先验尺度):**
|
||||
|
||||
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 并列率 |
|
||||
| --- | --- | ---: | ---: | ---: | ---: |
|
||||
| ±10 | B0 | 0.50 | 1.00 | 11 | 0.40 |
|
||||
| ±10 | R1d | 0.45 | 1.00 | 11 | 0.40 |
|
||||
| ±10 | R1y | 0.35 | 1.00 | 11 | 0.45 |
|
||||
| ±30 | B0 | 0.40 | 1.00 | 33 | 0.55 |
|
||||
| ±30 | R1d | 0.40 | 1.00 | 33 | 0.55 |
|
||||
| ±30 | R1y | 0.40 | 1.00 | 33 | 0.45 |
|
||||
| ±60 | B0 | 0.15 | 1.00 | 53 | 0.75 |
|
||||
| ±60 | R1d | 0.20 | 1.00 | 53 | 0.75 |
|
||||
| ±60 | R1y | 0.15 | 1.00 | 53 | 0.75 |
|
||||
|
||||
**读法。**
|
||||
|
||||
1. 每例平均只有 0.2–0.7 件日精度经历能计分,宽度中位三档都不动。多出来的 ±2 只改排序,而且往往改错:raw 口径下头名三档都下降。
|
||||
2. 真值从未被挤出区间(真值在区间保持 1.00):单件经历的 ±2 不够 8 分淘汰线。
|
||||
3. percent 口径的并列率高(0.40–0.75),是整数百分比把相近原始分压成同一个数所致,与本研究的方案无关;这是线上先验尺度(BUG-560)的另一面,只记录,不处理。
|
||||
|
||||
## R2 · 经历年月记错 ±1–3 个月
|
||||
|
||||
每例每档随机挑 1 件或 2 件日精度训练经历,把月份挪 ±1–3 个月(种子 `20260929`,每格重复 10 次),按 R1d 计分后回放。k=2 只在日精度训练经历 ≥2 的 14 例上跑。
|
||||
|
||||
| 口径 | 半径 | k | 回放次数 | 挤出真值比例 | 至少一次挤出的例子 | 宽度中位 |
|
||||
| --- | --- | ---: | ---: | ---: | ---: | ---: |
|
||||
| raw | ±10 | 1 / 2 | 200 / 140 | 0.0% / 2.9% | 0 / 1 | 15 / 15 |
|
||||
| raw | ±30 | 1 / 2 | 200 / 140 | 1.5% / 2.1% | 1 / 1 | 33 / 36 |
|
||||
| raw | ±60 | 1 / 2 | 200 / 140 | 0.0% / 0.0% | 0 / 0 | 56 / 54 |
|
||||
| percent | ±10 | 1 / 2 | 200 / 140 | 0.0% / 2.9% | 0 / 1 | 11 / 12 |
|
||||
| percent | ±30 | 1 / 2 | 200 / 140 | 1.5% / 2.9% | 1 / 1 | 33 / 34 |
|
||||
| percent | ±60 | 1 / 2 | 200 / 140 | 0.0% / 0.7% | 0 / 1 | 53 / 53 |
|
||||
|
||||
挤出率低,原因同 R1:挪错后的月份多半也分不开候选,等于没计分。R2 不改变 R1 的结论。
|
||||
|
||||
## R3 · 挡题范围缩到「同领域同年」
|
||||
|
||||
线上 `_existence_blocked_years` 会把已入账年份及同领域相邻年份(`EXISTENCE_NEARBY_YEARS`)的存在题挡掉。研究补丁改为只挡「同领域同一年」,其余不变。
|
||||
|
||||
| 口径 | 半径 | B0 头名 / 宽度 | R3 头名 / 宽度 | R3+R1d 头名 / 宽度 | R3 判定 |
|
||||
| --- | --- | --- | --- | --- | --- |
|
||||
| raw | ±10 | 0.75 / 15 | 0.60 / 14 | 0.55 / 14 | 头名降 |
|
||||
| raw | ±30 | 0.50 / 33 | 0.65 / 33 | 0.60 / 33 | 宽度未降 |
|
||||
| raw | ±60 | 0.35 / 56 | 0.45 / 55 | 0.40 / 54 | 过 |
|
||||
| percent | ±10 | 0.50 / 11 | 0.45 / 11 | 0.40 / 11 | 头名降 |
|
||||
| percent | ±30 | 0.40 / 33 | 0.50 / 31 | 0.50 / 30 | 过 |
|
||||
| percent | ±60 | 0.15 / 53 | 0.15 / 54 | 0.20 / 53 | 宽度升 |
|
||||
|
||||
真值在区间:除 percent ±60 R3+R1d 为 0.95 外全部 1.00。出题供给(截断到 `MAX_PROBES = 8` 之后)几乎不变(每例 5.1 / 7.1 / 7.2 → 5.1 / 7.1 / 7.15),但 60 格中有 45 格实际问到的前六题换了:放开相邻年份后,出题器会挑到别的年份。
|
||||
|
||||
**读法。** 宽窗(±30 / ±60)头名有 +0.10 到 +0.15 的改善,但样本只有 20 例(每例 0.05),±10 头名下降同样幅度,三档不同时过门。按任务书的门槛判 `no_benefit`。若产品想追,下一步应是另立研究:只在窗口 ≥30 分钟时放开,并先解释 ±10 下降是换题带来的随机波动还是系统性问题。
|
||||
|
||||
## 偏离与说明
|
||||
|
||||
- 基线头名与 09-26 研究略有差异(raw B0:本文 0.75 / 0.50 / 0.35,09-26 为 0.80 / 0.55 / 0.40;宽度 15 / 33 / 56 完全一致)。本脚本自己实现了回放循环(为了在六题之前插入经历计分),`top1_hit` 用的是相邻合并后的簇;各方案之间用的是同一条回放链路,比较有效,但数字不要与 09-26 逐格对照。
|
||||
- 任务书以「记得到月」为主设想;v4 没有月精度事件,日精度事件按所在月判定,等价于月精度判定。
|
||||
- 任务书写的 R1 年精度口径「只在跨年边界时计分」,实现为出题器现有的整年判定(`_evaluate_contexts(month=None)`,事件日取 7 月 1 日),它已经只在候选整年激活状态不同时才分边。
|
||||
@@ -0,0 +1,51 @@
|
||||
# PROGRESS · 打字经历按选择题规则计分(离线研究)+ 学业质量题措辞(2026-09-29)
|
||||
|
||||
- 任务书:`docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`
|
||||
- 执行:Claude(直接执行模式,fork 子代理),分支 `codex/rectification-typed-event-research-20260929`,基线 `origin/staging` @ `56722063`。未推送。
|
||||
- 研究报告:`docs/research/rectification_typed_event_scoring_2026_09_29.md` + `.json`
|
||||
|
||||
## 开工预检
|
||||
|
||||
- `python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45`:exit 0;python_runtime / fragment_scan / external_engine_adapters / focused_tests 全部 ok,remote_visibility `verified`。
|
||||
- 读 `docs/research/pre_work_error_ledger.md` ERR-110(冻结文件改动须重新冻结)、ERR-111(固定 `PYTHONHASHSEED`)。本机无 `.venv`,全部用系统 `python3`。
|
||||
|
||||
## R0 · 学业质量题措辞(BUG-1088)
|
||||
|
||||
- `scripts/rectification/event_probes.py`:`_display_date_label` 对 `precision == "year"` 不取月;`_quality_user_meaning` 学业分支按 `event_kind` 称呼(`EDUCATION_QUALITY_EVENT_NAME`:升学 / 学业变动 / 学业中断,兜底「学业经历」),句式「…那次X,结果更接近如愿、将就调剂、发挥失常还是说不清」。四个选项标签未改(仍含「录取」字样,属选项文案,留待产品决定)。
|
||||
- **计分身份不变**:`_event_month` 未改,`candidate_split_hash`、题目 `month` 字段仍带存储的月份;新增测试锁住这一点。
|
||||
- 同机 A/B(`scripts/research/quality_wording_ab.py`,`PYTHONHASHSEED=0`,v4 前 4 例 ±10 + 2 个虚构出生带学业质量事件):基线连跑两次逐字节一致;改后对比,**只有 2 个叶子字段类型变化**:`user_meaning`(2 处)与 `display_date_label`(2 处),均在虚构例的 `known_event_quality` 题上;引擎行分数、公开候选决策、全部题目的 id / split hash / outcomes 逐字节相同。
|
||||
- `2010 年 1 月那次上大学…` → `2010 年那次学业中断…`(年精度 + interruption)
|
||||
- `2011 年 1 月那次上大学…` → `2011 年那次升学…`(年精度 + start)
|
||||
- 重新冻结(ERR-110,不覆盖旧记录):
|
||||
- `scripts/research/sealed_holdout_rerun.py` `FREEZE/REPORT` → `docs/research/sealed_holdout_rerun_quality_wording_2026_09_29.{freeze.json,json}`;重放 20 例,metrics 与 20 条 trials 与 09-21 记录**完全相同**(top1 0.45 / top3 0.50 / MAE 6.45)。
|
||||
- `scripts/research/reported_offset_sweep.py` `FREEZE/REPORT` → `docs/research/reported_offset_quality_wording_2026_09_29.{freeze.json,json}`;900 trials 与 summary 与 09-21 **完全相同**。
|
||||
- `references/rectification_sealed_holdout.v1.json`:`current_tree_scorer`、`current_tree_fixed_protocol_rerun`、`current_tree_reported_offset_replay` 指向新记录;旧的 09-21 记录移入 `previous_fixed_protocol_rerun` / `previous_reported_offset_replay`;`evaluated_on` 2026-09-29。遗留 12 文件哈希 `c913c6f0…` 不变,只有扩展身份(event_probes.py 与两份研究脚本)变化。
|
||||
- `tests/test_sealed_holdout_contract_freshness.py` 的 `IDENTITY_FORBIDDEN_PREFIXES` 补两条新前缀(同 09-21 先例)。
|
||||
- 新测试:`tests/test_rectification_event_probes.py::QualityWordingTests`(4 条:三种类型不叫上大学、年精度无月、月精度保留月、split 月份不变)。
|
||||
|
||||
## R1–R3 · 离线研究(BUG-1089)
|
||||
|
||||
脚本 `scripts/research/typed_event_as_probe.py`,`PYTHONHASHSEED=0` 连跑两次 JSON 逐字节一致(各 1 分 25 秒)。结论全部 `no_benefit`,**不建议立实现单**。数字见研究报告;要点:
|
||||
|
||||
| 项 | raw 口径(±10 / ±30 / ±60) | 判定 |
|
||||
| --- | --- | --- |
|
||||
| B0 头名 / 宽度 | 0.75 / 0.50 / 0.35;15 / 33 / 56 | 基线 |
|
||||
| R1d 头名 / 宽度 | 0.70 / 0.45 / 0.30;15 / 33 / 56 | no_benefit |
|
||||
| R1y 头名 / 宽度 | 0.60 / 0.45 / 0.30;14 / 33 / 56 | no_benefit |
|
||||
| R3 头名 / 宽度 | 0.60 / 0.65 / 0.45;14 / 33 / 55 | no_benefit(±10 头名降) |
|
||||
| 能分边的日精度训练经历 | 4 / 9 / 14(共 57 件) | — |
|
||||
| R2 挤出真值(k=1 / k=2) | 0–1.5% / 2.1–2.9% | — |
|
||||
|
||||
percent 口径结论相同(判定全部 `no_benefit`)。
|
||||
|
||||
## 测试
|
||||
|
||||
- 定向:`tests/test_rectification_validation_integrity_gate.py`、`tests/test_sealed_holdout_contract_freshness.py`、`tests/test_reported_offset_research.py`、`tests/test_rectification_event_probes.py`:83 passed(改前重新冻结前为 7 失败,全部是冻结身份不匹配)。
|
||||
- 快速门 `python3 scripts/run_quality_gate.py --profile quick`(系统 python3,本机无 `.venv`):JSON / 审计 / 碎片 / BPHS 各步通过;pytest **1001 passed, 1 skipped**(含上面四个文件)。随后 `npm test` 步骤 exit 127:本工作树没有 `frontend/node_modules`,npm 找不到测试依赖。**前端三步(npm test / lint / build)未跑,记为环境缺口,不算通过。** 本轮没有改任何前端文件。
|
||||
- 隐私:`tests/test_repo_privacy_markers.py` 通过。
|
||||
|
||||
## 偏离
|
||||
|
||||
- 基线回放头名与 09-26 研究差 0.05(宽度一致),原因见研究报告「偏离与说明」;方案之间可比,不与 09-26 逐格对照。
|
||||
- v4 无月精度事件,日精度按所在月判定。
|
||||
- 未改质量题四个选项标签;前端 `collection-question-pool.ts` 也有「上大学」措辞(采集题示例),不在本单范围,未改。
|
||||
@@ -254,7 +254,7 @@
|
||||
| `TASK-self-edit-avatar-menu-20260928.md` | `PROGRESS-self-edit-avatar-menu-20260928.md` | **本人可编辑 + 各页头像菜单 + 报告页去说明**:`/people` 本人可编辑出生资料(BUG-1081,`ab6c55f8` 漏掉);次级页头像弹同一账户菜单、弹窗项跳 `/?account=`;我的报告删顶部说明与统计 | 已实现待验收(分支 `codex/self-edit-avatar-menu-20260928`,未推送) | BUG-1081 |
|
||||
| `TASK-mobile-chart-and-confirmed-edit-20260929.md` | `PROGRESS-mobile-chart-confirmed-edit-20260929.md` | **手机星盘 + 确认时间可改 + 天空定格**:iPhone 星盘被宽表撑出屏幕(BUG-1083);confirmed 状态改声明字段也重置排盘时间(推翻 BUG-264 约定);「那一刻的天空」改为重放汇聚→定格→右上角分享;去掉报告星图封面 | 已验收(Claude 2026-09-29),已推 staging | BUG-1083 |
|
||||
| `TASK-rectification-futile-collect-stop-20260929.md` | — | **生时校正停掉无效补经历循环(止血)**:打字经历不收窄(BUG-560 后果),流程却一路索要,真机 22 件整窗不动;采集只为开闸、门开后只问点选卡问完即出卡;交付正文去吻合率;多段时旁白误报「范围没变」;交付轮带采集题 | 待领取 | BUG-1084~1087 |
|
||||
| `TASK-rectification-typed-event-scoring-research-20260929.md` | — | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 待领取 | BUG-1088、1089 |
|
||||
| `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 待验收(Claude 子代理直接执行,分支 `codex/rectification-typed-event-research-20260929`,未推送;R1–R3 均 no_benefit) | BUG-1088、1089 |
|
||||
| `TASK-serif-headings-20260928.md` | `PROGRESS-serif-headings-20260928.md` | **全站标题改用自托管宋体、正文保持黑体**:产品推翻 BUG-737「CJK 不用衬线」结论(保留「声明的字体必须可加载」「不落系统宋体」两条);Noto Serif SC SemiBold 按通用规范汉字表 6500 字 unicode-range 切片自托管(改名 Jyotisha Serif SC),swap 不 preload;首页宋体流量 ≤300 KB;先于天空封面单 | 已实现待验收(分支 `codex/serif-headings-20260928`,未推送) | 不开新 BUG;BUG-737 追加说明 |
|
||||
| `TASK-cend-ui-claude-alignment-20260916.md` | `PROGRESS-cend-ui-r1/r2/r3-20260916.md` | **C 端界面向 claude.ai 产品界面对齐(三轮串行 R1→R2→R3,都动 `globals.css`,不得并行)**:根因是 `frontend/CLAUDE_DESIGN.md` 扒的是 **claude.com 营销官网**,它自己在 Known Gaps 里写明 claude.ai 产品界面不在范围内,而 `DESIGN.md:3` 把它当成了产品界面的实现契约。**R1**:`--font-display` 里 Tiempos Headline / StyreneB **从未加载**(无 `@font-face`、`public/` 无字体、`layout.tsx` 只 vendor 了 Inter),中文标题全站落到 **宋体 / SimSun**,波及 20 处含助手回答的 h2/h3(BUG-737);亮色强调色 `#85432f` 与暗色 `#d78064` 不同源,产品拍板亮色换 **Claude coral `#cc785c`**,**易漏点**是 `globals.css:16` 的 `--color-ring` 硬编码在 `@theme inline` 里不跟随 `:root`,另有第四个 `:root` 亮色块(`:4358`)必须同步(BUG-738);`.composer-footer` 常驻 44px + 顶栏 68px + `--composer-reserve` 148px,每屏固定吃掉 216px,模型选择器移进输入框内部、删掉底栏、顶栏收到 46px 并删「分析对象」副标题。**R2**:空状态是营销落地页(hero 卡 + 两张 132px 入口大卡 + 3 列 156px 主题卡),输入框被压在 **800px 以上**内容之下,重排成「问候 + 居中输入框 + 两枚入口 pill + 一排 chip」。**R3**:侧栏两个 `<details>` 拍平成一条「最近」、星盘的两个入口(侧栏分组 + 账户菜单)收敛到一处、删掉逐条助手头像。**决策记录 D3 推翻 DESIGN.md「报告强调色与应用同源」一句**(报告刻意保留深棕)。原型图 https://claude.ai/code/artifact/da275da6-2954-4f50-99aa-32bb8694d38b(三套画面 + 明暗,页面标题就是建议字体栈的实际渲染)。环境缺口:无登录态无 Chrome,四项真机观感留 `docs/testing/`。BUG 段 737–738 | 待领取 | — |
|
||||
| `TASK-cend-surfaces-claude-alignment-20260916.md` | `PROGRESS-cend-shell-20260916.md`、`PROGRESS-cend-report-20260916.md`、`PROGRESS-cend-rectification-20260916.md`、`PROGRESS-cend-chart-eph-20260916.md` | **次级页面对齐(上一单的续篇,R4→R5/R6,R7、R8 可并行)**:星盘 `/chart`、星历 `/ephemeris`、报告 `/reports` **各是脱离 app 外壳的独立全屏页**,顶部只有一个「返回对话」链接、侧栏整个消失,且三家各写了一套一模一样的 `*-shell`/`*-topbar`/`*-hero` 骨架——与上一单 E5 同根因(营销站 band 结构被套到产品界面)。**R4** 抽只读导航外壳 `AppNavRail`(只用现成的 `GET /api/sessions` + `GET /api/account`,会话行走 `sessionHref` 跳 `/?c=<uuid>`;**刻意不带**重命名/删除/收藏/归档——那套连着 `Home()` 的乐观更新与回滚,搬过来会撞 useState 增长门禁)。**R5** 星盘五 tab 下划线化 + 参数合表 + 行星表横向滚动;星历日期导航改 `‹ 日期 ›`。**R6** 报告中心卡片网格改行式列表;阅读页加常驻目录。**R7** 生时校正把可信区间从盘面板标题行提成常驻条(窄屏 `.is-compact` 下盘面板是 overlay,现在默认看不到区间),五个 `technique-audit` 折叠块收成两段。**R8** 设置内容区收窄(880px 弹窗里表单铺了 690px)、套餐卡三修饰符收敛成两态。**已解锁**:原挡路的设置单已于 `111b4a84`(BUG-698)合入。**两条不得回退**:BUG-698 的 `@supports (height: 1dvh)` 写法(重复声明回退会被 Lightning CSS 折叠)、BUG-616/617 的报告盘面 grid 实现。默认不占 BUG 号 | **R4–R8 全部已实现并验收合入** | R4:抽出 `AppNavRail`(只读,两个 GET,零写操作)+ `SecondaryShell`,三个次级页并入 app 外壳并删掉各自的 shell/topbar/hero;四个路由渲染标记**完全不变**(`/` `/chart` `/ephemeris` 仍 Static);CSS gzip −0.25%。`/reports/[reportId]` 留给 R6 与目录一起做。两处自身健壮性问题被测试抓到:`usePathname()` 可为 null、`fetch` 可能不存在。差点弄丢 BUG-717 的 eyebrow 文案(已放回)。R8:表单分区收窄到 440px(列表分区不变)、套餐卡三修饰符收敛成互斥的 `is-current` / `is-recommended`,`--highlighted` 删除改为滚动定位;手机端 `order:-1` 改挂 `[data-plan-alias]`(版位不是状态)。测试 3350→3354(净增 4),失败清单与基线逐条一致;`/` 仍 Static;我的干净构建实测 CSS gzip −3 字节。**遗留待产品拍板**:`?plan=` 深链现在完全没有视觉指向,只有滚动位置。R6:报告中心卡片网格改行式列表、阅读页并入外壳并把目录挪到右侧常驻。**任务书 E10 过期**——目录在 `cfcd369d` 就已存在,本轮是挪位置定稿而非从零加。挂外壳带出一个真实打印风险已处理:`.chat-app`/`.chat-panel` 是 `height:100%;overflow:hidden`,裸 `window.print()` 会把九节报告裁成一页,阅读页因此多挂一条只在挂载期生效的 print 样式解锁外壳。「生成中的分节进度」做不了——`REPORT_LIST_COLUMNS` 不返回节数,按 VOICE.md 不许前端编。R7:区间常驻条与盘面折叠收敛。**任务书 E9 也不准确**——对话区顶部早有常驻条 `RectificationTimeline` 且窄屏可见,真正只在盘面标题行的是**代表分钟**;因此没另造第二条,在既有条上补齐代表分钟与已答题数(与盘面同一次 `workingRectificationTime()` 调用)。折叠块实际是 **8 个**不是 5 个。**触发让步顺序第 5 条**:收窄进度未做——服务端无该字段,且 `candidate_range` 会放宽(BUG-572),前端相减会把一次放宽报成收窄,已写进 `BLOCKED.md`。顺带修掉一个**静默失效的旧断言**(`slice(indexOf(A), indexOf(B))` 在 B 改名后变成几乎整份文件,四条 `doesNotMatch` 假通过)|
|
||||
|
||||
@@ -0,0 +1,391 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Offline research for TASK-rectification-typed-event-scoring-research-20260929.
|
||||
|
||||
Question: should a dated typed event ("I changed jobs in 2019-10") move the
|
||||
candidate scores the same way an answered existence probe does?
|
||||
|
||||
Today an answered dated probe moves every candidate by ±2 through the
|
||||
Vimshottari/Narayana boundary-month split (`event_probes._evaluate_contexts`),
|
||||
while a typed event only reaches the engine prior, and the frontend skips
|
||||
`classified_from === "evidence"` yes answers (`core/build-state.ts`).
|
||||
|
||||
This script replays the v4 open holdout (20 public Rodden-AA cases, 7–8 events
|
||||
each; 59 day-precision and 84 year-precision events, no month-precision) with
|
||||
the same six-probe replay used by the 09-14 / 09-16 / 09-26 studies, and adds
|
||||
research arms:
|
||||
|
||||
* ``B0`` baseline: prior, then the first six public probes answered from truth.
|
||||
* ``R1d`` each training typed event with day precision is scored as an answered
|
||||
``yes`` probe (split from `_evaluate_contexts` at that year/month),
|
||||
then the same six probes.
|
||||
* ``R1y`` as R1d, plus year-precision events scored with the year-level split
|
||||
(`_evaluate_contexts(month=None)`, the existing year probe path).
|
||||
* ``R2k`` R1 with k ∈ {1, 2} typed day events moved by ±1–3 months (seeded);
|
||||
reports how often the truth is pushed out of the delivered range.
|
||||
* ``R3`` `_existence_blocked_years` narrowed to "same domain, same year"
|
||||
(probe supply count), alone and combined with R1d.
|
||||
|
||||
Two prior scales are reported: ``raw`` (engine row score, the research-lib
|
||||
convention of the earlier studies) and ``percent`` (proportional percent over
|
||||
public representatives, the production `relative_support` convention).
|
||||
|
||||
Nothing here changes a production default. Module attributes are patched for
|
||||
one call and restored; the end of `main` asserts they are the originals.
|
||||
This is an open-set replay, not a blind test: numbers are not accuracy.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import random
|
||||
import statistics
|
||||
import sys
|
||||
from contextlib import contextmanager
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
from typing import Any, Iterator, Sequence
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
if str(ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
import scripts.rectification.event_probes as event_probes # noqa: E402
|
||||
from scripts.active_rectification_event_engine import compute_candidate_static_contexts # noqa: E402
|
||||
from scripts.rectification.candidate_contrast import cluster_contexts_by_signature # noqa: E402
|
||||
from scripts.rectification.case_holdout import holdout_event_ids # noqa: E402
|
||||
from scripts.rectification.refinement_packet import window_scan # noqa: E402
|
||||
from scripts.rectification.scoring_service import ( # noqa: E402
|
||||
build_event_contribution_matrix,
|
||||
score_from_matrix,
|
||||
)
|
||||
from scripts.research.cluster_width_lib import ( # noqa: E402
|
||||
SEPARATION_LEAD,
|
||||
delivery_from_public,
|
||||
merge_adjacent_traced,
|
||||
public_from_clusters,
|
||||
raw_signature_clusters,
|
||||
still_valid_public,
|
||||
)
|
||||
from scripts.research.guided_collect_holdout_replay import ( # noqa: E402
|
||||
TODAY,
|
||||
_hhmm,
|
||||
_minutes,
|
||||
load_cases,
|
||||
precision_gate,
|
||||
)
|
||||
from scripts.research.minute_resolution_sweep import scoring_request_for # noqa: E402
|
||||
from scripts.research.probe_supply_after_six import ( # noqa: E402
|
||||
ASK_COUNT,
|
||||
apply_answer,
|
||||
optimal_answer,
|
||||
top1_hit,
|
||||
)
|
||||
|
||||
RADII = (10, 30, 60)
|
||||
PRIORS = ("raw", "percent")
|
||||
REPORT_JSON = ROOT / "docs" / "research" / "rectification_typed_event_scoring_2026_09_29.json"
|
||||
SHIFT_SEED = 20260929
|
||||
SHIFT_REPEATS = 10
|
||||
TYPED_SOURCE = "typed_event_research"
|
||||
|
||||
ORIGINAL_BLOCKED_YEARS = event_probes._existence_blocked_years
|
||||
|
||||
|
||||
@contextmanager
|
||||
def narrowed_blocking() -> Iterator[None]:
|
||||
"""R3: an already-known year blocks only itself in the same domain."""
|
||||
event_probes._existence_blocked_years = lambda _domain, known_years: set(known_years)
|
||||
try:
|
||||
yield
|
||||
finally:
|
||||
event_probes._existence_blocked_years = ORIGINAL_BLOCKED_YEARS
|
||||
|
||||
|
||||
def _percent(scores: dict[str, float]) -> dict[str, float]:
|
||||
total = sum(max(value, 0.0) for value in scores.values())
|
||||
if total <= 0:
|
||||
return {key: 0.0 for key in scores}
|
||||
# Integer percent like `_relative_support_proportional` (largest remainder is not
|
||||
# needed for ranking; plain rounding keeps the research copy independent).
|
||||
return {key: float(round(max(value, 0.0) / total * 100)) for key, value in scores.items()}
|
||||
|
||||
|
||||
def _event_month(event: dict[str, Any]) -> int | None:
|
||||
raw = str(event.get("date_start") or "")
|
||||
return int(raw[5:7]) if len(raw) >= 7 and raw[4] == "-" else None
|
||||
|
||||
|
||||
def _event_year(event: dict[str, Any]) -> int | None:
|
||||
raw = str(event.get("date_start") or "")
|
||||
return int(raw[:4]) if raw[:4].isdigit() else None
|
||||
|
||||
|
||||
def _shift_month(event: dict[str, Any], months: int) -> dict[str, Any]:
|
||||
year, month = _event_year(event), _event_month(event)
|
||||
if year is None or month is None:
|
||||
return event
|
||||
index = year * 12 + (month - 1) + months
|
||||
new_year, new_month = divmod(index, 12)
|
||||
stamp = f"{new_year:04d}-{new_month + 1:02d}-15"
|
||||
return {**event, "date_start": stamp, "date_end": stamp}
|
||||
|
||||
|
||||
def typed_event_probes(
|
||||
request: dict[str, Any],
|
||||
contexts: Sequence[dict[str, Any]],
|
||||
*,
|
||||
include_year: bool,
|
||||
events: Sequence[dict[str, Any]] | None = None,
|
||||
) -> tuple[list[dict[str, Any]], dict[str, int]]:
|
||||
"""Score each training typed event through the probe split, answered ``yes``."""
|
||||
full = [item for item in contexts if event_probes._context_time(item)]
|
||||
full.sort(key=lambda item: event_probes._clock(str(event_probes._context_time(item))))
|
||||
clusters = cluster_contexts_by_signature(full)
|
||||
reps = [cluster["representative"] for cluster in clusters if event_probes._scoreable(cluster["representative"])]
|
||||
if len(reps) < 2:
|
||||
reps = [item for item in full if event_probes._scoreable(item)]
|
||||
set_version = event_probes.candidate_set_version([cluster["times"] for cluster in clusters])
|
||||
source_events = list(events if events is not None else request["events"])
|
||||
holdout = set(holdout_event_ids(request["events"]))
|
||||
stats = {"training_day": 0, "training_year": 0, "split_day": 0, "split_year": 0}
|
||||
rows: list[dict[str, Any]] = []
|
||||
for event in source_events:
|
||||
if str(event.get("id")) in holdout:
|
||||
continue
|
||||
domain = event_probes.canonical_domain(event.get("domain"))
|
||||
if domain not in event_probes.DOMAIN_CATALOG:
|
||||
continue
|
||||
precision = str(event.get("precision") or "")
|
||||
year = _event_year(event)
|
||||
if year is None:
|
||||
continue
|
||||
if precision in {"day", "month"}:
|
||||
stats["training_day"] += 1
|
||||
month = _event_month(event)
|
||||
key = "split_day"
|
||||
elif precision == "year" and include_year:
|
||||
stats["training_year"] += 1
|
||||
month = None
|
||||
key = "split_year"
|
||||
else:
|
||||
continue
|
||||
row = event_probes._evaluate_contexts(
|
||||
reps,
|
||||
birth_date=str(request["birth_date"]),
|
||||
domain=domain,
|
||||
year=year,
|
||||
month=month,
|
||||
source=TYPED_SOURCE,
|
||||
clusters=clusters,
|
||||
set_version=set_version,
|
||||
)
|
||||
if row is None:
|
||||
continue
|
||||
stats[key] += 1
|
||||
rows.append(row)
|
||||
return rows, stats
|
||||
|
||||
|
||||
def replay(
|
||||
*,
|
||||
rows: Sequence[dict[str, Any]],
|
||||
contexts: Sequence[dict[str, Any]],
|
||||
typed: Sequence[dict[str, Any]],
|
||||
probes: Sequence[dict[str, Any]],
|
||||
true_time: str,
|
||||
prior_mode: str,
|
||||
) -> dict[str, Any]:
|
||||
raw = raw_signature_clusters(contexts)
|
||||
by_time = {stamp: row for row in rows if (stamp := _hhmm(row.get("time")))}
|
||||
merged, _trace = merge_adjacent_traced(raw, by_time)
|
||||
public = public_from_clusters(merged, rows)
|
||||
reps = [str(row["time"])[:5] for row in public]
|
||||
base = {stamp: float(row.get("score") or 0) for row in public if (stamp := _hhmm(row.get("time")))}
|
||||
scores = _percent(base) if prior_mode == "percent" else dict(base)
|
||||
conflicts = {time: 0 for time in reps}
|
||||
eliminated: set[str] = set()
|
||||
for probe in typed:
|
||||
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, "yes", reps)
|
||||
asked = 0
|
||||
for probe in list(probes)[:ASK_COUNT]:
|
||||
answer = optimal_answer(probe, true_time)
|
||||
if answer is None:
|
||||
continue
|
||||
asked += 1
|
||||
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, answer, reps)
|
||||
posterior = [{**row, "score": scores.get(str(row["time"])[:5], row.get("score") or 0)} for row in public]
|
||||
valid = still_valid_public(posterior, scores, eliminated, lead=SEPARATION_LEAD)
|
||||
delivery = delivery_from_public(valid)
|
||||
start, end = delivery.get("start"), delivery.get("end")
|
||||
inside = start is not None and end is not None and _minutes(start) <= _minutes(true_time) <= _minutes(end)
|
||||
active = [time for time in reps if time not in eliminated]
|
||||
gate = precision_gate(valid, scores)
|
||||
return {
|
||||
"width": delivery.get("width"),
|
||||
"truth_in_range": bool(inside),
|
||||
"top1": bool(top1_hit(scores, active, true_time, merged)),
|
||||
"tied": bool(gate["tied_for_first"]),
|
||||
"questions": asked,
|
||||
"typed_scored": len(typed),
|
||||
"probe_supply": len(probes),
|
||||
}
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], radius: int) -> dict[str, Any]:
|
||||
true_time = str(case["birth"]["time"])[:5]
|
||||
request = scoring_request_for(case, radius)
|
||||
contexts = compute_candidate_static_contexts(request)
|
||||
built = build_event_contribution_matrix(request, static_contexts=contexts)
|
||||
rows = score_from_matrix(request, built)
|
||||
times = [stamp for row in rows if (stamp := _hhmm(row.get("time")))]
|
||||
|
||||
def probes_now() -> list[dict[str, Any]]:
|
||||
return event_probes.discriminating_event_probes(
|
||||
{**request, "refresh_probes": False, "asked_probe_keys": []}, built,
|
||||
scan=window_scan(built), candidate_times=times, representative_time=true_time, today=TODAY,
|
||||
)
|
||||
|
||||
probes = probes_now()
|
||||
with narrowed_blocking():
|
||||
probes_narrow = probes_now()
|
||||
assert event_probes._existence_blocked_years is ORIGINAL_BLOCKED_YEARS
|
||||
|
||||
typed_day, stats = typed_event_probes(request, contexts, include_year=False)
|
||||
typed_all, stats_all = typed_event_probes(request, contexts, include_year=True)
|
||||
|
||||
rng = random.Random(f"{SHIFT_SEED}:{case.get('case_id')}:{radius}")
|
||||
holdout = set(holdout_event_ids(request["events"]))
|
||||
day_events = [
|
||||
index for index, event in enumerate(request["events"])
|
||||
if str(event.get("precision")) in {"day", "month"} and str(event.get("id")) not in holdout
|
||||
]
|
||||
shifted: dict[str, list[dict[str, Any]]] = {"1": [], "2": []}
|
||||
for k in (1, 2):
|
||||
if len(day_events) < k:
|
||||
continue
|
||||
for _ in range(SHIFT_REPEATS):
|
||||
picks = rng.sample(day_events, k)
|
||||
events = [
|
||||
_shift_month(event, rng.choice((-3, -2, -1, 1, 2, 3))) if index in picks else event
|
||||
for index, event in enumerate(request["events"])
|
||||
]
|
||||
moved, _ = typed_event_probes(request, contexts, include_year=False, events=events)
|
||||
shifted[str(k)].append({"typed": moved})
|
||||
|
||||
out: dict[str, Any] = {"case_id": case.get("case_id"), "radius": radius, "typed_stats": {**stats, **{
|
||||
"training_year": stats_all["training_year"], "split_year": stats_all["split_year"]}},
|
||||
"probe_supply": len(probes), "probe_supply_narrow": len(probes_narrow),
|
||||
"asked_keys_changed_narrow": [p["semantic_key"] for p in probes[:ASK_COUNT]]
|
||||
!= [p["semantic_key"] for p in probes_narrow[:ASK_COUNT]],
|
||||
"arms": {}}
|
||||
for prior in PRIORS:
|
||||
def run(typed: Sequence[dict[str, Any]], pool: Sequence[dict[str, Any]]) -> dict[str, Any]:
|
||||
return replay(rows=rows, contexts=contexts, typed=typed, probes=pool, true_time=true_time, prior_mode=prior)
|
||||
|
||||
arms = {
|
||||
"B0": run([], probes),
|
||||
"R1d": run(typed_day, probes),
|
||||
"R1y": run(typed_all, probes),
|
||||
"R3": run([], probes_narrow),
|
||||
"R3+R1d": run(typed_day, probes_narrow),
|
||||
}
|
||||
for k, trials in shifted.items():
|
||||
results = [run(item["typed"], probes) for item in trials]
|
||||
if results:
|
||||
arms[f"R2k{k}"] = {
|
||||
"trials": len(results),
|
||||
"truth_out": sum(1 for row in results if not row["truth_in_range"]),
|
||||
"top1": sum(1 for row in results if row["top1"]),
|
||||
"median_width": statistics.median(row["width"] for row in results if row["width"] is not None),
|
||||
}
|
||||
out["arms"][prior] = arms
|
||||
return out
|
||||
|
||||
|
||||
def summarize(rows: Sequence[dict[str, Any]], prior: str, radius: int, arm: str) -> dict[str, Any]:
|
||||
subset = [row["arms"][prior][arm] for row in rows if row.get("radius") == radius and not row.get("error")
|
||||
and arm in row["arms"][prior]]
|
||||
n = len(subset)
|
||||
if not n:
|
||||
return {"n": 0}
|
||||
if arm.startswith("R2k"):
|
||||
trials = sum(item["trials"] for item in subset)
|
||||
return {
|
||||
"n": n,
|
||||
"trials": trials,
|
||||
"truth_out_rate": round(sum(item["truth_out"] for item in subset) / trials, 4),
|
||||
"cases_with_truth_out": sum(1 for item in subset if item["truth_out"] > 0),
|
||||
"top1_rate": round(sum(item["top1"] for item in subset) / trials, 4),
|
||||
"median_width": statistics.median(item["median_width"] for item in subset),
|
||||
}
|
||||
return {
|
||||
"n": n,
|
||||
"top1": round(sum(item["top1"] for item in subset) / n, 4),
|
||||
"truth_in_range": round(sum(item["truth_in_range"] for item in subset) / n, 4),
|
||||
"median_width": statistics.median(item["width"] for item in subset if item["width"] is not None),
|
||||
"tie_rate": round(sum(item["tied"] for item in subset) / n, 4),
|
||||
"mean_questions": round(statistics.mean(item["questions"] for item in subset), 2),
|
||||
"mean_typed_scored": round(statistics.mean(item["typed_scored"] for item in subset), 2),
|
||||
"mean_probe_supply": round(statistics.mean(item["probe_supply"] for item in subset), 2),
|
||||
}
|
||||
|
||||
|
||||
def verdict(summary: dict[str, Any], prior: str, arm: str) -> str:
|
||||
"""Gate: top1 not lower AND truth-in-range not lower AND median width lower, all radii."""
|
||||
passes = []
|
||||
for radius in RADII:
|
||||
base = summary[prior][str(radius)]["B0"]
|
||||
cand = summary[prior][str(radius)][arm]
|
||||
if not cand.get("n"):
|
||||
return "uncertain"
|
||||
passes.append(
|
||||
cand["top1"] >= base["top1"]
|
||||
and cand["truth_in_range"] >= base["truth_in_range"]
|
||||
and cand["median_width"] < base["median_width"]
|
||||
)
|
||||
return "benefit" if all(passes) else "no_benefit"
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--limit", type=int, default=0)
|
||||
parser.add_argument("--json-out", default=str(REPORT_JSON))
|
||||
args = parser.parse_args()
|
||||
cases = load_cases()
|
||||
if args.limit:
|
||||
cases = cases[: args.limit]
|
||||
rows: list[dict[str, Any]] = []
|
||||
for case in cases:
|
||||
for radius in RADII:
|
||||
try:
|
||||
rows.append(evaluate_case(case, radius))
|
||||
except Exception as exc: # noqa: BLE001
|
||||
rows.append({"case_id": case.get("case_id"), "radius": radius, "error": f"{type(exc).__name__}: {exc}"})
|
||||
arms = ("B0", "R1d", "R1y", "R3", "R3+R1d", "R2k1", "R2k2")
|
||||
summary = {prior: {str(radius): {arm: summarize(rows, prior, radius, arm) for arm in arms}
|
||||
for radius in RADII} for prior in PRIORS}
|
||||
verdicts = {prior: {arm: verdict(summary, prior, arm) for arm in ("R1d", "R1y", "R3", "R3+R1d")} for prior in PRIORS}
|
||||
assert event_probes._existence_blocked_years is ORIGINAL_BLOCKED_YEARS
|
||||
payload = {
|
||||
"generated_for": "TASK-rectification-typed-event-scoring-research-20260929",
|
||||
"today": TODAY.isoformat(),
|
||||
"holdout": "references/real_case_calibration/minute_rectification_holdout_v4.json",
|
||||
"ask_count": ASK_COUNT,
|
||||
"separation_lead": SEPARATION_LEAD,
|
||||
"shift_seed": SHIFT_SEED,
|
||||
"shift_repeats": SHIFT_REPEATS,
|
||||
"open_set_not_blind": True,
|
||||
"summary": summary,
|
||||
"verdicts": verdicts,
|
||||
"rows": rows,
|
||||
"errors": [row for row in rows if row.get("error")],
|
||||
}
|
||||
out = Path(args.json_out)
|
||||
out.write_text(json.dumps(payload, ensure_ascii=False, indent=1, sort_keys=True) + "\n", encoding="utf-8")
|
||||
print(json.dumps({"summary": summary, "verdicts": verdicts}, ensure_ascii=False, indent=1))
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Reference in New Issue
Block a user