research(rectification): typed events scored as answered probes — no benefit (BUG-1089)

v4 open holdout, ±10/±30/±60, raw and percent priors: only 4/9/14 of 57
day-precision training events split candidates; widths unchanged, top-1
drops. Narrowed year blocking (R3) mixed. All arms no_benefit; no
implementation brief recommended. Two runs byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
Jesse_Chen
2026-09-29 10:06:13 +08:00
co-authored by Claude Opus 5.5
parent 56b51e2170
commit bcf0c86feb
6 changed files with 8763 additions and 8 deletions
+11 -7
View File
@@ -14647,25 +14647,29 @@
## BUG-1088 | 学业质量题一律问「那次上大学」,只记得年份的经历显示成「YYYY 年 1 月」
- 状态:investigating
- 状态:resolved(代码 + 回归测试 + 重新冻结;未部署,真机待合入 staging 后复核)
- 首次发现 / 最近更新:2026-09-29 / 2026-09-29
- 影响面:`scripts/rectification/event_probes.py::_quality_user_meaning`、`_display_date_label` / `_event_month`(冻结计分文件,ERR-110)。
- 影响面:`scripts/rectification/event_probes.py::_quality_user_meaning`、`_display_date_label`(冻结计分文件,ERR-110)。
- 用户现象:用户说的是某年一次艺考,卡片问「YYYY 年 1 月那次上大学,更接近如愿、将就调剂…」。
- 根因:学业模板写死「上大学」,不看 `event_kind`;年精度经历 `date_start = YYYY-01-01`,取月不看 `precision`。
- 修复:待执行(研究单 R0,改文字后按 ERR-110 用新 freeze 路径重新冻结)。
- 相关记录:ERR-110
- 触发条件:学业经历类型为 education_start / change / interruption,且出现质量题;年精度经历存为 `YYYY-01-01`。
- 根因:学业模板写死「上大学」,不看 `event_kind`;`_display_date_label` 用 `_event_month` 取月,不看 `precision`。
- 修复:学业分支按类型称呼(升学 / 学业变动 / 学业中断,兜底「学业经历」);`precision == "year"` 的标签只写「YYYY 年」。`_event_month` 与 split hash 不变,计分身份不变。
- 验证:`tests/test_rectification_event_probes.py::QualityWordingTests`(4 条);同机 A/B(`scripts/research/quality_wording_ab.py`,`PYTHONHASHSEED=0`)只有 `user_meaning` / `display_date_label` 文本变化,分数、候选决策、题目 id / split / outcomes 逐字节相同;按 ERR-110 新建冻结 `docs/research/sealed_holdout_rerun_quality_wording_2026_09_29.freeze.json`、`docs/research/reported_offset_quality_wording_2026_09_29.freeze.json`,重放 20 / 900 次结果与 09-21 完全一致;验证完整性门禁 83 passed。
- 防复发:年精度经历不得在任何用户文案里补月份;新增学业经历类型时同步 `EDUCATION_QUALITY_EVENT_NAME`。质量题四个选项仍含「录取」字样,未改。
- 相关记录:ERR-110、BUG-1089
- 复发自:无
- 修复版本:未修
- 修复版本:分支 `codex/rectification-typed-event-research-20260929`,未合入
## BUG-1089 | 带年月的打字经历不按选择题规则计分,还挡掉同领域相邻年份的选择题
- 状态:investigating(离线研究单待执行;不上线)
- 状态:investigating(离线研究已完成,结论不改代码;等止血单 BUG-1084 落地后由 Claude 决定是否关闭)
- 首次发现 / 最近更新:2026-09-29 / 2026-09-29
- 影响面:`core/build-state.ts::buildInferenceState`(evidence「是」跳过)、`event_probes.py::_existence_blocked_years`。
- 用户现象:同一件「某年某月工作变动」,点选卡答「是」动 ±2 分,自己打字说约等于 0;说了之后那一年附近的选择题也不再出。
- 已确认事实(虚构 12 例):12 件经历里去掉 4 件记得到月的,选择题 5/12 例变多,最终宽度中位 17 → 13 分钟。
- 根因:计分通道不对称,见 `docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`。
- 修复:未修。研究单量「打字经历按选择题边界规则计分」在 v4 开放集三档是否过门;过门才立实现单,合入前产品确认。
- 研究结果(2026-09-29,`docs/research/rectification_typed_event_scoring_2026_09_29.md`,v4 开放集 20 例,开放集回放不是准确率):按选择题同一判定函数给打字经历计分(R1),57 件日精度训练经历在 ±10 / ±30 / ±60 下只有 4 / 9 / 14 件能把候选分到两边,宽度中位三档不变,头名在 raw 口径下三档都降,判 `no_benefit`;挡题缩到同领域同年(R3)宽窗头名 +0.10~0.15 但 ±10 头名 −0.15,判 `no_benefit`;年月记错 ±1–3 月挤出真值 0–2.9%。结论:计分通道不对称在代码上属实,但打字经历本身多半分不开候选,改计分不划算;不立实现单。
- 防复发:不得在研究未过门时改先验尺度或淘汰线(BUG-560)。
- 相关记录:BUG-560、BUG-1048、BUG-1084
- 复发自:无
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,104 @@
# 打字经历按选择题规则计分:离线研究(2026-09-29)
- 任务书:`docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`;进度:`docs/tasks/PROGRESS-rectification-typed-event-research-20260929.md`
- 分支:`codex/rectification-typed-event-research-20260929`,基线 `origin/staging` @ `56722063`。
- 性质:**离线测量,不改线上默认值。** 研究补丁(R3 的 `_existence_blocked_years`)只在一次调用内替换,调用后与脚本末尾都断言已恢复原函数。
- 数据:`references/real_case_calibration/minute_rectification_holdout_v4.json`,20 例公开 Rodden-AA。v4 的事件是 59 件日精度、84 件年精度,**没有月精度**;每例留一件作 holdout,训练集为日精度 57 件、年精度 67 件。**这是开放集回放,不是盲测,下面的数字不是准确率。**
- 口径:ayanamsa `raman`,node mode `mean`,步长 2 分钟,半径 ±10 / ±30 / ±60;六题回放(`ASK_COUNT = 6`,按真值作答);交付区间 = 未淘汰且落后头名不足 8 分的簇的并集。两种先验都报:`raw`(引擎行原始分,09-14 / 09-26 研究口径)与 `percent`(公开代表按正比例换成整数百分比,线上 `relative_support` 口径)。
- 复跑:`PYTHONHASHSEED=0 python3 scripts/research/typed_event_as_probe.py`(约 1.5 分钟)→ `docs/research/rectification_typed_event_scoring_2026_09_29.json`。同机两次运行 JSON 逐字节一致。
- 判门(任务书):三档**同时**满足「头名命中不降、真值在区间不降、宽度中位下降」才算有收益。
## 结论
| 项 | 一句话结论 | 是否建议立实现单 |
| --- | --- | --- |
| R0 学业质量题措辞(BUG-1088) | 已修:按经历类型称呼(升学 / 学业变动 / 学业中断),年精度经历不再显示「1 月」;计分输出逐字节不变,已按 ERR-110 用新路径重新冻结 | 已在本分支实现 |
| R1 打字经历当作已答「是」的选择题 | **不过门。** 打字经历绝大多数分不开候选:±10 / ±30 / ±60 下,57 件日精度训练经历只有 4 / 9 / 14 件(7% / 16% / 25%)能把代表候选分到两边;宽度中位三档都不变,头名在 raw 口径下三档都降 | **不建议** |
| R2 经历年月记错 ±1–3 个月 | 挤出真值的比例 0–2.9%,与 09-26 答错两题(±30 约 7%)相比不高;但 R1 本身没有收益,R2 不改变结论 | 不适用 |
| R3 挡题范围缩到「同领域同年」 | **不过门,但宽窗有信号。** raw 口径下 ±30 头名 0.50 → 0.65、±60 0.35 → 0.45;±10 头名 0.75 → 0.60(宽度 15 → 14)。宽度三档基本不变。60 格中有 45 格实际问到的六题因此换了 | **不建议直接实现**;若要继续,需另立研究,先解释 ±10 的头名下降 |
**给产品的话:** 「打字说的经历不参与收窄」在代码里确实是两条不同的通道,但把打字经历改成按选择题规则计分**几乎没有用处**:一件你记得清清楚楚的事,多半恰好落在所有候选时刻都一样的那段大运里,本来就分不开候选。能分开候选的,是系统专门挑出来的「边界月份」选择题。这也解释了 09-26 研究 R3c 的发现:边界月份里只有 3–4.5% 能把代表候选分开。所以止血单「门开后不再索要经历、选择题问完就出卡」的方向是对的,本研究不支持再往打字经历上加计分。
## R1 · 打字经历当作已答选择题
**方法。** 对每件训练经历,用出题器的同一个判定函数 `event_probes._evaluate_contexts`(代表候选、签名聚类、同一套集合版本),以经历的领域和年月(日精度取所在月;R1y 的年精度取 `month=None`,即现有「整年」判定)生成一道「题」,按「是」作答(±2,强冲突计数同线上),然后照常回放六道选择题。函数返回 `None` 的经历(候选在那个月没有分边)不计分。
**能分开候选的经历(每档 20 例合计):**
| 半径 | 日精度训练经历 | 能分边的 | 年精度训练经历 | 能分边的 |
| --- | ---: | ---: | ---: | ---: |
| ±10 | 57 | 4 | 67 | 5 |
| ±30 | 57 | 9 | 67 | 10 |
| ±60 | 57 | 14 | 67 | 14 |
**raw 口径:**
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 并列率 | 每例提问 | 每例计分经历 |
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: |
| ±10 | B0 | 0.75 | 1.00 | 15 | 0.05 | 4.25 | 0 |
| ±10 | R1d | 0.70 | 1.00 | 15 | 0.05 | 4.25 | 0.20 |
| ±10 | R1y | 0.60 | 1.00 | 14 | 0.05 | 4.25 | 0.45 |
| ±30 | B0 | 0.50 | 1.00 | 33 | 0.05 | 5.40 | 0 |
| ±30 | R1d | 0.45 | 1.00 | 33 | 0.05 | 5.40 | 0.45 |
| ±30 | R1y | 0.45 | 1.00 | 33 | 0.05 | 5.40 | 0.95 |
| ±60 | B0 | 0.35 | 1.00 | 56 | 0.05 | 5.40 | 0 |
| ±60 | R1d | 0.30 | 1.00 | 56 | 0.05 | 5.40 | 0.70 |
| ±60 | R1y | 0.30 | 1.00 | 56 | 0.05 | 5.40 | 1.40 |
**percent 口径(线上先验尺度):**
| 半径 | 方案 | 头名命中 | 真值在区间 | 宽度中位 | 并列率 |
| --- | --- | ---: | ---: | ---: | ---: |
| ±10 | B0 | 0.50 | 1.00 | 11 | 0.40 |
| ±10 | R1d | 0.45 | 1.00 | 11 | 0.40 |
| ±10 | R1y | 0.35 | 1.00 | 11 | 0.45 |
| ±30 | B0 | 0.40 | 1.00 | 33 | 0.55 |
| ±30 | R1d | 0.40 | 1.00 | 33 | 0.55 |
| ±30 | R1y | 0.40 | 1.00 | 33 | 0.45 |
| ±60 | B0 | 0.15 | 1.00 | 53 | 0.75 |
| ±60 | R1d | 0.20 | 1.00 | 53 | 0.75 |
| ±60 | R1y | 0.15 | 1.00 | 53 | 0.75 |
**读法。**
1. 每例平均只有 0.2–0.7 件日精度经历能计分,宽度中位三档都不动。多出来的 ±2 只改排序,而且往往改错:raw 口径下头名三档都下降。
2. 真值从未被挤出区间(真值在区间保持 1.00):单件经历的 ±2 不够 8 分淘汰线。
3. percent 口径的并列率高(0.40–0.75),是整数百分比把相近原始分压成同一个数所致,与本研究的方案无关;这是线上先验尺度(BUG-560)的另一面,只记录,不处理。
## R2 · 经历年月记错 ±1–3 个月
每例每档随机挑 1 件或 2 件日精度训练经历,把月份挪 ±1–3 个月(种子 `20260929`,每格重复 10 次),按 R1d 计分后回放。k=2 只在日精度训练经历 ≥2 的 14 例上跑。
| 口径 | 半径 | k | 回放次数 | 挤出真值比例 | 至少一次挤出的例子 | 宽度中位 |
| --- | --- | ---: | ---: | ---: | ---: | ---: |
| raw | ±10 | 1 / 2 | 200 / 140 | 0.0% / 2.9% | 0 / 1 | 15 / 15 |
| raw | ±30 | 1 / 2 | 200 / 140 | 1.5% / 2.1% | 1 / 1 | 33 / 36 |
| raw | ±60 | 1 / 2 | 200 / 140 | 0.0% / 0.0% | 0 / 0 | 56 / 54 |
| percent | ±10 | 1 / 2 | 200 / 140 | 0.0% / 2.9% | 0 / 1 | 11 / 12 |
| percent | ±30 | 1 / 2 | 200 / 140 | 1.5% / 2.9% | 1 / 1 | 33 / 34 |
| percent | ±60 | 1 / 2 | 200 / 140 | 0.0% / 0.7% | 0 / 1 | 53 / 53 |
挤出率低,原因同 R1:挪错后的月份多半也分不开候选,等于没计分。R2 不改变 R1 的结论。
## R3 · 挡题范围缩到「同领域同年」
线上 `_existence_blocked_years` 会把已入账年份及同领域相邻年份(`EXISTENCE_NEARBY_YEARS`)的存在题挡掉。研究补丁改为只挡「同领域同一年」,其余不变。
| 口径 | 半径 | B0 头名 / 宽度 | R3 头名 / 宽度 | R3+R1d 头名 / 宽度 | R3 判定 |
| --- | --- | --- | --- | --- | --- |
| raw | ±10 | 0.75 / 15 | 0.60 / 14 | 0.55 / 14 | 头名降 |
| raw | ±30 | 0.50 / 33 | 0.65 / 33 | 0.60 / 33 | 宽度未降 |
| raw | ±60 | 0.35 / 56 | 0.45 / 55 | 0.40 / 54 | 过 |
| percent | ±10 | 0.50 / 11 | 0.45 / 11 | 0.40 / 11 | 头名降 |
| percent | ±30 | 0.40 / 33 | 0.50 / 31 | 0.50 / 30 | 过 |
| percent | ±60 | 0.15 / 53 | 0.15 / 54 | 0.20 / 53 | 宽度升 |
真值在区间:除 percent ±60 R3+R1d 为 0.95 外全部 1.00。出题供给(截断到 `MAX_PROBES = 8` 之后)几乎不变(每例 5.1 / 7.1 / 7.2 → 5.1 / 7.1 / 7.15),但 60 格中有 45 格实际问到的前六题换了:放开相邻年份后,出题器会挑到别的年份。
**读法。** 宽窗(±30 / ±60)头名有 +0.10 到 +0.15 的改善,但样本只有 20 例(每例 0.05),±10 头名下降同样幅度,三档不同时过门。按任务书的门槛判 `no_benefit`。若产品想追,下一步应是另立研究:只在窗口 ≥30 分钟时放开,并先解释 ±10 下降是换题带来的随机波动还是系统性问题。
## 偏离与说明
- 基线头名与 09-26 研究略有差异(raw B0:本文 0.75 / 0.50 / 0.35,09-26 为 0.80 / 0.55 / 0.40;宽度 15 / 33 / 56 完全一致)。本脚本自己实现了回放循环(为了在六题之前插入经历计分),`top1_hit` 用的是相邻合并后的簇;各方案之间用的是同一条回放链路,比较有效,但数字不要与 09-26 逐格对照。
- 任务书以「记得到月」为主设想;v4 没有月精度事件,日精度事件按所在月判定,等价于月精度判定。
- 任务书写的 R1 年精度口径「只在跨年边界时计分」,实现为出题器现有的整年判定(`_evaluate_contexts(month=None)`,事件日取 7 月 1 日),它已经只在候选整年激活状态不同时才分边。
@@ -0,0 +1,51 @@
# PROGRESS · 打字经历按选择题规则计分(离线研究)+ 学业质量题措辞(2026-09-29)
- 任务书:`docs/tasks/TASK-rectification-typed-event-scoring-research-20260929.md`
- 执行:Claude(直接执行模式,fork 子代理),分支 `codex/rectification-typed-event-research-20260929`,基线 `origin/staging` @ `56722063`。未推送。
- 研究报告:`docs/research/rectification_typed_event_scoring_2026_09_29.md` + `.json`
## 开工预检
- `python3 scripts/pre_work_check.py --remote-timeout 8 --command-timeout 45`:exit 0;python_runtime / fragment_scan / external_engine_adapters / focused_tests 全部 ok,remote_visibility `verified`。
- 读 `docs/research/pre_work_error_ledger.md` ERR-110(冻结文件改动须重新冻结)、ERR-111(固定 `PYTHONHASHSEED`)。本机无 `.venv`,全部用系统 `python3`。
## R0 · 学业质量题措辞(BUG-1088)
- `scripts/rectification/event_probes.py`:`_display_date_label` 对 `precision == "year"` 不取月;`_quality_user_meaning` 学业分支按 `event_kind` 称呼(`EDUCATION_QUALITY_EVENT_NAME`:升学 / 学业变动 / 学业中断,兜底「学业经历」),句式「…那次X,结果更接近如愿、将就调剂、发挥失常还是说不清」。四个选项标签未改(仍含「录取」字样,属选项文案,留待产品决定)。
- **计分身份不变**:`_event_month` 未改,`candidate_split_hash`、题目 `month` 字段仍带存储的月份;新增测试锁住这一点。
- 同机 A/B(`scripts/research/quality_wording_ab.py`,`PYTHONHASHSEED=0`,v4 前 4 例 ±10 + 2 个虚构出生带学业质量事件):基线连跑两次逐字节一致;改后对比,**只有 2 个叶子字段类型变化**:`user_meaning`(2 处)与 `display_date_label`(2 处),均在虚构例的 `known_event_quality` 题上;引擎行分数、公开候选决策、全部题目的 id / split hash / outcomes 逐字节相同。
- `2010 年 1 月那次上大学…` → `2010 年那次学业中断…`(年精度 + interruption)
- `2011 年 1 月那次上大学…` → `2011 年那次升学…`(年精度 + start)
- 重新冻结(ERR-110,不覆盖旧记录):
- `scripts/research/sealed_holdout_rerun.py` `FREEZE/REPORT` → `docs/research/sealed_holdout_rerun_quality_wording_2026_09_29.{freeze.json,json}`;重放 20 例,metrics 与 20 条 trials 与 09-21 记录**完全相同**(top1 0.45 / top3 0.50 / MAE 6.45)。
- `scripts/research/reported_offset_sweep.py` `FREEZE/REPORT` → `docs/research/reported_offset_quality_wording_2026_09_29.{freeze.json,json}`;900 trials 与 summary 与 09-21 **完全相同**。
- `references/rectification_sealed_holdout.v1.json`:`current_tree_scorer`、`current_tree_fixed_protocol_rerun`、`current_tree_reported_offset_replay` 指向新记录;旧的 09-21 记录移入 `previous_fixed_protocol_rerun` / `previous_reported_offset_replay`;`evaluated_on` 2026-09-29。遗留 12 文件哈希 `c913c6f0…` 不变,只有扩展身份(event_probes.py 与两份研究脚本)变化。
- `tests/test_sealed_holdout_contract_freshness.py` 的 `IDENTITY_FORBIDDEN_PREFIXES` 补两条新前缀(同 09-21 先例)。
- 新测试:`tests/test_rectification_event_probes.py::QualityWordingTests`(4 条:三种类型不叫上大学、年精度无月、月精度保留月、split 月份不变)。
## R1–R3 · 离线研究(BUG-1089)
脚本 `scripts/research/typed_event_as_probe.py`,`PYTHONHASHSEED=0` 连跑两次 JSON 逐字节一致(各 1 分 25 秒)。结论全部 `no_benefit`,**不建议立实现单**。数字见研究报告;要点:
| 项 | raw 口径(±10 / ±30 / ±60) | 判定 |
| --- | --- | --- |
| B0 头名 / 宽度 | 0.75 / 0.50 / 0.35;15 / 33 / 56 | 基线 |
| R1d 头名 / 宽度 | 0.70 / 0.45 / 0.30;15 / 33 / 56 | no_benefit |
| R1y 头名 / 宽度 | 0.60 / 0.45 / 0.30;14 / 33 / 56 | no_benefit |
| R3 头名 / 宽度 | 0.60 / 0.65 / 0.45;14 / 33 / 55 | no_benefit(±10 头名降) |
| 能分边的日精度训练经历 | 4 / 9 / 14(共 57 件) | — |
| R2 挤出真值(k=1 / k=2) | 0–1.5% / 2.1–2.9% | — |
percent 口径结论相同(判定全部 `no_benefit`)。
## 测试
- 定向:`tests/test_rectification_validation_integrity_gate.py`、`tests/test_sealed_holdout_contract_freshness.py`、`tests/test_reported_offset_research.py`、`tests/test_rectification_event_probes.py`:83 passed(改前重新冻结前为 7 失败,全部是冻结身份不匹配)。
- 快速门 `python3 scripts/run_quality_gate.py --profile quick`(系统 python3,本机无 `.venv`):JSON / 审计 / 碎片 / BPHS 各步通过;pytest **1001 passed, 1 skipped**(含上面四个文件)。随后 `npm test` 步骤 exit 127:本工作树没有 `frontend/node_modules`,npm 找不到测试依赖。**前端三步(npm test / lint / build)未跑,记为环境缺口,不算通过。** 本轮没有改任何前端文件。
- 隐私:`tests/test_repo_privacy_markers.py` 通过。
## 偏离
- 基线回放头名与 09-26 研究差 0.05(宽度一致),原因见研究报告「偏离与说明」;方案之间可比,不与 09-26 逐格对照。
- v4 无月精度事件,日精度按所在月判定。
- 未改质量题四个选项标签;前端 `collection-question-pool.ts` 也有「上大学」措辞(采集题示例),不在本单范围,未改。
+1 -1
View File
@@ -254,7 +254,7 @@
| `TASK-self-edit-avatar-menu-20260928.md` | `PROGRESS-self-edit-avatar-menu-20260928.md` | **本人可编辑 + 各页头像菜单 + 报告页去说明**:`/people` 本人可编辑出生资料(BUG-1081,`ab6c55f8` 漏掉);次级页头像弹同一账户菜单、弹窗项跳 `/?account=`;我的报告删顶部说明与统计 | 已实现待验收(分支 `codex/self-edit-avatar-menu-20260928`,未推送) | BUG-1081 |
| `TASK-mobile-chart-and-confirmed-edit-20260929.md` | `PROGRESS-mobile-chart-confirmed-edit-20260929.md` | **手机星盘 + 确认时间可改 + 天空定格**:iPhone 星盘被宽表撑出屏幕(BUG-1083);confirmed 状态改声明字段也重置排盘时间(推翻 BUG-264 约定);「那一刻的天空」改为重放汇聚→定格→右上角分享;去掉报告星图封面 | 已验收(Claude 2026-09-29),已推 staging | BUG-1083 |
| `TASK-rectification-futile-collect-stop-20260929.md` | — | **生时校正停掉无效补经历循环(止血)**:打字经历不收窄(BUG-560 后果),流程却一路索要,真机 22 件整窗不动;采集只为开闸、门开后只问点选卡问完即出卡;交付正文去吻合率;多段时旁白误报「范围没变」;交付轮带采集题 | 待领取 | BUG-1084~1087 |
| `TASK-rectification-typed-event-scoring-research-20260929.md` | — | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 待领取 | BUG-1088、1089 |
| `TASK-rectification-typed-event-scoring-research-20260929.md` | `PROGRESS-rectification-typed-event-research-20260929.md` | **打字经历按选择题规则计分(离线研究,不上线)**:计分通道不对称 + 已入账年份挡题;R0 学业质量题措辞 / 年精度显示成 1 月(冻结文件,需重新冻结) | 待验收(Claude 子代理直接执行,分支 `codex/rectification-typed-event-research-20260929`,未推送;R1–R3 均 no_benefit) | BUG-1088、1089 |
| `TASK-serif-headings-20260928.md` | `PROGRESS-serif-headings-20260928.md` | **全站标题改用自托管宋体、正文保持黑体**:产品推翻 BUG-737「CJK 不用衬线」结论(保留「声明的字体必须可加载」「不落系统宋体」两条);Noto Serif SC SemiBold 按通用规范汉字表 6500 字 unicode-range 切片自托管(改名 Jyotisha Serif SC),swap 不 preload;首页宋体流量 ≤300 KB;先于天空封面单 | 已实现待验收(分支 `codex/serif-headings-20260928`,未推送) | 不开新 BUG;BUG-737 追加说明 |
| `TASK-cend-ui-claude-alignment-20260916.md` | `PROGRESS-cend-ui-r1/r2/r3-20260916.md` | **C 端界面向 claude.ai 产品界面对齐(三轮串行 R1→R2→R3,都动 `globals.css`,不得并行)**:根因是 `frontend/CLAUDE_DESIGN.md` 扒的是 **claude.com 营销官网**,它自己在 Known Gaps 里写明 claude.ai 产品界面不在范围内,而 `DESIGN.md:3` 把它当成了产品界面的实现契约。**R1**:`--font-display` 里 Tiempos Headline / StyreneB **从未加载**(无 `@font-face`、`public/` 无字体、`layout.tsx` 只 vendor 了 Inter),中文标题全站落到 **宋体 / SimSun**,波及 20 处含助手回答的 h2/h3(BUG-737);亮色强调色 `#85432f` 与暗色 `#d78064` 不同源,产品拍板亮色换 **Claude coral `#cc785c`**,**易漏点**是 `globals.css:16` 的 `--color-ring` 硬编码在 `@theme inline` 里不跟随 `:root`,另有第四个 `:root` 亮色块(`:4358`)必须同步(BUG-738);`.composer-footer` 常驻 44px + 顶栏 68px + `--composer-reserve` 148px,每屏固定吃掉 216px,模型选择器移进输入框内部、删掉底栏、顶栏收到 46px 并删「分析对象」副标题。**R2**:空状态是营销落地页(hero 卡 + 两张 132px 入口大卡 + 3 列 156px 主题卡),输入框被压在 **800px 以上**内容之下,重排成「问候 + 居中输入框 + 两枚入口 pill + 一排 chip」。**R3**:侧栏两个 `<details>` 拍平成一条「最近」、星盘的两个入口(侧栏分组 + 账户菜单)收敛到一处、删掉逐条助手头像。**决策记录 D3 推翻 DESIGN.md「报告强调色与应用同源」一句**(报告刻意保留深棕)。原型图 https://claude.ai/code/artifact/da275da6-2954-4f50-99aa-32bb8694d38b(三套画面 + 明暗,页面标题就是建议字体栈的实际渲染)。环境缺口:无登录态无 Chrome,四项真机观感留 `docs/testing/`。BUG 段 737–738 | 待领取 | — |
| `TASK-cend-surfaces-claude-alignment-20260916.md` | `PROGRESS-cend-shell-20260916.md`、`PROGRESS-cend-report-20260916.md`、`PROGRESS-cend-rectification-20260916.md`、`PROGRESS-cend-chart-eph-20260916.md` | **次级页面对齐(上一单的续篇,R4→R5/R6,R7、R8 可并行)**:星盘 `/chart`、星历 `/ephemeris`、报告 `/reports` **各是脱离 app 外壳的独立全屏页**,顶部只有一个「返回对话」链接、侧栏整个消失,且三家各写了一套一模一样的 `*-shell`/`*-topbar`/`*-hero` 骨架——与上一单 E5 同根因(营销站 band 结构被套到产品界面)。**R4** 抽只读导航外壳 `AppNavRail`(只用现成的 `GET /api/sessions` + `GET /api/account`,会话行走 `sessionHref` 跳 `/?c=<uuid>`;**刻意不带**重命名/删除/收藏/归档——那套连着 `Home()` 的乐观更新与回滚,搬过来会撞 useState 增长门禁)。**R5** 星盘五 tab 下划线化 + 参数合表 + 行星表横向滚动;星历日期导航改 `‹ 日期 ›`。**R6** 报告中心卡片网格改行式列表;阅读页加常驻目录。**R7** 生时校正把可信区间从盘面板标题行提成常驻条(窄屏 `.is-compact` 下盘面板是 overlay,现在默认看不到区间),五个 `technique-audit` 折叠块收成两段。**R8** 设置内容区收窄(880px 弹窗里表单铺了 690px)、套餐卡三修饰符收敛成两态。**已解锁**:原挡路的设置单已于 `111b4a84`(BUG-698)合入。**两条不得回退**:BUG-698 的 `@supports (height: 1dvh)` 写法(重复声明回退会被 Lightning CSS 折叠)、BUG-616/617 的报告盘面 grid 实现。默认不占 BUG 号 | **R4–R8 全部已实现并验收合入** | R4:抽出 `AppNavRail`(只读,两个 GET,零写操作)+ `SecondaryShell`,三个次级页并入 app 外壳并删掉各自的 shell/topbar/hero;四个路由渲染标记**完全不变**(`/` `/chart` `/ephemeris` 仍 Static);CSS gzip −0.25%。`/reports/[reportId]` 留给 R6 与目录一起做。两处自身健壮性问题被测试抓到:`usePathname()` 可为 null、`fetch` 可能不存在。差点弄丢 BUG-717 的 eyebrow 文案(已放回)。R8:表单分区收窄到 440px(列表分区不变)、套餐卡三修饰符收敛成互斥的 `is-current` / `is-recommended`,`--highlighted` 删除改为滚动定位;手机端 `order:-1` 改挂 `[data-plan-alias]`(版位不是状态)。测试 3350→3354(净增 4),失败清单与基线逐条一致;`/` 仍 Static;我的干净构建实测 CSS gzip −3 字节。**遗留待产品拍板**:`?plan=` 深链现在完全没有视觉指向,只有滚动位置。R6:报告中心卡片网格改行式列表、阅读页并入外壳并把目录挪到右侧常驻。**任务书 E10 过期**——目录在 `cfcd369d` 就已存在,本轮是挪位置定稿而非从零加。挂外壳带出一个真实打印风险已处理:`.chat-app`/`.chat-panel` 是 `height:100%;overflow:hidden`,裸 `window.print()` 会把九节报告裁成一页,阅读页因此多挂一条只在挂载期生效的 print 样式解锁外壳。「生成中的分节进度」做不了——`REPORT_LIST_COLUMNS` 不返回节数,按 VOICE.md 不许前端编。R7:区间常驻条与盘面折叠收敛。**任务书 E9 也不准确**——对话区顶部早有常驻条 `RectificationTimeline` 且窄屏可见,真正只在盘面标题行的是**代表分钟**;因此没另造第二条,在既有条上补齐代表分钟与已答题数(与盘面同一次 `workingRectificationTime()` 调用)。折叠块实际是 **8 个**不是 5 个。**触发让步顺序第 5 条**:收窄进度未做——服务端无该字段,且 `candidate_range` 会放宽(BUG-572),前端相减会把一次放宽报成收窄,已写进 `BLOCKED.md`。顺带修掉一个**静默失效的旧断言**(`slice(indexOf(A), indexOf(B))` 在 B 改名后变成几乎整份文件,四条 `doesNotMatch` 假通过)|
+391
View File
@@ -0,0 +1,391 @@
#!/usr/bin/env python3
"""Offline research for TASK-rectification-typed-event-scoring-research-20260929.
Question: should a dated typed event ("I changed jobs in 2019-10") move the
candidate scores the same way an answered existence probe does?
Today an answered dated probe moves every candidate by ±2 through the
Vimshottari/Narayana boundary-month split (`event_probes._evaluate_contexts`),
while a typed event only reaches the engine prior, and the frontend skips
`classified_from === "evidence"` yes answers (`core/build-state.ts`).
This script replays the v4 open holdout (20 public Rodden-AA cases, 7–8 events
each; 59 day-precision and 84 year-precision events, no month-precision) with
the same six-probe replay used by the 09-14 / 09-16 / 09-26 studies, and adds
research arms:
* ``B0`` baseline: prior, then the first six public probes answered from truth.
* ``R1d`` each training typed event with day precision is scored as an answered
``yes`` probe (split from `_evaluate_contexts` at that year/month),
then the same six probes.
* ``R1y`` as R1d, plus year-precision events scored with the year-level split
(`_evaluate_contexts(month=None)`, the existing year probe path).
* ``R2k`` R1 with k ∈ {1, 2} typed day events moved by ±1–3 months (seeded);
reports how often the truth is pushed out of the delivered range.
* ``R3`` `_existence_blocked_years` narrowed to "same domain, same year"
(probe supply count), alone and combined with R1d.
Two prior scales are reported: ``raw`` (engine row score, the research-lib
convention of the earlier studies) and ``percent`` (proportional percent over
public representatives, the production `relative_support` convention).
Nothing here changes a production default. Module attributes are patched for
one call and restored; the end of `main` asserts they are the originals.
This is an open-set replay, not a blind test: numbers are not accuracy.
"""
from __future__ import annotations
import argparse
import json
import random
import statistics
import sys
from contextlib import contextmanager
from datetime import date
from pathlib import Path
from typing import Any, Iterator, Sequence
ROOT = Path(__file__).resolve().parents[2]
if str(ROOT) not in sys.path:
sys.path.insert(0, str(ROOT))
import scripts.rectification.event_probes as event_probes # noqa: E402
from scripts.active_rectification_event_engine import compute_candidate_static_contexts # noqa: E402
from scripts.rectification.candidate_contrast import cluster_contexts_by_signature # noqa: E402
from scripts.rectification.case_holdout import holdout_event_ids # noqa: E402
from scripts.rectification.refinement_packet import window_scan # noqa: E402
from scripts.rectification.scoring_service import ( # noqa: E402
build_event_contribution_matrix,
score_from_matrix,
)
from scripts.research.cluster_width_lib import ( # noqa: E402
SEPARATION_LEAD,
delivery_from_public,
merge_adjacent_traced,
public_from_clusters,
raw_signature_clusters,
still_valid_public,
)
from scripts.research.guided_collect_holdout_replay import ( # noqa: E402
TODAY,
_hhmm,
_minutes,
load_cases,
precision_gate,
)
from scripts.research.minute_resolution_sweep import scoring_request_for # noqa: E402
from scripts.research.probe_supply_after_six import ( # noqa: E402
ASK_COUNT,
apply_answer,
optimal_answer,
top1_hit,
)
RADII = (10, 30, 60)
PRIORS = ("raw", "percent")
REPORT_JSON = ROOT / "docs" / "research" / "rectification_typed_event_scoring_2026_09_29.json"
SHIFT_SEED = 20260929
SHIFT_REPEATS = 10
TYPED_SOURCE = "typed_event_research"
ORIGINAL_BLOCKED_YEARS = event_probes._existence_blocked_years
@contextmanager
def narrowed_blocking() -> Iterator[None]:
"""R3: an already-known year blocks only itself in the same domain."""
event_probes._existence_blocked_years = lambda _domain, known_years: set(known_years)
try:
yield
finally:
event_probes._existence_blocked_years = ORIGINAL_BLOCKED_YEARS
def _percent(scores: dict[str, float]) -> dict[str, float]:
total = sum(max(value, 0.0) for value in scores.values())
if total <= 0:
return {key: 0.0 for key in scores}
# Integer percent like `_relative_support_proportional` (largest remainder is not
# needed for ranking; plain rounding keeps the research copy independent).
return {key: float(round(max(value, 0.0) / total * 100)) for key, value in scores.items()}
def _event_month(event: dict[str, Any]) -> int | None:
raw = str(event.get("date_start") or "")
return int(raw[5:7]) if len(raw) >= 7 and raw[4] == "-" else None
def _event_year(event: dict[str, Any]) -> int | None:
raw = str(event.get("date_start") or "")
return int(raw[:4]) if raw[:4].isdigit() else None
def _shift_month(event: dict[str, Any], months: int) -> dict[str, Any]:
year, month = _event_year(event), _event_month(event)
if year is None or month is None:
return event
index = year * 12 + (month - 1) + months
new_year, new_month = divmod(index, 12)
stamp = f"{new_year:04d}-{new_month + 1:02d}-15"
return {**event, "date_start": stamp, "date_end": stamp}
def typed_event_probes(
request: dict[str, Any],
contexts: Sequence[dict[str, Any]],
*,
include_year: bool,
events: Sequence[dict[str, Any]] | None = None,
) -> tuple[list[dict[str, Any]], dict[str, int]]:
"""Score each training typed event through the probe split, answered ``yes``."""
full = [item for item in contexts if event_probes._context_time(item)]
full.sort(key=lambda item: event_probes._clock(str(event_probes._context_time(item))))
clusters = cluster_contexts_by_signature(full)
reps = [cluster["representative"] for cluster in clusters if event_probes._scoreable(cluster["representative"])]
if len(reps) < 2:
reps = [item for item in full if event_probes._scoreable(item)]
set_version = event_probes.candidate_set_version([cluster["times"] for cluster in clusters])
source_events = list(events if events is not None else request["events"])
holdout = set(holdout_event_ids(request["events"]))
stats = {"training_day": 0, "training_year": 0, "split_day": 0, "split_year": 0}
rows: list[dict[str, Any]] = []
for event in source_events:
if str(event.get("id")) in holdout:
continue
domain = event_probes.canonical_domain(event.get("domain"))
if domain not in event_probes.DOMAIN_CATALOG:
continue
precision = str(event.get("precision") or "")
year = _event_year(event)
if year is None:
continue
if precision in {"day", "month"}:
stats["training_day"] += 1
month = _event_month(event)
key = "split_day"
elif precision == "year" and include_year:
stats["training_year"] += 1
month = None
key = "split_year"
else:
continue
row = event_probes._evaluate_contexts(
reps,
birth_date=str(request["birth_date"]),
domain=domain,
year=year,
month=month,
source=TYPED_SOURCE,
clusters=clusters,
set_version=set_version,
)
if row is None:
continue
stats[key] += 1
rows.append(row)
return rows, stats
def replay(
*,
rows: Sequence[dict[str, Any]],
contexts: Sequence[dict[str, Any]],
typed: Sequence[dict[str, Any]],
probes: Sequence[dict[str, Any]],
true_time: str,
prior_mode: str,
) -> dict[str, Any]:
raw = raw_signature_clusters(contexts)
by_time = {stamp: row for row in rows if (stamp := _hhmm(row.get("time")))}
merged, _trace = merge_adjacent_traced(raw, by_time)
public = public_from_clusters(merged, rows)
reps = [str(row["time"])[:5] for row in public]
base = {stamp: float(row.get("score") or 0) for row in public if (stamp := _hhmm(row.get("time")))}
scores = _percent(base) if prior_mode == "percent" else dict(base)
conflicts = {time: 0 for time in reps}
eliminated: set[str] = set()
for probe in typed:
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, "yes", reps)
asked = 0
for probe in list(probes)[:ASK_COUNT]:
answer = optimal_answer(probe, true_time)
if answer is None:
continue
asked += 1
scores, conflicts, eliminated = apply_answer(scores, conflicts, eliminated, probe, answer, reps)
posterior = [{**row, "score": scores.get(str(row["time"])[:5], row.get("score") or 0)} for row in public]
valid = still_valid_public(posterior, scores, eliminated, lead=SEPARATION_LEAD)
delivery = delivery_from_public(valid)
start, end = delivery.get("start"), delivery.get("end")
inside = start is not None and end is not None and _minutes(start) <= _minutes(true_time) <= _minutes(end)
active = [time for time in reps if time not in eliminated]
gate = precision_gate(valid, scores)
return {
"width": delivery.get("width"),
"truth_in_range": bool(inside),
"top1": bool(top1_hit(scores, active, true_time, merged)),
"tied": bool(gate["tied_for_first"]),
"questions": asked,
"typed_scored": len(typed),
"probe_supply": len(probes),
}
def evaluate_case(case: dict[str, Any], radius: int) -> dict[str, Any]:
true_time = str(case["birth"]["time"])[:5]
request = scoring_request_for(case, radius)
contexts = compute_candidate_static_contexts(request)
built = build_event_contribution_matrix(request, static_contexts=contexts)
rows = score_from_matrix(request, built)
times = [stamp for row in rows if (stamp := _hhmm(row.get("time")))]
def probes_now() -> list[dict[str, Any]]:
return event_probes.discriminating_event_probes(
{**request, "refresh_probes": False, "asked_probe_keys": []}, built,
scan=window_scan(built), candidate_times=times, representative_time=true_time, today=TODAY,
)
probes = probes_now()
with narrowed_blocking():
probes_narrow = probes_now()
assert event_probes._existence_blocked_years is ORIGINAL_BLOCKED_YEARS
typed_day, stats = typed_event_probes(request, contexts, include_year=False)
typed_all, stats_all = typed_event_probes(request, contexts, include_year=True)
rng = random.Random(f"{SHIFT_SEED}:{case.get('case_id')}:{radius}")
holdout = set(holdout_event_ids(request["events"]))
day_events = [
index for index, event in enumerate(request["events"])
if str(event.get("precision")) in {"day", "month"} and str(event.get("id")) not in holdout
]
shifted: dict[str, list[dict[str, Any]]] = {"1": [], "2": []}
for k in (1, 2):
if len(day_events) < k:
continue
for _ in range(SHIFT_REPEATS):
picks = rng.sample(day_events, k)
events = [
_shift_month(event, rng.choice((-3, -2, -1, 1, 2, 3))) if index in picks else event
for index, event in enumerate(request["events"])
]
moved, _ = typed_event_probes(request, contexts, include_year=False, events=events)
shifted[str(k)].append({"typed": moved})
out: dict[str, Any] = {"case_id": case.get("case_id"), "radius": radius, "typed_stats": {**stats, **{
"training_year": stats_all["training_year"], "split_year": stats_all["split_year"]}},
"probe_supply": len(probes), "probe_supply_narrow": len(probes_narrow),
"asked_keys_changed_narrow": [p["semantic_key"] for p in probes[:ASK_COUNT]]
!= [p["semantic_key"] for p in probes_narrow[:ASK_COUNT]],
"arms": {}}
for prior in PRIORS:
def run(typed: Sequence[dict[str, Any]], pool: Sequence[dict[str, Any]]) -> dict[str, Any]:
return replay(rows=rows, contexts=contexts, typed=typed, probes=pool, true_time=true_time, prior_mode=prior)
arms = {
"B0": run([], probes),
"R1d": run(typed_day, probes),
"R1y": run(typed_all, probes),
"R3": run([], probes_narrow),
"R3+R1d": run(typed_day, probes_narrow),
}
for k, trials in shifted.items():
results = [run(item["typed"], probes) for item in trials]
if results:
arms[f"R2k{k}"] = {
"trials": len(results),
"truth_out": sum(1 for row in results if not row["truth_in_range"]),
"top1": sum(1 for row in results if row["top1"]),
"median_width": statistics.median(row["width"] for row in results if row["width"] is not None),
}
out["arms"][prior] = arms
return out
def summarize(rows: Sequence[dict[str, Any]], prior: str, radius: int, arm: str) -> dict[str, Any]:
subset = [row["arms"][prior][arm] for row in rows if row.get("radius") == radius and not row.get("error")
and arm in row["arms"][prior]]
n = len(subset)
if not n:
return {"n": 0}
if arm.startswith("R2k"):
trials = sum(item["trials"] for item in subset)
return {
"n": n,
"trials": trials,
"truth_out_rate": round(sum(item["truth_out"] for item in subset) / trials, 4),
"cases_with_truth_out": sum(1 for item in subset if item["truth_out"] > 0),
"top1_rate": round(sum(item["top1"] for item in subset) / trials, 4),
"median_width": statistics.median(item["median_width"] for item in subset),
}
return {
"n": n,
"top1": round(sum(item["top1"] for item in subset) / n, 4),
"truth_in_range": round(sum(item["truth_in_range"] for item in subset) / n, 4),
"median_width": statistics.median(item["width"] for item in subset if item["width"] is not None),
"tie_rate": round(sum(item["tied"] for item in subset) / n, 4),
"mean_questions": round(statistics.mean(item["questions"] for item in subset), 2),
"mean_typed_scored": round(statistics.mean(item["typed_scored"] for item in subset), 2),
"mean_probe_supply": round(statistics.mean(item["probe_supply"] for item in subset), 2),
}
def verdict(summary: dict[str, Any], prior: str, arm: str) -> str:
"""Gate: top1 not lower AND truth-in-range not lower AND median width lower, all radii."""
passes = []
for radius in RADII:
base = summary[prior][str(radius)]["B0"]
cand = summary[prior][str(radius)][arm]
if not cand.get("n"):
return "uncertain"
passes.append(
cand["top1"] >= base["top1"]
and cand["truth_in_range"] >= base["truth_in_range"]
and cand["median_width"] < base["median_width"]
)
return "benefit" if all(passes) else "no_benefit"
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--limit", type=int, default=0)
parser.add_argument("--json-out", default=str(REPORT_JSON))
args = parser.parse_args()
cases = load_cases()
if args.limit:
cases = cases[: args.limit]
rows: list[dict[str, Any]] = []
for case in cases:
for radius in RADII:
try:
rows.append(evaluate_case(case, radius))
except Exception as exc: # noqa: BLE001
rows.append({"case_id": case.get("case_id"), "radius": radius, "error": f"{type(exc).__name__}: {exc}"})
arms = ("B0", "R1d", "R1y", "R3", "R3+R1d", "R2k1", "R2k2")
summary = {prior: {str(radius): {arm: summarize(rows, prior, radius, arm) for arm in arms}
for radius in RADII} for prior in PRIORS}
verdicts = {prior: {arm: verdict(summary, prior, arm) for arm in ("R1d", "R1y", "R3", "R3+R1d")} for prior in PRIORS}
assert event_probes._existence_blocked_years is ORIGINAL_BLOCKED_YEARS
payload = {
"generated_for": "TASK-rectification-typed-event-scoring-research-20260929",
"today": TODAY.isoformat(),
"holdout": "references/real_case_calibration/minute_rectification_holdout_v4.json",
"ask_count": ASK_COUNT,
"separation_lead": SEPARATION_LEAD,
"shift_seed": SHIFT_SEED,
"shift_repeats": SHIFT_REPEATS,
"open_set_not_blind": True,
"summary": summary,
"verdicts": verdicts,
"rows": rows,
"errors": [row for row in rows if row.get("error")],
}
out = Path(args.json_out)
out.write_text(json.dumps(payload, ensure_ascii=False, indent=1, sort_keys=True) + "\n", encoding="utf-8")
print(json.dumps({"summary": summary, "verdicts": verdicts}, ensure_ascii=False, indent=1))
return 0
if __name__ == "__main__":
raise SystemExit(main())