research(rectification): offline R1/R2/R3 — answer-flip tolerance, V1/V2 rerun, dasha shift arithmetic
- R1: flipping 1 answer keeps truth in range 98-100% but cuts head hit by a third or more; 2 flips squeeze truth out in 7-10% of ±30/±60 replays (two flips = 8 points = SEPARATION_LEAD). - R2: weights do apply (research scorer == production at V0); V1/V2 are identity at ±30/±60 by construction and leave six-question metrics unchanged at ±10 -> no_benefit (measured). Supplementary V1n does not pass the gate. - R3: boundary shift is ~3.8 days/minute (1.3-5.9), not 1.1; the 45-day gate is ~8-34 minutes. The _representative_pairs hypothesis is refuted (all-pairs adds no dated probes); the bottleneck is monthly evaluation. New finding recorded as BUG-1048 (investigating): _boundary_windows year-straddle exemption and positional zip misalignment bypass the gate. - Dated errata appended (no deletions) to the 09-14/09-16 briefs and research docs; README board row -> 待验收. No production code, scoring, thresholds, gates or Skill changed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
7203d94e9c
commit
0f5442cea2
@@ -10789,6 +10789,7 @@
|
||||
- 修复:邀请语按候选数出话(2 → 「这两分钟」;>2 → 「这几个候选」;未知 → 不说这半句)。`exitAppendCalls()` 加入「收到 … 这 N 分钟 / 按现有信息分不开 / 说出来我接着算」,保持 `assert.equal(gates.length, 1)`,不加「只微调排序」。公开字段深比较补 `birth_time_source`。M1b 的 V1/V2 从 `no_benefit` 勘误为 `not_measured`。
|
||||
- 验证:`frontend/tests/rectification-open-collect-invite-20260914.test.ts`、`rectification-exhaustion-exit-20260906.test.ts`、`rectification-decision-authority.test.ts`。
|
||||
- 防复发:文案里出现的数字必须来自同一份数据,不得硬编码;改用户可见文案或公开字段时,必须同步跑 `npm test` 并把失败数与基线比对。
|
||||
- 补充(2026-09-26 离线研究 R2):M1b 的 V1/V2 已重跑。权重确实生效(研究计分器 V0 与线上逐分相同);V1/V2 在 ±30 / ±60 上按定义等于等权(0 个候选分数变化),在 ±10 上只微调分数、六题后指标不变。判定由 `not_measured` 改为 `no_benefit`(已实测)。见 `docs/research/rectification_offline_research_2026_09_26.md` §R2。
|
||||
- 相关记录:BUG-689、BUG-690、BUG-688、BUG-687、BUG-587
|
||||
- 复发自:BUG-689(邀请语复用两候选句)、BUG-690(公开字段未进深比较)
|
||||
- 修复版本:待发布
|
||||
@@ -14036,3 +14037,20 @@
|
||||
- 相关记录:BUG-721、BUG-722、BUG-723、BUG-724、BUG-725、BUG-726、BUG-388、BUG-1046。
|
||||
- 复发自:无。
|
||||
- 修复版本:`ab0f01a9`,已在 staging,deploy-staging run 2930(`8a409434`)部署(2026-09-26 对账)。
|
||||
|
||||
## BUG-1048 | 出题闸门 `_boundary_windows` 两处绕过:跨 1 月 1 日的边界不卡最小间隔,列表错位时比的是不对应的边界
|
||||
|
||||
- 状态:investigating(离线研究确认了代码行为,未确认对真人作答的影响;未改代码)
|
||||
- 首次发现 / 最近更新:2026-09-26 / 2026-09-26
|
||||
- 影响面:`scripts/rectification/event_probes.py` `_boundary_windows`,经 `_union_boundary_dates` 影响所有带日期的自动点选题(`source = dasha_boundary`),初始与刷新两阶段;Vimshottari 与 Narayana 两条轨道都走这个函数。`event_probes.py` 属于冻结评分文件(ERR-110)。
|
||||
- 用户现象:无直接用户报告。离线研究(公开 AA 开放集 v4 与虚构时刻)发现,文档所说的「候选间大运边界差 < 45 天(刷新 30 天)不出题」并没有对所有配对生效。
|
||||
- 触发条件:(1)两个候选的同一条边界分别落在 12 月底和次年 1 月初:代码条件是 `one.year == two.year and 差 < 阈值` 才跳过,跨年的配对即使只差几天也放行。(2)两个候选的边界列表按位置 `zip`:某条边界跨出 `[出生+5, 出生+80]` 的年份边缘时,两边列表长度差一到两个元素(MD 起点与其第一个 AD 起点是同一天)。此后每一对比较的都是不对应的边界,间隔很大,于是全部放行。
|
||||
- 根因:闸门只实现了同年分支;按位置配对假设两边列表逐项对应,而 Vimshottari 的真实关系是整体平移同一天数。
|
||||
- 已确认事实:20 分钟虚构配对(MD+AD)有 10.7% 错位;v4 刷新阶段(含 PD)线上配对有 35–44% 错位。研究补丁把闸门改为对所有配对生效、错位时对齐后:±10 初始阶段没有带日期题的例子从 4 例变成 8 例,刷新阶段从 3 例变成 4 例,带日期题均值 4.85 → 4.10;±30 / ±60 基本不变。窄窗上现有的一部分带日期题依赖这两个漏洞。
|
||||
- 未确认:这些题问的是「几个候选的边界只差几天」的那个月份,按月精度回答能否可靠区分,没有评估;修掉漏洞会让窄窗更早出现「没有带日期题」,是否可以接受需要产品决定。
|
||||
- 修复:未修。任何修改都要按 ERR-110 重新冻结 sealed holdout / reported offset 记录。
|
||||
- 验证:`python3 scripts/research/representative_pairs_probe.py`(`strict_gate` 对照)、`python3 scripts/research/dasha_shift_per_minute.py`(`repo_raw_zip_misaligned_share`)。
|
||||
- 防复发:待定。改闸门时,同时用同年、跨年、列表错位三类虚构配对做回归。
|
||||
- 相关记录:BUG-689、BUG-740、BUG-1047;ERR-110。
|
||||
- 复发自:无。
|
||||
- 修复版本:未修。
|
||||
|
||||
Reference in New Issue
Block a user