Files
Jyotisha/docs/research/jev_intent_2026_09_19.md
T
jesse-ux 507ef959f5
Independent Staging Quality Gate / validate (push) Failing after 9m3s
Independent Staging Quality Gate / publish (push) Skipped
research(jev-intent): fix2 把来源 B 现行与高置信错误补进报告
离线从 cache 聚合,不重跑模型。无焦点层 Flash 69.7% 低于 Jev 78.8%。采集层相对门槛标不可判。
2026-09-19 12:28:43 +08:00

211 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TypeSafe Jev 接管校正意图分类 · 离线对照(2026-09-19)
- 任务:`docs/tasks/TASK-rectification-jev-intent-classifier-research-20260919.md`
- 基线:`origin/staging` @ `69a44fe7`
- Jev 模型:`jev-1.13.0`(不用 jev-latest)
- 真值 sha256:synthetic `79d53a958b3b42ed18c79e3bbcff4b31deb22b65eedcd89b39326a7fefc9b0e8`;simulated `53f13012fdc8f3721335bbe2961edf8c506a2f705b62bd9d302fce83ce2ef88d`;disputed `e80c5cd69c521a099d743e4a4632ad6bc103c3b5a505f5b376a7c9cebedb7646`
- 来源 B:157 条已标注(本地,不提交原文)
## 模型
- 生成模型:`deepseek-flash` / 版本 `deepseek-flash` / 不是线上会话模型(只用于造来源 C)
- 复核模型:`deepseek-flash` / 版本 `deepseek-flash` / 不是线上会话模型(只用于独立复核,看不到目标标签)
- 对照模型:`deepseek-flash` / 版本 `deepseek-flash` / **= 线上会话模型**(DeepSeek Flash,顶生产 `classifyRectificationTurnIntent` 提示词)
- Jev:`jev-1.13.0` / 不是线上会话模型
## 结论
**缺数据**。来源 B 与来源 C 同层 intent 准确率差 > 10pp,模拟语料不代表真人,来源 C 门槛结论降为缺数据。来源 B 与来源 C 同层 intent 准确率:choice B 94.1% vs C 99.0%(差 4.9%,n_B=17);collect B 94.4% vs C 92.2%(差 2.1%,n_B=107);none B 78.8% vs C 96.0%(差 17.2%,n_B=33)。 有层差 > 10pp,结论降为缺数据。 同时来源 C 绝对门槛未过:低置信召回 23.1% / 45.5% / 25.0% < 60%(错了却仍高置信)。 相对 −3pp(同一样本):choice Jev 99.0% vs 现行 98.8%(门槛 95.8%,过);collect Jev 92.2% vs 现行 97.5%(门槛 94.5%,不可判(复核与对照同源));none Jev 96.0% vs 现行 94.0%(门槛 91.0%,过)。
若接:不得上线。先补真机样本或重造更像真人的来源 C,再测。
## T3 指标(Jev,来源 C)
| 层 | n | intent | answer_class | dated | 高置信错误 | 低置信覆盖 | 低置信召回 | no/unsure 互判 | 自洽率 | 中位 ms | P95 ms | 次均 input tok |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| choice | 400 | 99.0% | 95.8% | 100.0% | 0.0% | 2.2% | 23.1% | 0 | 99.8% | 908 | 1194 | 1127 |
| collect | 400 | 92.2% | 95.7% | 98.8% | 2.8% | 8.0% | 45.5% | 6 | 98.0% | 909 | 1086 | 991 |
| none | 100 | 96.0% | — | 100.0% | 1.0% | 2.0% | 25.0% | 0 | 100.0% | 909 | 981 | 922 |
来源 A(测试夹具,n=10)intent 90.0%,answer_class 85.7%,高置信错误 0.0%。
## 现行模型对照
模型:`deepseek-flash`。DeepSeek Flash = 线上会话模型,顶现行 `classifyRectificationTurnIntent` 提示词;来源 C 全量第一次,第二次分层 33%(seed 20260920);来源 A 全量;来源 B 已标注全量一次。
| 层 | n | intent | answer_class | dated | 自洽率 | 中位 ms |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| choice | 400 | 98.8% | 99.5% | 100.0% | 100.0% | 685 |
| collect | 400 | 97.5% | 99.6% | 99.0% | 99.2% | 686 |
| none | 100 | 94.0% | — | 100.0% | 100.0% | 672 |
同一样本上 Jev vs 现行(相对门槛用这一表):
| 层 | n | Jev intent | 现行 intent | 差(Jev−现行) | 门槛现行−3pp | 判定 |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| choice | 400 | 99.0% | 98.8% | 0.2% | 95.8% | 过 |
| collect | 400 | 92.2% | 97.5% | -5.2% | 94.5% | 不可判(复核与对照同源) |
| none | 100 | 96.0% | 94.0% | 2.0% | 91.0% | 过 |
## 代表性检验(来源 B vs 来源 C)
来源 B 与来源 C 同层 intent 准确率:choice B 94.1% vs C 99.0%(差 4.9%,n_B=17);collect B 94.4% vs C 92.2%(差 2.1%,n_B=107);none B 78.8% vs C 96.0%(差 17.2%,n_B=33)。 有层差 > 10pp,结论降为缺数据。
来源 B 是真人 + 人工标注 + 线上模型三者齐备的唯一一组。现行无置信度,置信度三列为空。
| 范围 | n | Jev intent | 现行 intent | Jev 高置信错误 | Jev 低置信召回 | Jev 自洽 |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| 全集 | 157 | 91.1% | 89.2% | 6.4% | 24.1% | 96.2% |
| choice | 17 | 94.1% | 88.2% | 0.0% | 66.7% | 94.1% |
| collect | 107 | 94.4% | 95.3% | 5.6% | 18.2% | 99.1% |
| none | 33 | 78.8% | 69.7% | 12.1% | 20.0% | 87.9% |
无焦点层现行 intent 69.7%、Jev 78.8%。现行更低,说明 78.8% 主要是这 33 条本身难,不是单 Jev 不行。
无焦点层 gold × 预测混淆计数(只有计数,无原文)。预测出现 `answer_current_focus` 是因为模型把无焦点句当成在回答采集题。
### gold × Jev(n=33)
| gold \ pred | provide_new_evidence | stop_rectification | ask_about_result | unclear | answer_current_focus |
| --- | ---: | ---: | ---: | ---: | ---: |
| provide_new_evidence | 24 | 0 | 0 | 0 | 0 |
| stop_rectification | 0 | 0 | 0 | 0 | 0 |
| ask_about_result | 0 | 0 | 1 | 0 | 0 |
| unclear | 0 | 0 | 0 | 1 | 7 |
| answer_current_focus | 0 | 0 | 0 | 0 | 0 |
### gold × 现行(n=33)
| gold \ pred | provide_new_evidence | stop_rectification | ask_about_result | unclear | answer_current_focus |
| --- | ---: | ---: | ---: | ---: | ---: |
| provide_new_evidence | 21 | 0 | 0 | 0 | 3 |
| stop_rectification | 0 | 0 | 0 | 0 | 0 |
| ask_about_result | 0 | 0 | 0 | 1 | 0 |
| unclear | 1 | 0 | 0 | 2 | 5 |
| answer_current_focus | 0 | 0 | 0 | 0 | 0 |
## 置信度–准确率曲线与 θ
推荐 θ = 0.9。点:
```json
[
{
"theta": 0.3,
"jev_share": 0.9911111111111112,
"fallback_share": 0.008888888888888889,
"jev_acc": 0.9405829596412556,
"high_conf_errors_in_used": 12
},
{
"theta": 0.4,
"jev_share": 0.9744444444444444,
"fallback_share": 0.025555555555555557,
"jev_acc": 0.9498289623717218,
"high_conf_errors_in_used": 12
},
{
"theta": 0.5,
"jev_share": 0.9522222222222222,
"fallback_share": 0.04777777777777778,
"jev_acc": 0.9568261376896149,
"high_conf_errors_in_used": 12
},
{
"theta": 0.6,
"jev_share": 0.9222222222222223,
"fallback_share": 0.07777777777777778,
"jev_acc": 0.9698795180722891,
"high_conf_errors_in_used": 12
},
{
"theta": 0.7,
"jev_share": 0.8977777777777778,
"fallback_share": 0.10222222222222223,
"jev_acc": 0.9752475247524752,
"high_conf_errors_in_used": 12
},
{
"theta": 0.8,
"jev_share": 0.8222222222222222,
"fallback_share": 0.17777777777777778,
"jev_acc": 0.9837837837837838,
"high_conf_errors_in_used": 12
},
{
"theta": 0.9,
"jev_share": 0.7033333333333334,
"fallback_share": 0.2966666666666667,
"jev_acc": 0.9921011058451816,
"high_conf_errors_in_used": 5
}
]
```
## 错例画像(改写后,无法对应真实会话)
### cjk_colloquial
- `C-choice-0149` / choice / 「嗯,那会儿好像有点变动,但印象不深。另外2019年搬过一次家。」 gold={'answer_class': 'weak_yes', 'has_new_dated_event': True, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.65}
- `C-choice-0152` / choice / 「好像有吧但印象不深,对了2019年搬过家。」 gold={'answer_class': 'weak_yes', 'has_new_dated_event': True, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.76}
- `C-collect-0286` / collect / 「差不多是2021年夏天那会儿,七月份前后吧,确实有明显变化。」 gold={'answer_class': 'yes', 'has_new_dated_event': False, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'weak_yes', 'has_new_dated_event': False, 'confidence': 0.71}
### literal
- `A10` / collect / 「2016年3月入学」 gold={'answer_class': 'yes', 'has_new_dated_event': False, 'intent': 'answer_current_focus'} jev={'intent': 'provide_new_evidence', 'answer_class': None, 'has_new_dated_event': True, 'confidence': 0.42}
- `C-choice-0386` / choice / 「这个真不好说,另外2019年我换过工作搬了家。」 gold={'answer_class': None, 'has_new_dated_event': True, 'intent': 'provide_new_evidence'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.64}
- `C-choice-0341` / choice / 「对了,1995年那会儿我搬家了,你问的这个我真说不好。」 gold={'answer_class': None, 'has_new_dated_event': True, 'intent': 'provide_new_evidence'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.54}
### other
- `C-choice-0148` / choice / 「好像有点小变化,但记不太深。另外2019年换过工作。」 gold={'answer_class': 'weak_yes', 'has_new_dated_event': True, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.73}
- `C-choice-0154` / choice / 「好像有那么点,但印象不深。对了,2002年换过一次工作。」 gold={'answer_class': 'weak_yes', 'has_new_dated_event': True, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.59}
- `C-choice-0162` / choice / 「好像有点印象,但很淡。另外2019年搬过一次家。」 gold={'answer_class': 'weak_yes', 'has_new_dated_event': True, 'intent': 'answer_current_focus'} jev={'intent': 'answer_current_focus', 'answer_class': 'unsure', 'has_new_dated_event': True, 'confidence': 0.61}
## T1 判据对照
见 `scripts/research/jev_intent_questions.py` 的 `CRITERION_MAP`。丢掉的语义:生产提示里的「通常」「不要按 A/B/C/D 猜」「不要按关键词表」;结构不变量改由 `enforce_combo` 强制。
| 现行提示原句 | Jev 落点 | 丢失 |
| --- | --- | --- |
| 结合当前问题和动态选项判断用户是在回答当前问题、提供新的带时间经历、要求停止整个校正、询问结果,还是语义不清。 | intent Choice with five literal criteria, including unclear as the residual option. | — |
| 若是在回答当前问题,answer_class 必须使用某个选项提供的 answer_class;否则 answer_class 必须为 null。 | answer_class Choice is built only from the live options; code nulls it when intent != answer_current_focus. | — |
| has_new_dated_event 仅在用户同一句里除了回答当前问题之外,还提供了新的、带大概时间的经历时为 true。 | has_new_dated_event Noul true/false criteria, same split. | — |
| 单纯的否定或单纯的选项回答必须为 false。 | Noul false: 单纯否定、单纯记不清不算新经历。 | — |
| 若同一句话既回答了当前问题又补充了新的带时间经历,intent 仍为 answer_current_focus,has_new_dated_event 为 true。 | intent and Noul are separate questions; code keeps answer_current_focus when both fire. | Jev does not itself enforce this pair; code does. |
| “当前方面没有、那段时间没有变化”通常是回答当前问题,不是停止整个流程。 | stop_rectification criterion: only an explicit stop of the whole flow. 当前方面没有 is answer_current_focus. | The hedge 通常 is dropped; the stop option is written as an exclusive literal. |
| 只有用户明确要求停止整个校正时才分类为 stop_rectification。 | stop_rectification criterion, same literal. | — |
| 不要按 A/B/C/D 的位置猜语义,只按选项 label 与 answer_class 判断。 | answer_class criteria keyed by answer_class, not by A/B/C/D. | The prohibition itself is not sent; Jev is not given A/B/C/D as the choice keys. |
| 「没有、没发生过、这方面没什么」→ intent 为 answer_current_focus,answer_class 为 no。 | Collect answer_class criterion for no. | — |
| 「记不清、不记得、忘了、想不起来、以后再说」→ answer_class 为 unsure。 | Collect answer_class criterion for unsure. | — |
| 用户用带大概年月的经历直接回答当前采集题 → intent 为 answer_current_focus,answer_class 为 yes(程度较弱时为 weak_yes),has_new_dated_event 为 false。 | Collect yes/weak_yes criteria + Noul false for a dated answer that is the collect answer. | — |
| 不要把「没有」或「记不清」标成 yes。 | Separate no and unsure criteria; yes requires a dated or affirmative answer. | The 'do not' wording is dropped. |
| 若既没有否定、也没有说记不清、也没有给出带年月经历,intent 为 unclear,answer_class 必须为 null。 | unclear residual on intent Choice; code nulls answer_class. | — |
| 若用户只在补充带时间的经历、并没有回答当前采集题,intent 为 provide_new_evidence。 | provide_new_evidence criterion. | — |
| 若同一句话既明确否定当前采集题又补充了新的带时间经历,intent 仍为 answer_current_focus 且 answer_class 为 no,不要改成 provide_new_evidence。 | Code keeps no + Noul true. Jev intent/Noul are independent. | Jev may split this pair; code does not re-vote intent from the Noul. |
| 不要按关键词表或正则猜测,只根据当前问题与用户这句话的语义分类。 | Not sent. Jev has no keyword table in the question. | The meta-instruction is dropped (Jev answers the written question, not the intended one). |
## 限制
1. **复核 ≈ 生产提示,且复核模型 = 对照模型。** `REVIEW_RUBRIC` 与生产 `COLLECT_INSTRUCTIONS` 逐句对应;生成 / 复核 / 对照都是 `deepseek-flash`。进入测试集的 900 条是「Flash 用近生产提示能答对目标标签」的那 900 条,被剔的 36 条恰是 Flash 不同意的。因此来源 C 上现行 97.5% / 98.8% / 94.0% 是构造出来的上界。采集层「Jev 92.2% 未过相对门槛 94.5%」**不可当作 Jev 输给现行的证据**。
2. **采集层 intent 错例的 gold 有争议。** 来源 C 采集层 Jev intent 错例 31 条:`provide_new_evidence → answer_current_focus` 17、`unclear → answer_current_focus` 9、`stop → unclear` 4、`ask → unclear` 1。17 条 pne 几乎全是「另外 2019 年我换工作搬了家」句式,若干尾句落在当前题域,按生产提示可读成 `answer_current_focus + unsure`。9 条 unclear(「一时半会儿真捋不明白」)按「记不清 → unsure」也读得通。两类合计 ≥ 20 条,占该层 intent 错例约 2/3。gold 来自「生成目标 + Flash 复核同意」,不等于人工真值。改写示例:
- 「另外 2019 年换过工作,感情那会儿真没细想。」gold=provide_new_evidence;可读成在回答感情采集题。
- 「另外 2019 年搬了家,工作那摊子反而没顾上细想。」gold=provide_new_evidence;可读成在回答工作采集题。
- 「这事我一时真说不上来。」gold=unclear;按生产提示是 unsure。
3. **语料仍不像真人。** 修复单 1 只把长度和人设写成硬红线(已过)。原单还要求按来源 B 的标点 / 语气词比例约束,未进红线:
| 指标 | 来源 B(真人) | 来源 C(模拟) |
| --- | ---: | ---: |
| 含标点 | 13% | 98.9% |
| 含语气词(吧/呢/啊/嗯/哦/额/emm) | 4% | 25.6% |
| 含年份或月份 | 81% | 47.8% |
| 带年份句里写「2019」 | — | 236 / 429 = 55% |
| `provide_new_evidence` 里是搬家/换工作 | — | 152 / 157 |
根因:生成脚本 `ALT_EVENT_HINTS` 给七个领域的「另一件事」全是搬家/换工作,年份未约束,模型收敛到「另外 2019 年搬过家」。这解释了模拟语料不代表真人的一部分,也解释了采集层错例为何长得一样。本单不修,留给产品决定是否再造一轮。
## 回退
任何上线方案必须保留回退到现行会话模型的路径。官方限流会动态调整。