refactor(consult): chat binds the method's truth boundaries, not the engine manual (BUG-1256 step 2)
Independent Staging Quality Gate / validate (push) Successful in 14m35s
Independent Staging Quality Gate / publish (push) Successful in 3m37s

The method block bound into every chat turn quoted most of SKILL.md (benchmark
scores, CLI flags, oracle queues, file paths) plus the shared method, whose
output contract asks for JSON verdicts, A/B/C/D confidence, audit tables, raw
data and web verification - the opposite of the chat shape. Chat now quotes
only the truth-boundary sections and adds CHAT_METHOD_BOUNDARIES for the
limits from the dropped sections it still has to keep. SKILL.md and the shared
method file are unchanged for reports and skill_read.

Natal system prompt 55,242 -> 26,591 characters. DeepSeek A/B on public
golden charts (flash 10 questions): input 24,560 -> 14,025 tokens, 37.8 ->
30.3 s, no regression in automatic checks or reading; a high-rigor request
still says plainly that no external check was done. Full suite 4970 / fail
24, identical to b8385adb; build keeps / Static.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
Jesse_Chen
2026-10-07 08:48:04 +08:00
co-authored by Claude Opus 5.5
parent 6a372bc3a8
commit c76fdda7ab
6 changed files with 142 additions and 50 deletions
@@ -0,0 +1,47 @@
# 模型对比 · 普通对话提示词精简(2026-10-07,BUG-1256)
对比对象:改前 = staging `b8385adb` 的提示词;改后 = 分支 `codex/prompt-slim-20261007`(去重 + 方法书精简)。
做法:用仓库里真实引擎的公开名人 golden 盘(Steve Jobs / Barack Obama / Elizabeth Taylor)构造工具结果,与线上同一套拼法(系统提示 + 用户轮指令 + 工具调用 + 工具结果)直连 DeepSeek API,`max_tokens` 与线上一致(24,576)。两边的问题、盘面、工具结果完全相同,只换系统提示。脚本在会话临时目录,不入库;回答原文不入库,下面只记数字和结论。
## 体量
| 项 | 改前 | 改后 |
| --- | --- | --- |
| 本命系统提示 | 55,242 字符(约 2.2 万 token) | 26,591 字符(约 1.06 万 token) |
| 英文形状摘要出现次数 | 4 | 1 |
| 方法书摘录 + 共享方法 | 25,872 字符 | 约 2,500 字符 |
| 每轮实际输入(API 计量,含工具结果) | 平均 24,560 token | 平均 14,025 token(−43%) |
## deepseek-flash,10 题(事业 / 婚恋 / 财运 / 父母 / 子女 / 健康 / 学业 / 是非题 ×2 / 问时间)
| 指标 | 改前 | 改后 |
| --- | --- | --- |
| 平均耗时 | 37.8 s | 30.3 s(−20%) |
| 平均思考 token | 6,975 | 5,683(−19%) |
| 空回答 | 0(`max_tokens` 改为线上值后) | 0 |
| 自动检查(标题、列表、加粗、清单层名、数宫法、自造词 / 音译、置信标签、断定读心、后台词、空话、语气词) | 全 0 | 全 0 |
| 英文术语命中 | 1 题 | 1 题 |
| 「确定性措辞」命中 | 0 | 1 题,人工看是误报(原句「不是哪一天一定会发生什么」) |
| 是非题首句 | 「看情况」「不是」 | 「看情况」「不是」 |
人工通读 1、3、8 题两版:结构、分段、读心、括号依据一致;改后第 8 题主动说明「另一套推运没对上,所以不给具体某一天」(精简后补的边界要点生效)。
## 高严谨请求(第 11 题:「拉满三大引擎和外部案例验证,精确到月,说明核对过哪些外部资料」)
两版都如实说明未做外部核验。改后措辞更直白(「没有做外部引擎比对,没有联网核验,也没有拿真实案例回放校准……不假装做过」),并以两套推运未同向为由不给确定月份。耗时 42.3 s → 28.7 s。
## deepseek-v4-pro,抽 3 题
| 指标 | 改前 | 改后 |
| --- | --- | --- |
| 平均耗时 | 92.5 s | 33.6 s |
| 平均思考 token | 11,644 | 3,588 |
| 空回答 | 1(第 8 题思考 24,563 token 耗尽上限,181 s) | 1(第 7 题 4 s 返回空,重跑 3/3 正常,判为偶发) |
第 7 题再各跑:改后 3 次、改前 2 次,均正常作答,首句为「不是 / 看情况」。
## 结论与缺口
- 精简后体量减半,速度更快、思考更少,回答质量与诚实边界未见退化。
- 样本小(flash 10 题 + 1 题、pro 3 题),只用 3 张公开名人盘;真机仍需按 `consult-voice-sister-20261007.md` 与 `consult-readable-20261006.md` 两份清单走。