refactor(consult): chat binds the method's truth boundaries, not the engine manual (BUG-1256 step 2)
The method block bound into every chat turn quoted most of SKILL.md (benchmark scores, CLI flags, oracle queues, file paths) plus the shared method, whose output contract asks for JSON verdicts, A/B/C/D confidence, audit tables, raw data and web verification - the opposite of the chat shape. Chat now quotes only the truth-boundary sections and adds CHAT_METHOD_BOUNDARIES for the limits from the dropped sections it still has to keep. SKILL.md and the shared method file are unchanged for reports and skill_read. Natal system prompt 55,242 -> 26,591 characters. DeepSeek A/B on public golden charts (flash 10 questions): input 24,560 -> 14,025 tokens, 37.8 -> 30.3 s, no regression in automatic checks or reading; a high-rigor request still says plainly that no external check was done. Full suite 4970 / fail 24, identical to b8385adb; build keeps / Static. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
6a372bc3a8
commit
c76fdda7ab
@@ -16809,3 +16809,17 @@
|
||||
- 防复发:测试锁人设关键句、读心三条、禁用断定句、受冲只给范围,且 `index.ts` 不再有「concisely」、voice 不再有「restate / 嘴有点毒 / 锋利」。
|
||||
- 相关记录:BUG-1070(被部分推翻)、BUG-1182(受冲口径保留)、BUG-1244(行动不强制)、BUG-1253。
|
||||
- 修复版本:staging `b8385adb`(2026-10-07 08:13 部署核对:health gitCommit=b8385adb、各项 ok,`/api/account` 401,`/login` 200;门禁 run 1750 validate 15 分钟一次通过)。真机未走,状态保持 fixed-pending-verify
|
||||
|
||||
## BUG-1256 | 普通对话系统提示 5.5 万字符:同一形状说 4 遍、方法书带进工程记录与冲突的输出合同
|
||||
|
||||
- 状态:fixed-pending-verify(分支 `codex/prompt-slim-20261007`;部署与真机前不标 resolved)
|
||||
- 首次发现 / 最近更新:2026-10-07 / 2026-10-07
|
||||
- 来源:产品问「提示词规则 2 万多字能不能优化」,Claude 实测体量后产品授权直接执行,并提供临时模型 key 做改前改后对比。
|
||||
- 影响面:本命对话与申报时段对话的系统提示(`mastra/index.ts`、`product-voice.ts` 合同、`consultation-thinking-plan.ts` 的 natal 标题规则、`skill-binding.ts` 的方法块)。不改证据卡、工具、领域清单、用户轮指令、模型与时钟。
|
||||
- 现象:本命系统提示 55,242 字符(约 2.2 万 token),每轮实际输入约 2.46 万 token。英文形状摘要在合同、index、natal 标题规则、skill 骨架里各一份;合同又逐条复述中文 ANSWER SHAPE;方法书摘录 1.74 万字符里多为引擎跑分、CLI 参数、oracle 队列、文件路径;共享方法(全谱调用合同 + 事件判定骨架)要求输出 JSON verdict、A/B/C/D 置信度、审计表、原始数据与联网验证,与聊天形状冲突;一批通用回答政策被放在「When reference_transparency is present」标题下。
|
||||
- 根因:形状与方法块分轮叠加,每轮只加不删;方法块按 SKILL.md 的整节标题摘取,SKILL.md 主要是引擎与报告手册。
|
||||
- 修复:第一步(去重、不改意思):形状摘要只留合同一份,删合同里复述中文形状的条目,通用政策挪到「Always」标题下。第二步:聊天方法块只摘 SKILL.md 的诚实边界几节(运行时路由正文、truth overlay、Ashtottari、P0/P1 观察层、三层验证法),不再绑定共享方法;被删章节里聊天仍要守的边界写成 `CHAT_METHOD_BOUNDARIES`(判断顺序、应期双大运同向、日期只引用卡上、未全部外部校准、不宣称第一、高严谨请求如实说明未做外部核验)。SKILL.md 与共享方法文件不动,报告与 skill_read 照用。
|
||||
- 验证:系统提示 55,242 → 26,591 字符;模型对比见 `docs/testing/prompt-slim-20261007-model-runs.md`(flash 10 题平均输入 24,560 → 14,025 token、耗时 37.8 → 30.3 s、思考 6,975 → 5,683 token,自动检查与人工通读无退化;高严谨请求如实说明未做外部核验)。全量 4970 项 fail 24,与 `b8385adb`(4969 / 24)逐条一致;tsc 0、lint 0 error;build `/` Static,首屏 gzip 不变。改断言 2 个文件(三栏写在测试里),新增 2 条。
|
||||
- 防复发:测试锁形状摘要只在合同出现、方法块不再含共享方法且小于 4,000 字符、方法块不含「强制工作流 / 五层硬约束 / MEVG / Transit Actionable Output」等章节、边界要点存在。
|
||||
- 相关记录:BUG-1253(清单原句被念给用户,同属「给模型分析用的文字进了说话层」)、BUG-1255。
|
||||
- 修复版本:待发布
|
||||
|
||||
Reference in New Issue
Block a user