From c76fdda7ab2fd3cf489dc2b110cab6adc4114e6a Mon Sep 17 00:00:00 2001 From: Jesse_Chen Date: Wed, 7 Oct 2026 08:48:04 +0800 Subject: [PATCH] refactor(consult): chat binds the method's truth boundaries, not the engine manual (BUG-1256 step 2) The method block bound into every chat turn quoted most of SKILL.md (benchmark scores, CLI flags, oracle queues, file paths) plus the shared method, whose output contract asks for JSON verdicts, A/B/C/D confidence, audit tables, raw data and web verification - the opposite of the chat shape. Chat now quotes only the truth-boundary sections and adds CHAT_METHOD_BOUNDARIES for the limits from the dropped sections it still has to keep. SKILL.md and the shared method file are unchanged for reports and skill_read. Natal system prompt 55,242 -> 26,591 characters. DeepSeek A/B on public golden charts (flash 10 questions): input 24,560 -> 14,025 tokens, 37.8 -> 30.3 s, no regression in automatic checks or reading; a high-rigor request still says plainly that no external check was done. Full suite 4970 / fail 24, identical to b8385adb; build keeps / Static. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_017eEAG8HD3mm8gsKXgk8uU8 --- CHANGELOG.md | 6 ++ docs/BUG_HISTORY.md | 14 ++++ docs/tasks/README.md | 1 + .../prompt-slim-20261007-model-runs.md | 47 +++++++++++ frontend/src/mastra/skill-binding.ts | 80 ++++++++++--------- frontend/tests/skill-binding.test.ts | 44 ++++++---- 6 files changed, 142 insertions(+), 50 deletions(-) create mode 100644 docs/testing/prompt-slim-20261007-model-runs.md diff --git a/CHANGELOG.md b/CHANGELOG.md index 329589f7..78a1aed3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,11 @@ # 印度占星 Skill 更新日志 +## 2026-10-07 — 普通对话每轮带给模型的规则减半,回答更快(未上线) + +- 每轮发给模型的固定规则从约 5.5 万字符减到约 2.7 万:同一段回答形状不再重复 4 遍;占星方法书只带诚实边界几节,不再带引擎跑分、脚本命令和写给报告的输出格式。 +- 用 DeepSeek 实测同一组问题:每轮输入少 43%,平均回答快约 20%,回答质量与诚实边界未见退化(BUG-1256)。 +- Skill 文件本身不改,报告照用;不改数据库、模型、答题时钟和计费。 + ## 2026-10-07 — 回答换成知心姐姐的口气,可以说出你和对方可能的心情(未上线) - 回答的人设从「嘴有点毒的同事」改成懂行、说话温和的姐姐:不顺的地方照实说,但说法顾着你;不说教、不催;不用语气词卖萌。和首页开场语的口气统一。 diff --git a/docs/BUG_HISTORY.md b/docs/BUG_HISTORY.md index e14da184..bbfddbda 100644 --- a/docs/BUG_HISTORY.md +++ b/docs/BUG_HISTORY.md @@ -16809,3 +16809,17 @@ - 防复发:测试锁人设关键句、读心三条、禁用断定句、受冲只给范围,且 `index.ts` 不再有「concisely」、voice 不再有「restate / 嘴有点毒 / 锋利」。 - 相关记录:BUG-1070(被部分推翻)、BUG-1182(受冲口径保留)、BUG-1244(行动不强制)、BUG-1253。 - 修复版本:staging `b8385adb`(2026-10-07 08:13 部署核对:health gitCommit=b8385adb、各项 ok,`/api/account` 401,`/login` 200;门禁 run 1750 validate 15 分钟一次通过)。真机未走,状态保持 fixed-pending-verify + +## BUG-1256 | 普通对话系统提示 5.5 万字符:同一形状说 4 遍、方法书带进工程记录与冲突的输出合同 + +- 状态:fixed-pending-verify(分支 `codex/prompt-slim-20261007`;部署与真机前不标 resolved) +- 首次发现 / 最近更新:2026-10-07 / 2026-10-07 +- 来源:产品问「提示词规则 2 万多字能不能优化」,Claude 实测体量后产品授权直接执行,并提供临时模型 key 做改前改后对比。 +- 影响面:本命对话与申报时段对话的系统提示(`mastra/index.ts`、`product-voice.ts` 合同、`consultation-thinking-plan.ts` 的 natal 标题规则、`skill-binding.ts` 的方法块)。不改证据卡、工具、领域清单、用户轮指令、模型与时钟。 +- 现象:本命系统提示 55,242 字符(约 2.2 万 token),每轮实际输入约 2.46 万 token。英文形状摘要在合同、index、natal 标题规则、skill 骨架里各一份;合同又逐条复述中文 ANSWER SHAPE;方法书摘录 1.74 万字符里多为引擎跑分、CLI 参数、oracle 队列、文件路径;共享方法(全谱调用合同 + 事件判定骨架)要求输出 JSON verdict、A/B/C/D 置信度、审计表、原始数据与联网验证,与聊天形状冲突;一批通用回答政策被放在「When reference_transparency is present」标题下。 +- 根因:形状与方法块分轮叠加,每轮只加不删;方法块按 SKILL.md 的整节标题摘取,SKILL.md 主要是引擎与报告手册。 +- 修复:第一步(去重、不改意思):形状摘要只留合同一份,删合同里复述中文形状的条目,通用政策挪到「Always」标题下。第二步:聊天方法块只摘 SKILL.md 的诚实边界几节(运行时路由正文、truth overlay、Ashtottari、P0/P1 观察层、三层验证法),不再绑定共享方法;被删章节里聊天仍要守的边界写成 `CHAT_METHOD_BOUNDARIES`(判断顺序、应期双大运同向、日期只引用卡上、未全部外部校准、不宣称第一、高严谨请求如实说明未做外部核验)。SKILL.md 与共享方法文件不动,报告与 skill_read 照用。 +- 验证:系统提示 55,242 → 26,591 字符;模型对比见 `docs/testing/prompt-slim-20261007-model-runs.md`(flash 10 题平均输入 24,560 → 14,025 token、耗时 37.8 → 30.3 s、思考 6,975 → 5,683 token,自动检查与人工通读无退化;高严谨请求如实说明未做外部核验)。全量 4970 项 fail 24,与 `b8385adb`(4969 / 24)逐条一致;tsc 0、lint 0 error;build `/` Static,首屏 gzip 不变。改断言 2 个文件(三栏写在测试里),新增 2 条。 +- 防复发:测试锁形状摘要只在合同出现、方法块不再含共享方法且小于 4,000 字符、方法块不含「强制工作流 / 五层硬约束 / MEVG / Transit Actionable Output」等章节、边界要点存在。 +- 相关记录:BUG-1253(清单原句被念给用户,同属「给模型分析用的文字进了说话层」)、BUG-1255。 +- 修复版本:待发布 diff --git a/docs/tasks/README.md b/docs/tasks/README.md index 7a7d6b10..316962b2 100644 --- a/docs/tasks/README.md +++ b/docs/tasks/README.md @@ -429,3 +429,4 @@ | —(直接执行,无任务书) | `PROGRESS-consult-readable-20261006.md` | **回答铺满一行、输出不抖、不念分析清单**(10-06 真机截图):答案正文去 `text-wrap: pretty`(WebKit 整段等长 + 逐字重排,BUG-1252);清单开头写明只用于分析、婚恋三层改生活说法、字段白话对照、括号最多两条(BUG-1253) | 已验收 | Claude 直接执行;真机清单 `docs/testing/consult-readable-20261006.md` 待产品 | | —(直接执行,无任务书) | —(记录在 BUG-1254) | **门禁偶发红:数据库测试排队等槽位超时**:46 个 fixture 文件抢 2 个槽位、固定等 5 分钟,排在后面的随机失败;改为槽位易手即续期、连续 10 分钟无进展才失败(BUG-1254) | 已验收 | Claude 直接执行;门禁实际效果待看 | | —(直接执行,无任务书) | —(记录在 BUG-1255) | **回答口吻改知心姐姐、放开读心**(10-07 产品决定):人设统一;读心守三条(先答再读、可能/多半一句、当事人要有依据且未受冲);删冲突的旧规矩。推翻 BUG-1070 读心部分 | 已验收 | Claude 直接执行;真机清单 `docs/testing/consult-voice-sister-20261007.md` 待产品 | +| —(直接执行,无任务书) | —(记录在 BUG-1256) | **普通对话提示词减半**(10-07 产品授权 + 临时 key 对比):形状说明 4 份合 1、方法书只摘诚实边界、共享方法不再进聊天、边界要点改写补回;系统提示 55,242 → 26,591 字符,模型对比无退化 | 已验收 | Claude 直接执行;对比记录 `docs/testing/prompt-slim-20261007-model-runs.md` | diff --git a/docs/testing/prompt-slim-20261007-model-runs.md b/docs/testing/prompt-slim-20261007-model-runs.md new file mode 100644 index 00000000..fd6b0f71 --- /dev/null +++ b/docs/testing/prompt-slim-20261007-model-runs.md @@ -0,0 +1,47 @@ +# 模型对比 · 普通对话提示词精简(2026-10-07,BUG-1256) + +对比对象:改前 = staging `b8385adb` 的提示词;改后 = 分支 `codex/prompt-slim-20261007`(去重 + 方法书精简)。 + +做法:用仓库里真实引擎的公开名人 golden 盘(Steve Jobs / Barack Obama / Elizabeth Taylor)构造工具结果,与线上同一套拼法(系统提示 + 用户轮指令 + 工具调用 + 工具结果)直连 DeepSeek API,`max_tokens` 与线上一致(24,576)。两边的问题、盘面、工具结果完全相同,只换系统提示。脚本在会话临时目录,不入库;回答原文不入库,下面只记数字和结论。 + +## 体量 + +| 项 | 改前 | 改后 | +| --- | --- | --- | +| 本命系统提示 | 55,242 字符(约 2.2 万 token) | 26,591 字符(约 1.06 万 token) | +| 英文形状摘要出现次数 | 4 | 1 | +| 方法书摘录 + 共享方法 | 25,872 字符 | 约 2,500 字符 | +| 每轮实际输入(API 计量,含工具结果) | 平均 24,560 token | 平均 14,025 token(−43%) | + +## deepseek-flash,10 题(事业 / 婚恋 / 财运 / 父母 / 子女 / 健康 / 学业 / 是非题 ×2 / 问时间) + +| 指标 | 改前 | 改后 | +| --- | --- | --- | +| 平均耗时 | 37.8 s | 30.3 s(−20%) | +| 平均思考 token | 6,975 | 5,683(−19%) | +| 空回答 | 0(`max_tokens` 改为线上值后) | 0 | +| 自动检查(标题、列表、加粗、清单层名、数宫法、自造词 / 音译、置信标签、断定读心、后台词、空话、语气词) | 全 0 | 全 0 | +| 英文术语命中 | 1 题 | 1 题 | +| 「确定性措辞」命中 | 0 | 1 题,人工看是误报(原句「不是哪一天一定会发生什么」) | +| 是非题首句 | 「看情况」「不是」 | 「看情况」「不是」 | + +人工通读 1、3、8 题两版:结构、分段、读心、括号依据一致;改后第 8 题主动说明「另一套推运没对上,所以不给具体某一天」(精简后补的边界要点生效)。 + +## 高严谨请求(第 11 题:「拉满三大引擎和外部案例验证,精确到月,说明核对过哪些外部资料」) + +两版都如实说明未做外部核验。改后措辞更直白(「没有做外部引擎比对,没有联网核验,也没有拿真实案例回放校准……不假装做过」),并以两套推运未同向为由不给确定月份。耗时 42.3 s → 28.7 s。 + +## deepseek-v4-pro,抽 3 题 + +| 指标 | 改前 | 改后 | +| --- | --- | --- | +| 平均耗时 | 92.5 s | 33.6 s | +| 平均思考 token | 11,644 | 3,588 | +| 空回答 | 1(第 8 题思考 24,563 token 耗尽上限,181 s) | 1(第 7 题 4 s 返回空,重跑 3/3 正常,判为偶发) | + +第 7 题再各跑:改后 3 次、改前 2 次,均正常作答,首句为「不是 / 看情况」。 + +## 结论与缺口 + +- 精简后体量减半,速度更快、思考更少,回答质量与诚实边界未见退化。 +- 样本小(flash 10 题 + 1 题、pro 3 题),只用 3 张公开名人盘;真机仍需按 `consult-voice-sister-20261007.md` 与 `consult-readable-20261006.md` 两份清单走。 diff --git a/frontend/src/mastra/skill-binding.ts b/frontend/src/mastra/skill-binding.ts index b2d83107..41415a69 100644 --- a/frontend/src/mastra/skill-binding.ts +++ b/frontend/src/mastra/skill-binding.ts @@ -5,7 +5,6 @@ import { resolveLiveJyotishSkill, resolveLiveJyotishSkillRuntimePath, } from "../lib/skill-package-registry.ts"; -import { sharedConsultationMethodMarkdown } from "../lib/consultation-methodology.ts"; const skill = resolveLiveJyotishSkill(); @@ -18,24 +17,41 @@ export const jyotishSkillPackage = skill; export const jyotishSkillRuntimePath = resolveLiveJyotishSkillRuntimePath(skill); /** - * Headings from the commercial SKILL.md that govern answering a natal chart. - * The rest of that file is a local-agent / maintainer manual (CLI indexes, - * construction rules, celebrity case catalogs) and stays in the package for - * skill_read rather than being stuffed into every consultation prompt. + * What a chat answer takes from the commercial SKILL.md (BUG-1256). The file is + * mostly the engine and report manual: benchmark scores, CLI flags, oracle + * queues, file paths, and output contracts written for reports ("only output + * verdict + confidence + audit + raw evidence", "every transit prediction gives + * a period + action + [A]/[B]/[C]", "every reading must be web-verified"). + * Bound into every chat turn, those contradicted the chat shape and made up + * most of a 55k-character system prompt. Chat keeps the truth-boundary + * sections only; the full file stays in the package for skill_read. + * `body: true` keeps the section's own text (blockquotes and file-reading + * lists dropped); `subsections` are the `###` parts kept under it. */ -const RUNTIME_METHOD_HEADINGS = [ - "商业运行时路由", - "关联技法完整调取", - "强制工作流", - "五层硬约束", - "强制规则", - "当前最硬的未闭环点", - "核心方法论", - "强制规范速查", - "注意事项", +const CHAT_METHOD_SELECTION = [ + { section: "商业运行时路由", body: true, subsections: ["Skill truth overlay 硬边界", "Ashtottari"] }, + { section: "关联技法完整调取", body: false, subsections: ["P0/P1 观察层"] }, + { section: "核心方法论", body: false, subsections: ["三层验证法"] }, ] as const; -const DROPPED_RUNTIME_SUBHEADINGS = ["开源复用边界冻结"] as const; +/** Lines of a kept section body that only tell an agent which files to open. */ +function isFilePointerLine(line: string): boolean { + return line.startsWith(">") || /^\d+\.\s/.test(line) || line.includes("渐进读取"); +} + +/** + * The boundaries from the dropped sections that a chat answer still has to + * keep, in the product's own words (BUG-1256). Outside the verbatim skill + * block, so the excerpt stays a faithful quote of the live file. + */ +export const CHAT_METHOD_BOUNDARIES = ` +以下是方法书里没有摘进来的章节中,对话回答仍要守的边界(全文仍在技能目录里): +- 判断顺序:先看本命有没有这件事的底子,再看大运有没有把它激活,再看够不够落到现实,最后才谈时间;不从一段大运或一次行运直接跳到「会发生」。 +- 谈应期至少要 Vimshottari 与 Narayana 两套大运同向;两套冲突时降低把握,不给确定说法。 +- 大运起止日期只引用卡上给出的,不自己推算;大运精细边界和 Shadbala 绝对值并未全部完成外部校准,说日期和力量时不说「已校准」「精确到」。 +- 不宣称全球第一、所有技法已封顶或完美精度。 +- 用户要求「高严谨、拉满三大引擎、联网核验、拿真实案例验证」时:这次对话里没有做外部引擎比对、全网资料核验和真实案例校准,照实说明,结论降一级,不假装做过。 +`; function publishedSkillBody(): string { const raw = readFileSync(resolve(skill.resolvedPath, "SKILL.md"), "utf8"); @@ -43,21 +59,6 @@ function publishedSkillBody(): string { return (frontmatter ? raw.slice(frontmatter[0].length) : raw).trim(); } -function dropSubheadings(section: string, dropped: readonly string[]): string { - const lines = section.split("\n"); - const kept: string[] = []; - let skipping = false; - for (const line of lines) { - if (line.startsWith("### ")) { - skipping = dropped.some((heading) => line.includes(heading)); - } else if (line.startsWith("## ")) { - skipping = false; - } - if (!skipping) kept.push(line); - } - return kept.join("\n").trim(); -} - /** * Mastra answers a skill activation with the entrypoint *and* a flat listing of * every file in the package, which for this package measured 129,651 bytes: the @@ -97,8 +98,17 @@ function boundMethod(): string { const kept = sections.flatMap((section) => { const heading = section[0] ?? ""; - if (!RUNTIME_METHOD_HEADINGS.some((wanted) => heading.includes(wanted))) return []; - const text = dropSubheadings(section.join("\n"), DROPPED_RUNTIME_SUBHEADINGS); + const wanted = CHAT_METHOD_SELECTION.find((entry) => heading.includes(entry.section)); + if (!wanted) return []; + const parts: string[][] = [[]]; + for (const line of section) { + if (line.startsWith("### ")) parts.push([line]); + else parts[parts.length - 1]!.push(line); + } + const [own = [], ...subsections] = parts; + const body = wanted.body ? own.filter((line) => !isFilePointerLine(line)) : own.slice(0, 1); + const keptSubsections = subsections.filter((part) => wanted.subsections.some((name) => (part[0] ?? "").includes(name))); + const text = [...body, ...keptSubsections.flat()].join("\n").replace(/\n{3,}/g, "\n\n").trim(); return text.length > 0 ? [text] : []; }); @@ -118,9 +128,7 @@ function methodBlock(reportSkeleton: string) { ${BOUND_METHOD_MARKER} ${boundMethod()} - -${sharedConsultationMethodMarkdown()} -`; +${CHAT_METHOD_BOUNDARIES}`; } /** Method body + marker, without the natal Level 2 report skeleton. Window agents use this. */ diff --git a/frontend/tests/skill-binding.test.ts b/frontend/tests/skill-binding.test.ts index 7383fc00..e48efe6c 100644 --- a/frontend/tests/skill-binding.test.ts +++ b/frontend/tests/skill-binding.test.ts @@ -59,13 +59,20 @@ test("the bound method is the runtime excerpt, not the maintainer manual", () => const body = liveSkillBody(); const bound = boundMethodText(); - for (const heading of ["商业运行时路由", "关联技法完整调取", "强制工作流", "五层硬约束", "核心方法论", "注意事项", "commercial_skill_truth_overlay"]) { + // 原值: 摘录含「商业运行时路由 / 关联技法完整调取 / 强制工作流 / 五层硬约束 / 核心方法论 / 注意事项 / commercial_skill_truth_overlay」与「全谱系真实调用」「静态分析10步」 + // 新值: 只含诚实边界几节(运行时路由正文、truth overlay、Ashtottari、P0/P1 观察层、三层验证法);工程与报告用的章节留在技能目录 + // 原因: 2026-10-07 聊天提示词精简(BUG-1256 第二步):方法书只摘诚实边界几节,工程记录、跑分、脚本、文件路径和与聊天形状冲突的输出合同不再进每轮系统提示 + for (const heading of ["商业运行时路由", "关联技法完整调取", "核心方法论"]) { assert.ok(body.includes(heading), heading); assert.ok(bound.includes(heading), heading); } - for (const phrase of ["全谱系真实调用", "P0/P1 观察层", "网页对话的技法审计表由界面折叠展示", "静态分析10步"]) { + for (const phrase of ["P0/P1 观察层", "网页对话的技法审计表由界面折叠展示", "候选出生时间不得写成 confirmed", "Skill truth overlay 硬边界", "三层验证法"]) { assert.ok(bound.includes(phrase), phrase); } + for (const heading of ["强制工作流", "五层硬约束", "注意事项", "强制规范速查", "当前最硬的未闭环点", "MEVG 强制外部验证门控", "Transit Actionable Output", "全谱系真实调用", "静态分析10步"]) { + assert.ok(body.includes(heading), heading); + assert.equal(bound.includes(heading), false, heading); + } for (const heading of ["施工判断原则", "冲顶路线", "37大子命令", "验证与错题体系", "开源复用边界冻结"]) { assert.ok(body.includes(heading), heading); assert.equal(bound.includes(heading), false, heading); @@ -296,16 +303,25 @@ test("only chart-answering agents bind the method", () => { } }); -test("the system block binds the shared method sections verbatim from the package", () => { - const shared = sharedConsultationMethodMarkdown(); - const opened = jyotishSkillMethodBlock.indexOf(""); - const closed = jyotishSkillMethodBlock.indexOf(""); - assert.ok(opened >= 0 && closed > opened); - const boundShared = jyotishSkillMethodBlock.slice( - opened + "".length, - closed, - ).trim(); - assert.equal(boundShared, shared.trim()); - assert.match(shared, /Full-Spectrum Invocation Contract/); - assert.ok(shared.includes(readFileSync(resolve(jyotishSkillPackage.resolvedPath, "references/event_judgment_skeleton.md"), "utf8").trim())); +test("chat binds the method boundaries, not the report-shaped shared method (BUG-1256)", async () => { + // 原值: 系统块逐字绑定共享方法(Full-Spectrum Invocation Contract + event_judgment_skeleton.md) + // 新值: 不再绑定共享方法;改为 ,用产品自己的话保留聊天仍要守的边界 + // 原因: 2026-10-07 聊天提示词精简(BUG-1256 第二步):方法书只摘诚实边界几节,工程记录、跑分、脚本、文件路径和与聊天形状冲突的输出合同不再进每轮系统提示;共享方法要求输出 JSON verdict + A/B/C/D 置信度 + 审计表 + 原始数据 + 联网验证,与聊天形状冲突 + const { CHAT_METHOD_BOUNDARIES } = await import("../src/mastra/skill-binding.ts"); + for (const block of [jyotishSkillMethodBlock, jyotishSkillMethodCoreBlock]) { + assert.equal(block.includes(""), false); + assert.equal(block.includes("Full-Spectrum Invocation Contract"), false); + assert.ok(block.includes(CHAT_METHOD_BOUNDARIES)); + } + for (const kept of ["不从一段大运或一次行运直接跳到「会发生」", "Vimshottari 与 Narayana 两套大运同向", "大运起止日期只引用卡上给出的", "不宣称全球第一", "没有做外部引擎比对、全网资料核验和真实案例校准,照实说明"]) { + assert.ok(CHAT_METHOD_BOUNDARIES.includes(kept), kept); + } + // The shared method is still the package's own file, for reports and skill_read. + assert.match(sharedConsultationMethodMarkdown(), /Full-Spectrum Invocation Contract/); +}); + +test("the chat method block stays small (BUG-1256)", () => { + // It was 25,872 characters, most of the 55k natal system prompt. + assert.ok(jyotishSkillMethodBlock.length < 4_000, `${jyotishSkillMethodBlock.length}`); + assert.ok(jyotishSkillMethodCoreBlock.length < 4_000, `${jyotishSkillMethodCoreBlock.length}`); });