feat(research): consult biography backtest harness and pre-change baseline
TASK-consult-no-presupposition-and-backtest-20261001 T4 infrastructure only
(T1-T3 untouched). New files only, so it merges cleanly with the parallel
evidence-card brief.
- capture_consult_biography_backtest_golden.py: same handler/body/trim as the
evidence-card golden; nine public AA charts x parents/marriage/health/career.
- consult-biography-backtest-golden.json: real engine output, byte-reproducible.
- consult_biography_backtest_rubric.json: facts with sources, must_not,
expected_signals per figure x domain.
- consult-biography-backtest.mts: runs the product's real agent, tools, card,
methodology, user-turn shape and streamAgentResponse; deterministic checks only.
- Baseline on 9b937c4a: 39/72 severe biography conflicts; all five audit cases
reproduce in both runs.
BUG-1170 is registered when brief 3 lands.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
e5199106a7
commit
691ee440e3
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,41 @@
|
||||
# PROGRESS · 名人生平回测基础设施与基线(2026-10-01)
|
||||
|
||||
- 任务书:`docs/tasks/TASK-consult-no-presupposition-and-backtest-20261001.md`,本轮只做 **T4 的基础设施 + 改动前基线**;T1~T3(形状、示例、追问轮)不在本轮。
|
||||
- 执行方式:产品负责人要求子代理直接执行。
|
||||
- 基线:`origin/staging` `9b937c4a`
|
||||
- 分支 / worktree:`codex/consult-biography-backtest-20261001` / `.worktrees/consult-biography-backtest-20261001`(未推送)
|
||||
- 并行:另一会话在改证据卡(`TASK-consult-card-affliction-data-20261001`)。本轮**只新增文件、不改任何已有文件**,两边可以干净合并;`docs/tasks/README.md` 索引也没动,由第三单执行方登记。
|
||||
|
||||
## 交付
|
||||
|
||||
| 文件 | 内容 |
|
||||
| --- | --- |
|
||||
| `scripts/research/capture_consult_biography_backtest_golden.py` | 复用 `capture_consult_evidence_card_golden.py` 的 `trim()` 与 `consult_evidence_card_run` 的 handler、请求体;9 位 × 4 条引擎路由(parents 走 family);路由间相同的 chart 字段存一份 `shared_chart`,不同的(今天是 `modules.yogas`、`modules.chara_dasha`)留在各路由下。`PYTHONHASHSEED=0 JYOTISH_API_CHART_CACHE_TTL_SECONDS=0` 下两次输出逐字节相同 |
|
||||
| `frontend/tests/fixtures/consult-biography-backtest-golden.json` | 真实引擎输出,2.87 MB;不并入现有合同 golden |
|
||||
| `docs/research/consult_biography_backtest_rubric.json` | 9 位 × 父母 / 婚姻 / 健康 / 事业:生平事实(带来源)、`must_not`(含原文筛网)、`expected_signals`(引用卡字段)、评分口径 |
|
||||
| `frontend/scripts/research/consult-biography-backtest.mts` | 用产品真实模块组装(见测试文档「与真实路由的差异」);`--dump-cards` / `--recheck`;确定性检查只做禁句、性别未知代词、must_not 原文 |
|
||||
| `docs/testing/consult-affliction-backtest-20261001.md` | 名单、保真度缺口、基线表与合计、代表句、成本 |
|
||||
|
||||
## 基线结论(详见测试文档)
|
||||
|
||||
72 份里严重冲突 39 份;排查记录的 5 份严重冲突 10 / 10 次复现;新增 4 位严重冲突 20 / 32(通过线 ≤ 1);对照组误报 1 份(轻)。新发现:事业题 16 / 18 份把人写成「专业人 / 顾问 / 幕后」,与盘无关,第三单不覆盖,建议另开单。
|
||||
|
||||
## 门禁
|
||||
|
||||
| 项 | 结果 |
|
||||
| --- | --- |
|
||||
| `tsc --noEmit`(`tsconfig.json` 的 include 含 `**/*.mts`,脚本在范围内) | 0 错 |
|
||||
| `eslint`(`npm run lint` 的配置覆盖 `.mts`) | 0 error |
|
||||
| 既有测试 | 未改任何测试文件 |
|
||||
| `npm test`(Node 22) | 4868 项,pass 4803,fail 24,cancelled 0;24 项全是需要 Docker / PostgreSQL / 部署环境的套件(`startPostgresFixture` 等),本机无 Docker,属已知环境缺口,与本轮新增文件无关 |
|
||||
| Python 快速门 `run_quality_gate.py --profile quick`(本机无 `.venv`,用系统 python3) | pytest 1036 passed / 1 skipped;门禁内的 `npm test` 被系统 Node 20 跑,70 项因 Node 20 不支持模块 mock 失败(环境),用 Node 22 单跑结果见上一行 |
|
||||
| `tests/test_repo_privacy_markers.py` | 通过 |
|
||||
|
||||
## BUG 编号
|
||||
|
||||
本轮是基础设施,不登记 BUG。BUG-1170(T4 名人生平回测)在第三单(`TASK-consult-no-presupposition-and-backtest-20261001`)合入时登记,引用本进度记录与测试文档的基线。
|
||||
|
||||
## 环境备注
|
||||
|
||||
- 模型 key 只从环境变量读,没有写入任何文件、日志或提交。
|
||||
- 模型输出在 `scratch/consult-biography-backtest/baseline-20261001/`(gitignore),未提交。
|
||||
@@ -0,0 +1,119 @@
|
||||
# 普通对话名人生平回测(2026-10-01)
|
||||
|
||||
- 任务书:`docs/tasks/TASK-consult-no-presupposition-and-backtest-20261001.md` T4(D5:普通对话提示词或清单改动的固定验收)
|
||||
- 排查记录:`docs/research/consult_affliction_audit_2026_10_01.md`
|
||||
- 名单与评分表:`docs/research/consult_biography_backtest_rubric.json`
|
||||
- golden:`frontend/tests/fixtures/consult-biography-backtest-golden.json`(`scripts/research/capture_consult_biography_backtest_golden.py` 生成,与 `capture_consult_evidence_card_golden.py` 同一 handler、同一请求体)
|
||||
- 脚本:`frontend/scripts/research/consult-biography-backtest.mts`
|
||||
- 模型输出不入库,默认写到 `scratch/consult-biography-backtest/<时间戳>/`(已 gitignore);本文只摘短句。
|
||||
|
||||
## 怎么跑
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
DS_KEY=<临时 key> ./node_modules/.bin/tsx scripts/research/consult-biography-backtest.mts --repeats 2 --concurrency 6
|
||||
# 只看模型拿到的卡,不调模型:
|
||||
./node_modules/.bin/tsx scripts/research/consult-biography-backtest.mts --dump-cards --out /tmp/cards
|
||||
# 改了评分表后,对旧输出重算确定性检查:
|
||||
./node_modules/.bin/tsx scripts/research/consult-biography-backtest.mts --recheck <输出目录>
|
||||
```
|
||||
|
||||
需要 Node 22(本机 `/exec-daemon/node`)。脚本不进 CI,需要外网与模型 key。key 只从环境变量 `DS_KEY` / `BACKTEST_MODEL_KEY` 读。
|
||||
|
||||
## 名单
|
||||
|
||||
| 名人 | 测什么 | 出生资料 |
|
||||
| --- | --- | --- |
|
||||
| Steve Jobs | 出生即送养;晚年重病(胰腺肿瘤,56 岁去世);顶级企业家 | AA,`minute_rectification_development_v1.json` |
|
||||
| Barack Obama | 父亲两岁离开、外祖父母带大;婚姻稳定、健康良好(这两项作对照) | AA,`minute_rectification_holdout_v4.json` |
|
||||
| Elizabeth Taylor | 八次婚姻、丧偶;终身多病、数十次手术;童星 | AA,同上 |
|
||||
| Marilyn Monroe | 生父不明、母亲精神病住院、寄养家庭与孤儿院;三次离婚;慢性病与药物依赖 | AA,`minute_rectification_holdout_v5.json` |
|
||||
| Judy Garland | 父亲在她 13 岁去世、强势母亲;五次婚姻;终身药物依赖;童星 | AA,同上 |
|
||||
| Édith Piaf | 被母亲遗弃、祖母带大;两次婚姻、挚爱空难;车祸与吗啡依赖 | AA,同上 |
|
||||
| Frida Kahlo | 小儿麻痹、车祸、约 30 次手术、截肢;与同一人离婚又复婚;与父亲亲近、与母亲冷淡 | AA,同上 |
|
||||
| George W. Bush(对照) | 父母都在且亲近、1977 年至今一段婚姻、健康总体良好 | AA,同上 |
|
||||
| Zinedine Zidane(对照) | 父母都在且亲近、1994 年至今一段婚姻、运动员健康 | AA,同上 |
|
||||
|
||||
Bill Clinton 不在仓库案例库里,没有加;「父亲缺席」由 Obama、Monroe、Garland、Jobs 覆盖。
|
||||
|
||||
## 与真实路由的差异(保真度缺口)
|
||||
|
||||
脚本直接用产品的 `getJyotishAgent`(系统提示、绑定的技能方法块、工具)、Mastra `Agent.stream` + `consultationNatalPrepareStep` + `consultationGenerationSettings`、`createConsultationTools`(卡片、方法论清单、`read-consultation-evidence` 查卡)、`consultationUserTurnContent` + `natalUserTurnShape`,以及 `streamAgentResponse`(按 step 取答案、Pass 4、合同重试、空答重试、长度续写)。差异:
|
||||
|
||||
1. 引擎调用换成 golden:同 capture 路径、raman 岁差(产品默认也是 raman)、mean 交点、参考日 2026-09-27,外部 VedAstro 不调用。
|
||||
2. 用户回合不带「用户称呼」:真实路由会带档案名;这里故意不给名人名字,免得模型凭记忆答生平。
|
||||
3. `natalToolInstruction` 一句和 `currentTimeContext` 是 `route.ts` 内部函数、未导出,脚本按原文复刻(两处字符串)。改了这两处要同步改脚本。
|
||||
4. 没有内容审核、计费、标题旁路事件、历史摘要、请求时钟(真实路由有工具期与 5 分钟答题钟;本轮最长一份 87 s,不触发);单轮、无历史。
|
||||
5. 性别一律按档案未填(性别未知)。
|
||||
6. 模型经 Mastra 的 OpenAI 兼容通道直连 DeepSeek(`deepseek-flash`),不经产品数据库的模型目录;产品的 staging 模型是 `deepseek-v4-flash`,两者是否同一模型以供应商为准。
|
||||
7. 计划用 `createConsultationPlan({ userIntent, theme })`;真实路由用预留阶段的计划(含点数与深度)。
|
||||
8. 模型若自己多选一个领域(本轮 1 次:Obama 父母 run1 多选了 `family`),`family` / `children` 一律回放 `parents` 的 golden(引擎同为 family 路由,只差问题文本)。
|
||||
|
||||
## 基线(改动前)
|
||||
|
||||
- 代码:`origin/staging` `9b937c4a`(第一、二、三单都未合入)
|
||||
- 模型:`deepseek-flash`,thinking 开,输出上限 24576(产品设置)
|
||||
- 9 位 × 4 领域 × 2 次 = 72 份,失败 0 份
|
||||
- 评分:Claude 逐份阅读,按评分表 `how_to_score`:通过 / 漏读 / 读反 / 严重冲突 / 误报
|
||||
|
||||
| 名人 | 父母 r1 / r2 | 婚姻 r1 / r2 | 健康 r1 / r2 | 事业 r1 / r2 |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Steve Jobs | 严重冲突 / 严重冲突 | 通过 / 通过 | 读反 / 严重冲突 | 严重冲突 / 严重冲突 |
|
||||
| Barack Obama | 严重冲突 / 严重冲突 | 通过 / 通过(预设未婚) | 通过 / 通过 | 严重冲突 / 严重冲突 |
|
||||
| Elizabeth Taylor | 通过 / 通过 | 严重冲突 / 严重冲突 | 严重冲突 / 严重冲突 | 严重冲突 / 严重冲突 |
|
||||
| Marilyn Monroe | 严重冲突 / 严重冲突 | 严重冲突 / 严重冲突 | 读反 / 读反 | 严重冲突 / 严重冲突 |
|
||||
| Judy Garland | 严重冲突 / 严重冲突 | 通过(弱) / 严重冲突 | 严重冲突 / 严重冲突 | 严重冲突 / 严重冲突 |
|
||||
| Édith Piaf | 严重冲突 / 严重冲突 | 通过(弱) / 严重冲突 | 读反 / 读反 | 通过 / 严重冲突 |
|
||||
| Frida Kahlo | 漏读 / 漏读 | 读反 / 读反 | 严重冲突 / 严重冲突 | 漏读 / 严重冲突 |
|
||||
| George W. Bush(对照) | 通过 / 通过 | 通过 / 通过 | 通过 / 通过 | 严重冲突 / 严重冲突 |
|
||||
| Zinedine Zidane(对照) | 通过 / 通过 | 误报(轻) / 通过 | 通过 / 通过 | 严重冲突 / 严重冲突 |
|
||||
|
||||
### 合计
|
||||
|
||||
| 项 | 数 |
|
||||
| --- | --- |
|
||||
| 严重冲突 | 39 / 72 份 |
|
||||
| 读反 | 7 |
|
||||
| 漏读 | 3 |
|
||||
| 通过 | 22 |
|
||||
| 对照组误报 | 1(轻,Zidane 婚姻 run1「旧联系重新变近、分开又见面」) |
|
||||
| 两次都通过的格子 | 9 / 36 |
|
||||
| 排查记录里的 5 份严重冲突(Obama 父母、Jobs 父母、Taylor 婚姻 / 健康 / 事业) | 10 / 10 份仍冲突,原样复现 |
|
||||
| 新增 4 位(Monroe、Garland、Piaf、Kahlo)严重冲突 | 20 / 32 份(通过线 ≤ 1) |
|
||||
| 确定性检查:禁句(逆行…反复/来回) | 7 处,6 份 |
|
||||
| 确定性检查:性别未知时婚恋题用他/她 | 5 处(3 处指伴侣,2 处是把金星称「她」) |
|
||||
| 确定性检查:must_not 原文 | 27 处(是筛网,最终以阅读为准;对照组 Bush 健康「慢性病」一处属正常描述) |
|
||||
|
||||
按领域:父母 10 份严重冲突(缺席的父亲一律写成「话少、分量重、靠做事表达」);婚姻 6 份(多次婚姻写成「定下来之后很稳」);健康 7 份严重 + 5 份读反(终身重病写成「底子不差、恢复力强」);事业 16 / 18 份严重冲突,连两位对照组也冲突。
|
||||
|
||||
### 代表句(每条只摘短句)
|
||||
|
||||
- 父母·缺席的父亲:Obama「他话不多,存在感却稳……你做事的方式里有一部分就是照着他的样子长出来的」;Monroe「爸爸话少、分量重,是你衡量自己的那把尺」;Garland「父亲这条线表面上不显眼,但分量不轻」;Jobs「他们一直在,而且投入得很深」「挑一件具体的事去问爸爸的意见」。
|
||||
- 父母·被遗弃的母亲:Piaf「她是你家里管事最多的那个人」。
|
||||
- 婚姻:Taylor「你的婚姻是有的,而且底子偏硬……定下来之后反而稳」;Monroe「最后落到的那段关系……事情交给他你放心」;Garland「婚姻是往后走、慢慢稳下来的类型」。
|
||||
- 健康:Taylor「身体底子属于中上:扛得住,也恢复得回来」;Kahlo「你不是一张虚弱的盘」;Garland「这不是有严重疾病迹象」;Jobs「你不是被病拖垮的类型」。
|
||||
- 事业:Taylor「偏技术、分析和系统……纯靠嘴皮子和人缘往前推的事,你做起来会别扭」;Obama「顶是深度,不是高度……不是管一个大组织那种顶」;Zidane「真正的手艺在嘴和脑子上:写、讲、分析、算、谈判」;Monroe「靠一门具体的手艺慢慢被人认出来」。
|
||||
- 杂项:Taylor 事业 run1 正文出现「加上海王……更正一下」。
|
||||
|
||||
### 基线读出来的东西(给第三单和验收)
|
||||
|
||||
1. 排查记录的 C 类(形状预设对象在场)在 4 位新名人上全部复现:只要问父母,缺席的一方就被写成「话少、靠做事」。第三单 T1 的新句正对这一条。
|
||||
2. 事业题出现排查记录没有单列的一类:16 / 18 份把人写成「专业人 / 顾问 / 幕后」,与盘无关(总统、球星、电影明星都一样)。来源看起来是同一形状要求加上「接触层 / 结构层 / 落地层」三层清单,模型把「不预设上班族」理解成了「默认专业服务」。AL 落 10 宫、10 宫 SAV 高这些成名信号几乎没被读出来。第三单不覆盖这一条,需要另开单或在第三单决策记录里补。
|
||||
3. 健康题默认开头「底子不差」,8 宫、6 宫主落 8 宫等受冲只降级成「慢性消耗」。第一单(健康卡加 1 / 12 宫)和第二单(受冲读法)是对口的修法。
|
||||
4. 婚姻题有清单,表现最好(9 / 18 通过),但多次婚姻的三位仍有 6 份写成稳定。
|
||||
5. 对照组只有 1 份轻度误报:改动后要重点看对照组有没有被「先说受冲、请确认」带出新的误报。
|
||||
|
||||
## 改动后
|
||||
|
||||
(第三单 T1~T3 合入后,用同一命令重跑,按上表格式填;通过线见任务书 T4 第 5 条。)
|
||||
|
||||
## 成本(一次全量)
|
||||
|
||||
| 项 | 数 |
|
||||
| --- | --- |
|
||||
| 份数 | 72(9 × 4 × 2) |
|
||||
| 墙钟 | 664 s,并发 6;单份中位 54 s,最长 87 s |
|
||||
| 输入 token | 3,480,792,其中命中缓存 3,187,584(约 92%) |
|
||||
| 输出 token | 761,012,其中推理 664,649 |
|
||||
| 单份平均 | 输入约 48k、输出约 10.6k |
|
||||
| 查卡工具 | 72 份里 19 次调用了 `read-consultation-evidence` |
|
||||
@@ -0,0 +1,537 @@
|
||||
/**
|
||||
* Biography backtest for the ordinary (natal) consultation answer
|
||||
* (TASK-consult-no-presupposition-and-backtest-20261001 T4, D5).
|
||||
*
|
||||
* Runs the product's real consultation agent against public Rodden AA charts
|
||||
* and writes each answer for a human / Claude to score against
|
||||
* docs/research/consult_biography_backtest_rubric.json. Nothing here copies
|
||||
* prompt text: the system prompt, the skill method block, the tools, the
|
||||
* evidence card, the methodology checklist, the user-turn shape and the answer
|
||||
* stream handling all come from the product modules, so a run always tests the
|
||||
* prompts on the current checkout.
|
||||
*
|
||||
* What is real Where it comes from
|
||||
* system prompt + bound skill method getJyotishAgent (src/mastra/index.ts)
|
||||
* tool loop, prepareStep, settings Mastra Agent.stream + consultationNatalPrepareStep
|
||||
* + consultationGenerationSettings (route.ts natalStreamOptions)
|
||||
* tool result (contract, card, createConsultationTools -> toModelEvidenceView
|
||||
* methodology checklist) (engine call replaced by the golden workflow)
|
||||
* evidence lookup tool the real read-consultation-evidence
|
||||
* user turn consultationUserTurnContent + natalUserTurnShape
|
||||
* answer extraction, pass 4, retries, streamAgentResponse (stepScopedAnswer, pass4Mode,
|
||||
* length continuation retry / retryForAnswer / continueAfterLength as route.ts)
|
||||
*
|
||||
* Gaps against app/api/consult/route.ts are listed in GAPS below and in
|
||||
* docs/testing/consult-affliction-backtest-20261001.md.
|
||||
*
|
||||
* Usage (from frontend/, Node 22):
|
||||
* DS_KEY=... ./node_modules/.bin/tsx scripts/research/consult-biography-backtest.mts \
|
||||
* [--out <dir>] [--repeats 2] [--figures a,b] [--domains parents,marriage,health,career] \
|
||||
* [--model deepseek-flash] [--base-url https://api.deepseek.com] [--concurrency 4] [--dump-cards]
|
||||
*
|
||||
* --dump-cards writes the model-visible tool result per figure x domain and
|
||||
* calls no model. --recheck <dir> re-runs the deterministic checks of an
|
||||
* earlier run against the current rubric and rewrites its summary. The key is read from DS_KEY (or BACKTEST_MODEL_KEY) only and
|
||||
* is never written anywhere. Output defaults to <repo>/scratch/ (gitignored);
|
||||
* only the summary may be quoted into docs.
|
||||
*/
|
||||
import { mkdirSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import { dirname, join, resolve } from "node:path";
|
||||
import { fileURLToPath } from "node:url";
|
||||
|
||||
import { getJyotishAgent } from "../../src/mastra/index.ts";
|
||||
import {
|
||||
AGENT_MAX_STEPS,
|
||||
consultationContinueGenerationSettings,
|
||||
consultationGenerationSettings,
|
||||
consultationNatalPrepareStep,
|
||||
consultationStepBudgetReceipt,
|
||||
createConsultationAgentContext,
|
||||
createConsultationRuntimeState,
|
||||
publicConsultationRuntimeSteps,
|
||||
type ConsultationRuntimeState,
|
||||
} from "../../src/mastra/consultation-tools.ts";
|
||||
import {
|
||||
consultationWorkflowResponseSchema,
|
||||
minuteSensitiveThemesFromBirthTimeSensitivity,
|
||||
} from "../../src/mastra/consultation-workflow.ts";
|
||||
import type { ResolvedLanguageModel } from "../../src/mastra/model.ts";
|
||||
import { createConsultationPlan } from "../../src/lib/consultation-plan.ts";
|
||||
import { consultationUserTurnContent } from "../../src/lib/consultation-session-history.ts";
|
||||
import { consultationContinueMessages, natalUserTurnShape } from "../../src/lib/consultation-thinking-plan.ts";
|
||||
import { streamAgentResponse } from "../../src/lib/stream-agent-response.ts";
|
||||
import type { ConsultationDomain } from "../../src/lib/consultation-domain-registry.ts";
|
||||
import type { ServerChartConsultation } from "../../src/lib/consultation-route-service.ts";
|
||||
|
||||
const HERE = dirname(fileURLToPath(import.meta.url));
|
||||
const FRONTEND = resolve(HERE, "../..");
|
||||
const ROOT = resolve(FRONTEND, "..");
|
||||
const GOLDEN = join(FRONTEND, "tests/fixtures/consult-biography-backtest-golden.json");
|
||||
const RUBRIC = join(ROOT, "docs/research/consult_biography_backtest_rubric.json");
|
||||
|
||||
/** Where this harness differs from the live route; repeated in the summary. */
|
||||
export const GAPS = [
|
||||
"引擎调用换成 golden(同 capture 路径、raman 岁差、mean 交点、参考日 2026-09-27);产品用用户档案的岁差设置。",
|
||||
"用户回合不带「用户称呼」:真实路由会带档案名,这里故意不给名人名字,免得模型凭记忆答生平。",
|
||||
"natalToolInstruction 一句与 currentTimeContext 是 route.ts 内部函数,未导出,脚本里按原文复刻(见 routeToolInstruction / routeCurrentTime)。",
|
||||
"没有内容审核(moderateOutput)、计费、标题旁路事件、历史摘要;单轮、无历史。",
|
||||
"性别一律按档案未填(性别未知)处理。",
|
||||
"模型经 Mastra 的 OpenAI 兼容通道直连 DeepSeek(model id 由参数给),不经产品数据库的模型目录。",
|
||||
] as const;
|
||||
|
||||
// route.ts (natalToolInstruction, entrypoint undefined = ordinary chat). Not exported there.
|
||||
const routeToolInstruction = "如需新的个人星盘结论,必须调用服务器绑定的排盘工具。";
|
||||
|
||||
// route.ts currentTimeContext. Not exported there.
|
||||
function routeCurrentTime(now: Date) {
|
||||
const chinaTime = new Date(now.getTime() + 8 * 60 * 60 * 1000).toISOString().replace("T", " ").slice(0, 19);
|
||||
return `服务端当前时间(权威):${now.toISOString()};中国标准时间(UTC+8):${chinaTime}。涉及“现在、今天、今年、未来几个月”等相对时间时,以此为准。`;
|
||||
}
|
||||
|
||||
// The golden is computed for this reference date; the request clock matches it.
|
||||
const REQUEST_TIME = new Date("2026-09-27T04:00:00.000Z");
|
||||
|
||||
type Json = Record<string, unknown>;
|
||||
type GoldenRoute = { domain: string; engine_route: string; question: string };
|
||||
type GoldenFigure = {
|
||||
id: string;
|
||||
label: string;
|
||||
rodden_rating: string;
|
||||
shared_chart: Json;
|
||||
routes: Record<string, Json & { chart?: Json; chart_modules?: Json }>;
|
||||
};
|
||||
type Golden = { routes: GoldenRoute[]; figures: GoldenFigure[] };
|
||||
|
||||
type RubricFigure = {
|
||||
id: string;
|
||||
domains: Record<string, { must_not?: { claim: string; literals?: string[] }[] }>;
|
||||
};
|
||||
|
||||
function parseArgs(argv: readonly string[]) {
|
||||
const value = (name: string) => {
|
||||
const index = argv.indexOf(`--${name}`);
|
||||
return index >= 0 ? argv[index + 1] : undefined;
|
||||
};
|
||||
const stamp = new Date().toISOString().replace(/[:.]/g, "-");
|
||||
return {
|
||||
out: resolve(value("out") ?? join(ROOT, "scratch/consult-biography-backtest", stamp)),
|
||||
repeats: Math.max(1, Number(value("repeats") ?? 2)),
|
||||
figures: value("figures")?.split(",").filter(Boolean) ?? null,
|
||||
domains: value("domains")?.split(",").filter(Boolean) ?? ["parents", "marriage", "health", "career"],
|
||||
model: value("model") ?? "deepseek-flash",
|
||||
baseUrl: value("base-url") ?? "https://api.deepseek.com",
|
||||
concurrency: Math.max(1, Number(value("concurrency") ?? 4)),
|
||||
dumpCards: argv.includes("--dump-cards"),
|
||||
recheck: value("recheck"),
|
||||
};
|
||||
}
|
||||
|
||||
/** One route's engine response, with the figure's shared chart put back exactly as captured. */
|
||||
function routeWorkflow(figure: GoldenFigure, domain: string): Json {
|
||||
const route = figure.routes[domain];
|
||||
if (!route) throw new Error(`golden has no ${domain} route for ${figure.id}`);
|
||||
const { chart: chartOverrides, chart_modules: moduleOverrides, ...rest } = structuredClone(route);
|
||||
const shared = structuredClone(figure.shared_chart);
|
||||
const chart = {
|
||||
...shared,
|
||||
...(chartOverrides ?? {}),
|
||||
modules: { ...(shared.modules as Json), ...(moduleOverrides ?? {}) },
|
||||
};
|
||||
return { ...rest, chart };
|
||||
}
|
||||
|
||||
// The engine route each product domain runs as (consultation-domain-registry):
|
||||
// parents / children / family all run the engine's family route.
|
||||
const GOLDEN_ROUTE_FOR_DOMAIN: Record<string, string> = {
|
||||
parents: "parents",
|
||||
family: "parents",
|
||||
children: "parents",
|
||||
marriage: "marriage",
|
||||
health: "health",
|
||||
career: "career",
|
||||
};
|
||||
|
||||
function stubRunWorkflow(figure: GoldenFigure, executed: string[]) {
|
||||
return async (input: { theme: string }) => {
|
||||
const key = GOLDEN_ROUTE_FOR_DOMAIN[input.theme];
|
||||
if (!key) throw new Error(`backtest golden has no engine route for domain ${input.theme}`);
|
||||
executed.push(input.theme);
|
||||
// runConsultationWorkflow: schema parse, then minute-sensitive themes on the policy.
|
||||
const data = consultationWorkflowResponseSchema.parse(routeWorkflow(figure, key));
|
||||
const consumer = data.consumer_context as Json;
|
||||
return {
|
||||
...data,
|
||||
consumer_context: {
|
||||
...consumer,
|
||||
answer_policy: {
|
||||
...(consumer.answer_policy as Json),
|
||||
minute_sensitive_themes: minuteSensitiveThemesFromBirthTimeSensitivity(data.birth_time_sensitivity),
|
||||
},
|
||||
},
|
||||
} as unknown as typeof data;
|
||||
};
|
||||
}
|
||||
|
||||
function serverChart(figure: GoldenFigure): ServerChartConsultation {
|
||||
const birth = (figure.shared_chart.birth ?? {}) as Json;
|
||||
const route = routeWorkflow(figure, "parents");
|
||||
const body = ((route.chart as Json).birth ?? birth) as Json;
|
||||
const num = (key: string, fallback = 0) => (typeof body[key] === "number" ? body[key] as number : fallback);
|
||||
// Birth fields only shape the tool input schema; the stub ignores them.
|
||||
const toolInput = {
|
||||
year: num("year", 1950),
|
||||
month: num("month", 1),
|
||||
day: num("day", 1),
|
||||
hour: num("hour"),
|
||||
minute: num("minute"),
|
||||
city: "public chart",
|
||||
lat: num("lat", num("latitude")),
|
||||
lon: num("lon", num("longitude")),
|
||||
tz: num("tz", num("timezone")),
|
||||
ayanamsa: "raman",
|
||||
declared_accuracy: "minute",
|
||||
time_source: "birth_record",
|
||||
} as unknown as ServerChartConsultation["toolInput"];
|
||||
return {
|
||||
name: "backtest",
|
||||
toolInput,
|
||||
truth: {
|
||||
birthDate: "1950-01-01",
|
||||
reportedBirthTime: null,
|
||||
activeBirthTime: null,
|
||||
selectedTimeKind: "reported",
|
||||
birthTimeSource: "birth_record",
|
||||
birthTimeStatus: "reported",
|
||||
placeLabel: "public chart",
|
||||
placeCodes: { countryCode: null, provinceCode: null, cityCode: null, districtCode: null },
|
||||
placeId: null,
|
||||
placeType: null,
|
||||
placeProvider: null,
|
||||
timezoneId: null,
|
||||
timezoneSource: null,
|
||||
latitude: toolInput.lat,
|
||||
longitude: toolInput.lon,
|
||||
timezoneOffset: toolInput.tz,
|
||||
},
|
||||
gender: null,
|
||||
} as ServerChartConsultation;
|
||||
}
|
||||
|
||||
function model(args: ReturnType<typeof parseArgs>, apiKey: string): ResolvedLanguageModel {
|
||||
return {
|
||||
id: `backtest-${args.model}`,
|
||||
label: args.model,
|
||||
description: "biography backtest",
|
||||
creditCost: 0,
|
||||
isDefault: false,
|
||||
mode: "compatible",
|
||||
model: { providerId: "deepseek", modelId: args.model, url: args.baseUrl, apiKey } as ResolvedLanguageModel["model"],
|
||||
};
|
||||
}
|
||||
|
||||
type Job = { figure: GoldenFigure; domain: string; question: string; run: number };
|
||||
|
||||
type Usage = { inputTokens?: number; outputTokens?: number; reasoningTokens?: number; cachedInputTokens?: number };
|
||||
|
||||
function addUsage(total: Usage, usage: unknown) {
|
||||
const row = (usage ?? {}) as Record<string, unknown>;
|
||||
for (const key of ["inputTokens", "outputTokens", "reasoningTokens", "cachedInputTokens"] as const) {
|
||||
const value = row[key];
|
||||
if (typeof value === "number") total[key] = (total[key] ?? 0) + value;
|
||||
}
|
||||
}
|
||||
|
||||
function agentContext(job: Job, state: ConsultationRuntimeState, executed: string[]) {
|
||||
return createConsultationAgentContext({
|
||||
userId: "biography-backtest",
|
||||
sessionId: `backtest-${job.figure.id}`,
|
||||
requestId: `backtest-${job.figure.id}-${job.domain}-${job.run}`,
|
||||
consultationMode: "verified_chart",
|
||||
plan: createConsultationPlan({ userIntent: job.question, theme: job.domain as ConsultationDomain }),
|
||||
theme: job.domain as ConsultationDomain,
|
||||
followUpTurn: false,
|
||||
serverChart: serverChart(job.figure),
|
||||
state,
|
||||
runWorkflow: stubRunWorkflow(job.figure, executed) as never,
|
||||
});
|
||||
}
|
||||
|
||||
function userTurn(question: string) {
|
||||
return consultationUserTurnContent({
|
||||
currentTime: routeCurrentTime(REQUEST_TIME),
|
||||
instruction: `${routeToolInstruction}${natalUserTurnShape({ history: [] })}`,
|
||||
question,
|
||||
});
|
||||
}
|
||||
|
||||
async function runJob(job: Job, resolved: ResolvedLanguageModel) {
|
||||
const state = createConsultationRuntimeState();
|
||||
const executed: string[] = [];
|
||||
const ctx = agentContext(job, state, executed);
|
||||
const agent = getJyotishAgent(resolved, ctx);
|
||||
const runId = ctx.requestId;
|
||||
const baseMessages = [{ role: "user" as const, content: userTurn(job.question) }];
|
||||
const streamOptions = {
|
||||
runId,
|
||||
maxSteps: AGENT_MAX_STEPS,
|
||||
...consultationGenerationSettings(resolved.model),
|
||||
};
|
||||
const natalStreamOptions = { ...streamOptions, prepareStep: consultationNatalPrepareStep };
|
||||
const usages: Promise<unknown>[] = [];
|
||||
const startedAt = Date.now();
|
||||
let output: string | null = null;
|
||||
let failure: string | null = null;
|
||||
const first = await agent.stream(baseMessages, natalStreamOptions);
|
||||
usages.push(first.totalUsage);
|
||||
const response = streamAgentResponse({
|
||||
runId,
|
||||
requestId: runId,
|
||||
state,
|
||||
stream: first.fullStream as never,
|
||||
requireTool: true,
|
||||
retry: async () => {
|
||||
const retried = await agent.stream([
|
||||
...baseMessages,
|
||||
{ role: "user" as const, content: "运行合同不完整:本次尚未取得服务器计算结果。请调用 run-jyotish-consultation 完成计算,再据此回答;不要在工具参数中添加出生资料。" },
|
||||
], natalStreamOptions);
|
||||
usages.push(retried.totalUsage);
|
||||
return retried.fullStream as never;
|
||||
},
|
||||
retryForAnswer: async (retryHint?: string) => {
|
||||
const retried = await agent.stream([
|
||||
...baseMessages,
|
||||
{ role: "user" as const, content: `服务器计算已经完成,但上一轮没有输出任何回答文本。请重新取回本次计算结果,然后直接给出回答;不要只描述过程或工具调用。${retryHint ? `\n${retryHint}` : ""}` },
|
||||
], natalStreamOptions);
|
||||
usages.push(retried.totalUsage);
|
||||
return retried.fullStream as never;
|
||||
},
|
||||
continueAfterLength: async (text: string, evidence?: unknown) => {
|
||||
const continued = await agent.stream(consultationContinueMessages(baseMessages, text, evidence), {
|
||||
...streamOptions,
|
||||
...consultationContinueGenerationSettings(resolved.model),
|
||||
});
|
||||
usages.push(continued.totalUsage);
|
||||
return continued.fullStream as never;
|
||||
},
|
||||
stepScopedAnswer: true,
|
||||
pass4Mode: "verified_chart",
|
||||
toolStatus: () => {
|
||||
const status = state.workflowReceipt?.status;
|
||||
return status === "ready" || status === "degraded" ? status : "blocked";
|
||||
},
|
||||
receipt: () => ({
|
||||
runId,
|
||||
runtime: "mastra-agentic",
|
||||
skill: {
|
||||
name: "jyotish-vedic-astrology",
|
||||
loaded: state.jyotishSkillBound,
|
||||
referenceReads: state.skillReferenceReadCount,
|
||||
methodologySections: state.methodologySectionCount,
|
||||
},
|
||||
steps: publicConsultationRuntimeSteps(state),
|
||||
stepBudget: consultationStepBudgetReceipt(state),
|
||||
workflow: state.workflowReceipt ?? { route: job.domain, status: "blocked", preciseTiming: "blocked", missingLayers: [], domains: [job.domain as ConsultationDomain] },
|
||||
techniqueTruth: state.techniqueTruth ?? "unknown",
|
||||
}) as never,
|
||||
onComplete: (text: string) => {
|
||||
output = text;
|
||||
},
|
||||
onError: (error: unknown) => {
|
||||
failure = error instanceof Error ? error.message : String(error);
|
||||
},
|
||||
});
|
||||
// Drain the public event stream; that is what drives the run.
|
||||
await response.text();
|
||||
const usage: Usage = {};
|
||||
for (const item of await Promise.allSettled(usages)) {
|
||||
if (item.status === "fulfilled") addUsage(usage, item.value);
|
||||
}
|
||||
// Assigned inside the stream callbacks, which TypeScript cannot see.
|
||||
const settledOutput = output as string | null;
|
||||
const settledFailure = failure as string | null;
|
||||
return {
|
||||
output: settledOutput ?? "",
|
||||
failure: settledFailure,
|
||||
executed,
|
||||
seconds: (Date.now() - startedAt) / 1000,
|
||||
usage,
|
||||
cardChars: state.evidenceCardChars ?? null,
|
||||
visibleChars: state.modelVisibleChars ?? null,
|
||||
methodologySections: state.methodologySectionCount,
|
||||
lookups: state.evidenceLookupCallCount,
|
||||
finish: state.composeFinishReason ?? state.modelFinishReason ?? null,
|
||||
};
|
||||
}
|
||||
|
||||
// ---- deterministic checks (judgment scoring is done by reading the answers) ----
|
||||
|
||||
const BANNED = [
|
||||
{ id: "retrograde_repeat", pattern: /逆行[^。!?\n]{0,30}(反复|打回来|来回|拉回|折返|回头)/gu },
|
||||
{ id: "repeat_retrograde", pattern: /(反复|打回来|来回)[^。!?\n]{0,12}逆行/gu },
|
||||
];
|
||||
// 他 / 她 when the card says 性别未知: only for the partner (marriage) or children.
|
||||
const PRONOUN = /(?<![其吉])[他她](?![们人])/gu;
|
||||
|
||||
export function deterministicChecks(answer: string, domain: string, mustNot: { claim: string; literals?: string[] }[]) {
|
||||
const context = (index: number) => answer.slice(Math.max(0, index - 14), index + 16).replace(/\s+/g, " ");
|
||||
const banned = BANNED.flatMap((rule) => [...answer.matchAll(rule.pattern)].map((hit) => ({ rule: rule.id, text: hit[0] })));
|
||||
const pronouns = domain === "marriage" || domain === "children"
|
||||
? [...answer.matchAll(PRONOUN)].map((hit) => context(hit.index ?? 0))
|
||||
: [];
|
||||
const literalHits = mustNot.flatMap((item) => (item.literals ?? [])
|
||||
.filter((literal) => answer.includes(literal))
|
||||
.map((literal) => ({ claim: item.claim, literal })));
|
||||
return { banned, pronouns, literalHits };
|
||||
}
|
||||
|
||||
async function pool<T, R>(items: readonly T[], size: number, fn: (item: T) => Promise<R>) {
|
||||
const results: R[] = new Array(items.length);
|
||||
let next = 0;
|
||||
await Promise.all(Array.from({ length: Math.min(size, items.length) }, async () => {
|
||||
while (next < items.length) {
|
||||
const index = next++;
|
||||
results[index] = await fn(items[index]!);
|
||||
}
|
||||
}));
|
||||
return results;
|
||||
}
|
||||
|
||||
async function dumpCards(jobs: readonly Job[], out: string) {
|
||||
for (const job of jobs.filter((item) => item.run === 1)) {
|
||||
const state = createConsultationRuntimeState();
|
||||
const ctx = agentContext(job, state, []);
|
||||
const fake = model({ ...parseArgs([]), model: "card-dump" }, "unused");
|
||||
const agent = getJyotishAgent(fake, ctx);
|
||||
const tools = await agent.listTools();
|
||||
const tool = tools["run-jyotish-consultation"] as unknown as { execute: (input: unknown, context: unknown) => Promise<unknown> };
|
||||
const view = await tool.execute({ question: job.question, domains: [job.domain] }, {});
|
||||
const path = join(out, "cards", `${job.figure.id}-${job.domain}.json`);
|
||||
mkdirSync(dirname(path), { recursive: true });
|
||||
writeFileSync(path, `${JSON.stringify(view, null, 1)}\n`);
|
||||
console.log(`card ${path} (${JSON.stringify(view).length} chars)`);
|
||||
}
|
||||
}
|
||||
|
||||
// The product logs every reasoning delta (consultation-budget.ts); a batch run
|
||||
// would drown its own progress lines in them.
|
||||
const productInfo = console.info.bind(console);
|
||||
console.info = (...items: unknown[]) => {
|
||||
if (typeof items[0] === "string" && items[0].startsWith("[consult-reasoning]")) return;
|
||||
productInfo(...items);
|
||||
};
|
||||
|
||||
type Row = { figure: string; domain: string; run: number; chars: number; seconds: number; checks: ReturnType<typeof deterministicChecks> };
|
||||
|
||||
function writeSummary(out: string, summary: Json & { rows: Row[] }, modelId: string) {
|
||||
const rows = summary.rows;
|
||||
summary.banned_hits = rows.reduce((sum, row) => sum + row.checks.banned.length, 0);
|
||||
summary.pronoun_hits = rows.reduce((sum, row) => sum + row.checks.pronouns.length, 0);
|
||||
summary.must_not_literal_hits = rows.reduce((sum, row) => sum + row.checks.literalHits.length, 0);
|
||||
writeFileSync(join(out, "summary.json"), `${JSON.stringify(summary, null, 1)}\n`);
|
||||
writeFileSync(join(out, "summary.md"), [
|
||||
`# 名人生平回测 · ${modelId}`,
|
||||
"",
|
||||
`开始 ${summary.started_at},${rows.length} 份,墙钟 ${Number(summary.wall_seconds).toFixed(0)} s,失败 ${summary.failures} 份。`,
|
||||
`用量合计:${JSON.stringify(summary.usage)}`,
|
||||
`确定性检查:禁句 ${summary.banned_hits},性别代词(婚恋/子女)${summary.pronoun_hits},must_not 原文 ${summary.must_not_literal_hits}。判断类评分需逐份阅读。`,
|
||||
"",
|
||||
"| 名人 | 领域 | run | 字数 | 秒 | 禁句 | 代词 | must_not |",
|
||||
"| --- | --- | --- | --- | --- | --- | --- | --- |",
|
||||
...rows.map((row) => `| ${row.figure} | ${row.domain} | ${row.run} | ${row.chars} | ${row.seconds.toFixed(0)} | ${row.checks.banned.length} | ${row.checks.pronouns.length} | ${row.checks.literalHits.length} |`),
|
||||
"",
|
||||
"## must_not 原文命中",
|
||||
"",
|
||||
...rows.flatMap((row) => row.checks.literalHits.map((hit) => `- ${row.figure} · ${row.domain} · run${row.run}:「${hit.literal}」(${hit.claim})`)),
|
||||
"",
|
||||
"## 与真实路由的差异",
|
||||
"",
|
||||
...GAPS.map((gap) => `- ${gap}`),
|
||||
"",
|
||||
].join("\n"));
|
||||
}
|
||||
|
||||
function recheck(out: string, rubric: { figures: RubricFigure[] }) {
|
||||
const summary = JSON.parse(readFileSync(join(out, "summary.json"), "utf8")) as Json & { rows: Row[]; model: string };
|
||||
for (const row of summary.rows) {
|
||||
const text = readFileSync(join(out, row.figure, `${row.domain}-run${row.run}.md`), "utf8");
|
||||
const answer = text.slice(text.indexOf("\n---\n\n") + 6).trim();
|
||||
const mustNot = rubric.figures.find((item) => item.id === row.figure)?.domains[row.domain]?.must_not ?? [];
|
||||
row.checks = deterministicChecks(answer === "(无回答)" ? "" : answer, row.domain, mustNot);
|
||||
}
|
||||
writeSummary(out, summary, summary.model);
|
||||
console.log(`rechecked: ${join(out, "summary.md")}`);
|
||||
}
|
||||
|
||||
async function main() {
|
||||
const args = parseArgs(process.argv.slice(2));
|
||||
const golden = JSON.parse(readFileSync(GOLDEN, "utf8")) as Golden;
|
||||
const rubric = JSON.parse(readFileSync(RUBRIC, "utf8")) as { figures: RubricFigure[] };
|
||||
if (args.recheck) {
|
||||
recheck(resolve(args.recheck), rubric);
|
||||
return;
|
||||
}
|
||||
const questions = new Map(golden.routes.map((route) => [route.domain, route.question]));
|
||||
const figures = golden.figures.filter((figure) => !args.figures || args.figures.includes(figure.id));
|
||||
const jobs: Job[] = figures.flatMap((figure) => args.domains.flatMap((domain) => {
|
||||
const question = questions.get(domain);
|
||||
if (!question) throw new Error(`no golden question for domain ${domain}`);
|
||||
return Array.from({ length: args.repeats }, (_, index) => ({ figure, domain, question, run: index + 1 }));
|
||||
}));
|
||||
mkdirSync(args.out, { recursive: true });
|
||||
if (args.dumpCards) {
|
||||
await dumpCards(jobs, args.out);
|
||||
return;
|
||||
}
|
||||
const apiKey = process.env.DS_KEY ?? process.env.BACKTEST_MODEL_KEY;
|
||||
if (!apiKey) throw new Error("Set DS_KEY (or BACKTEST_MODEL_KEY); without a key the backtest is an environment gap.");
|
||||
const resolved = model(args, apiKey);
|
||||
const started = Date.now();
|
||||
const rows = await pool(jobs, args.concurrency, async (job) => {
|
||||
let result: Awaited<ReturnType<typeof runJob>>;
|
||||
try {
|
||||
result = await runJob(job, resolved);
|
||||
} catch (error) {
|
||||
result = { output: "", failure: error instanceof Error ? error.message : String(error), executed: [], seconds: 0, usage: {}, cardChars: null, visibleChars: null, methodologySections: 0, lookups: 0, finish: null };
|
||||
}
|
||||
const mustNot = rubric.figures.find((item) => item.id === job.figure.id)?.domains[job.domain]?.must_not ?? [];
|
||||
const checks = deterministicChecks(result.output, job.domain, mustNot);
|
||||
const file = join(args.out, job.figure.id, `${job.domain}-run${job.run}.md`);
|
||||
mkdirSync(dirname(file), { recursive: true });
|
||||
writeFileSync(file, [
|
||||
`# ${job.figure.label} · ${job.domain} · run ${job.run}`,
|
||||
"",
|
||||
`- 问题:${job.question}`,
|
||||
`- 模型:${args.model};用时 ${result.seconds.toFixed(1)} s;结束:${result.finish ?? "?"};查卡 ${result.lookups} 次;清单段 ${result.methodologySections}`,
|
||||
`- 执行领域:${result.executed.join(", ") || "无"};卡 ${result.cardChars ?? "?"} 字符,工具结果 ${result.visibleChars ?? "?"} 字符`,
|
||||
`- 用量:${JSON.stringify(result.usage)}`,
|
||||
`- 确定性检查:禁句 ${checks.banned.length},性别代词 ${checks.pronouns.length},must_not 原文 ${checks.literalHits.length}`,
|
||||
...(result.failure ? [`- 失败:${result.failure}`] : []),
|
||||
...checks.banned.map((hit) => ` - 禁句 ${hit.rule}:「${hit.text}」`),
|
||||
...checks.pronouns.map((hit) => ` - 代词:「${hit}」`),
|
||||
...checks.literalHits.map((hit) => ` - must_not:「${hit.literal}」(${hit.claim})`),
|
||||
"",
|
||||
"---",
|
||||
"",
|
||||
result.output || "(无回答)",
|
||||
"",
|
||||
].join("\n"));
|
||||
console.log(`${job.figure.id} ${job.domain} run${job.run}: ${result.output.length} chars, ${result.seconds.toFixed(0)} s${result.failure ? `, FAILED ${result.failure}` : ""}`);
|
||||
return { figure: job.figure.id, domain: job.domain, run: job.run, chars: result.output.length, ...result, output: undefined, checks };
|
||||
});
|
||||
const total: Usage = {};
|
||||
for (const row of rows) addUsage(total, row.usage);
|
||||
const summary = {
|
||||
model: args.model,
|
||||
started_at: new Date(started).toISOString(),
|
||||
wall_seconds: (Date.now() - started) / 1000,
|
||||
jobs: rows.length,
|
||||
failures: rows.filter((row) => row.failure || row.chars === 0).length,
|
||||
usage: total,
|
||||
gaps: GAPS,
|
||||
rows,
|
||||
};
|
||||
writeSummary(args.out, summary as unknown as Json & { rows: Row[] }, args.model);
|
||||
console.log(`summary: ${join(args.out, "summary.md")}`);
|
||||
}
|
||||
|
||||
await main();
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,179 @@
|
||||
"""Capture real engine consultation responses for the biography backtest.
|
||||
|
||||
Same path as `capture_consult_evidence_card_golden.py` (same handler, same
|
||||
request body, same research reference date, raman ayanamsa, mean nodes, the
|
||||
external VedAstro stand-in), extended to nine public Rodden AA charts and to the
|
||||
four engine routes the backtest asks about: `family` (the engine route the
|
||||
product's `parents` domain runs as), `marriage`, `health` and `career`. Each
|
||||
response is trimmed by key only with the same `trim()` as the evidence-card
|
||||
golden: every kept value is the engine's own value, unchanged.
|
||||
|
||||
The chart block (`workflow.chart`) is almost route-independent: the keys all
|
||||
four routes agree on are stored once per figure as `shared_chart`, and the
|
||||
keys that differ (today `modules.yogas` and `modules.chara_dasha`) stay with
|
||||
each route under `chart` / `chart_modules`. The TypeScript reader overlays
|
||||
them back (`frontend/scripts/research/consult-biography-backtest.mts`), so
|
||||
every route's chart is exactly what the engine returned.
|
||||
|
||||
JYOTISH_API_CHART_CACHE_TTL_SECONDS=0 PYTHONHASHSEED=0 \
|
||||
python3 scripts/research/capture_consult_biography_backtest_golden.py
|
||||
|
||||
Birth data only from the repository's public case libraries (Astro-Databank
|
||||
AA); no user chart is ever read.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path[:0] = [str(ROOT), str(ROOT / "scripts"), str(ROOT / "scripts" / "research")]
|
||||
|
||||
from consult_evidence_card_lib import AYANAMSA, NODE_MODE, REFERENCE_DATE # noqa: E402
|
||||
import consult_evidence_card_run as runner # noqa: E402
|
||||
from capture_consult_evidence_card_golden import trim # noqa: E402
|
||||
|
||||
OUT = ROOT / "frontend" / "tests" / "fixtures" / "consult-biography-backtest-golden.json"
|
||||
CASES = "references/real_case_calibration"
|
||||
|
||||
# id, label, case library, case id. The rubric
|
||||
# (docs/research/consult_biography_backtest_rubric.json) carries why each is here.
|
||||
FIGURES = (
|
||||
("steve_jobs", "Steve Jobs", f"{CASES}/minute_rectification_development_v1.json", "steve_jobs_1955_development"),
|
||||
("barack_obama", "Barack Obama", f"{CASES}/minute_rectification_holdout_v4.json", "barack_obama_1961_aa_v4_holdout"),
|
||||
("elizabeth_taylor", "Elizabeth Taylor", f"{CASES}/minute_rectification_holdout_v4.json", "elizabeth_taylor_1932_aa_v4_holdout"),
|
||||
("marilyn_monroe", "Marilyn Monroe", f"{CASES}/minute_rectification_holdout_v5.json", "marilyn_monroe_1926_aa_v5_holdout"),
|
||||
("judy_garland", "Judy Garland", f"{CASES}/minute_rectification_holdout_v5.json", "judy_garland_1922_aa_v5_holdout"),
|
||||
("edith_piaf", "Édith Piaf", f"{CASES}/minute_rectification_holdout_v5.json", "edith_piaf_1915_aa_v5_holdout"),
|
||||
("frida_kahlo", "Frida Kahlo", f"{CASES}/minute_rectification_holdout_v5.json", "frida_kahlo_1907_aa_v5_holdout"),
|
||||
("george_w_bush", "George W. Bush", f"{CASES}/minute_rectification_holdout_v5.json", "george_w_bush_1946_aa_v5_holdout"),
|
||||
("zinedine_zidane", "Zinedine Zidane", f"{CASES}/minute_rectification_holdout_v5.json", "zinedine_zidane_1972_aa_v5_holdout"),
|
||||
)
|
||||
|
||||
# Product domain -> engine route and the question the backtest asks. The
|
||||
# product sends a card-only domain (parents) as its engine route's contract
|
||||
# (frontend/src/lib/consultation-workflow-request.ts).
|
||||
ROUTES = (
|
||||
{"domain": "parents", "engine_route": "family", "question": "我和父母关系如何,他们怎么对待我?"},
|
||||
{"domain": "marriage", "engine_route": "marriage", "question": "我的婚姻和感情会是什么样?"},
|
||||
{"domain": "health", "engine_route": "health", "question": "我的身体底子怎么样,健康上要注意什么?"},
|
||||
{"domain": "career", "engine_route": "career", "question": "我的事业会往什么方向走,能做到什么程度?"},
|
||||
)
|
||||
|
||||
|
||||
def load_chart(spec: tuple[str, str, str, str]) -> dict[str, Any]:
|
||||
figure_id, label, path, case_id = spec
|
||||
payload = json.loads((ROOT / path).read_text(encoding="utf-8"))
|
||||
case = next(item for item in payload["cases"] if item["case_id"] == case_id)
|
||||
birth = case["birth"]
|
||||
year, month, day = (int(part) for part in birth["date"].split("-"))
|
||||
hour, minute = (int(part) for part in birth["time"].split(":")[:2])
|
||||
source = birth.get("source") or {}
|
||||
return {
|
||||
"id": figure_id,
|
||||
"label": label,
|
||||
"source": path,
|
||||
"case_id": case_id,
|
||||
"rodden_rating": source.get("rodden_rating"),
|
||||
"birth_source_url": source.get("url"),
|
||||
# Same body as consult_evidence_card_lib.load_public_charts.
|
||||
"body": {
|
||||
"year": year,
|
||||
"month": month,
|
||||
"day": day,
|
||||
"hour": hour,
|
||||
"minute": minute,
|
||||
"second": 0,
|
||||
"lat": birth["latitude"],
|
||||
"lon": birth["longitude"],
|
||||
"tz": birth["timezone_offset"],
|
||||
"city": birth.get("place") or label,
|
||||
"ayanamsa": AYANAMSA,
|
||||
"node_mode": NODE_MODE,
|
||||
"today": REFERENCE_DATE,
|
||||
"entry_mode": "direct_chart",
|
||||
"defer_optional_external_evidence": True,
|
||||
"declared_accuracy": "minute",
|
||||
"birth_time_accuracy": "confirmed",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def split_shared_chart(routes: dict[str, dict[str, Any]]) -> dict[str, Any]:
|
||||
"""Lift the chart keys every route agrees on into one copy.
|
||||
|
||||
A key (or a `modules` key) whose value differs between routes stays with
|
||||
each route under `chart` (or `chart_modules`); the reader overlays it on
|
||||
the shared copy, so every route's chart comes back exactly as captured.
|
||||
"""
|
||||
charts = [value.pop("chart") for value in routes.values()]
|
||||
shared: dict[str, Any] = {}
|
||||
keys = {key for chart in charts for key in chart if key != "modules"}
|
||||
for key in sorted(keys):
|
||||
values = [chart.get(key) for chart in charts]
|
||||
if all(key in chart for chart in charts) and all(item == values[0] for item in values):
|
||||
shared[key] = values[0]
|
||||
else:
|
||||
for value, chart in zip(routes.values(), charts):
|
||||
if key in chart:
|
||||
value.setdefault("chart", {})[key] = chart[key]
|
||||
modules = [chart.get("modules") or {} for chart in charts]
|
||||
shared["modules"] = {}
|
||||
for key in sorted({key for item in modules for key in item}):
|
||||
values = [item.get(key) for item in modules]
|
||||
if all(key in item for item in modules) and all(entry == values[0] for entry in values):
|
||||
shared["modules"][key] = values[0]
|
||||
else:
|
||||
for value, item in zip(routes.values(), modules):
|
||||
if key in item:
|
||||
value.setdefault("chart_modules", {})[key] = item[key]
|
||||
return shared
|
||||
|
||||
|
||||
def main() -> int:
|
||||
if os.environ.get("PYTHONHASHSEED") != "0":
|
||||
raise SystemExit("Set PYTHONHASHSEED=0 before starting this process (ERR-111).")
|
||||
only = {item for item in os.environ.get("BACKTEST_FIGURES", "").split(",") if item}
|
||||
runner._block_external_vedastro()
|
||||
figures = []
|
||||
for spec in FIGURES:
|
||||
if only and spec[0] not in only:
|
||||
continue
|
||||
chart = load_chart(spec)
|
||||
routes: dict[str, dict[str, Any]] = {}
|
||||
for route in ROUTES:
|
||||
question = {"id": route["domain"], "domain": route["engine_route"], "question": route["question"]}
|
||||
routes[route["domain"]] = trim(runner._run_workflow(runner._workflow_body(chart, question)))
|
||||
print(f"{chart['id']} {route['domain']} ok", flush=True)
|
||||
shared = split_shared_chart(routes)
|
||||
figures.append({
|
||||
"id": chart["id"],
|
||||
"label": chart["label"],
|
||||
"source": chart["source"],
|
||||
"case_id": chart["case_id"],
|
||||
"rodden_rating": chart["rodden_rating"],
|
||||
"birth_source_url": chart["birth_source_url"],
|
||||
"shared_chart": shared,
|
||||
"routes": routes,
|
||||
})
|
||||
payload = {
|
||||
"source": "scripts/research/capture_consult_biography_backtest_golden.py",
|
||||
"note": "Real engine consultation_workflow responses for nine public AA charts x four routes, trimmed by key only (same trim as consult-evidence-card-golden.json).",
|
||||
"reference_date": REFERENCE_DATE,
|
||||
"ayanamsa": AYANAMSA,
|
||||
"node_mode": NODE_MODE,
|
||||
"routes": list(ROUTES),
|
||||
"external_vedastro": "not_called",
|
||||
"figures": figures,
|
||||
}
|
||||
OUT.write_text(json.dumps(payload, ensure_ascii=False, separators=(",", ":")) + "\n", encoding="utf-8")
|
||||
print(f"wrote {OUT} ({OUT.stat().st_size} bytes)")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Reference in New Issue
Block a user