A 👎 now saves, server side, the rated turn plus the context window the model read for it (reconstructed from the stored session with the same consultationHistoryWindow the consult route uses), model and run facts. 👍 is only counted. Switching to 👍 or clearing deletes the snapshot. Bodies are blanked after 90 days; the row cascades on session delete and on account deletion. - Optional 不满意原因 panel under the answer after a 👎 (five reasons, 200-char note, "会把这一轮对话发给我们排查"). - Admin 对话质量记录: 👍/👎 stats by day and model, list without text, audited snapshot open, 处理状态 + note (support.quality.read/write). - Privacy draft: what a 👎 keeps, why, 90 days, deletion. - Migration 20260930050000 is add-only; set_reply_rating() replaced with the same signature. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N4f2nya58RoRu4yEmJgRGE
43 lines
4.4 KiB
Markdown
43 lines
4.4 KiB
Markdown
# PROGRESS · 点踩收集 + 对话质量记录 · 2026-09-30
|
||
|
||
> 产品 2026-09-30「点赞点踩后台需要收集,点踩要收集这轮对话以便排查」,选定:快照 = 这一轮 + 模型当时看到的上下文;保留 90 天。执行:Claude fork 子代理 E。分支 `codex/reply-quality-20260930`,基于 `dbac4196`(合规轮 A–D 之后)。术语按 `CONTEXT.md`:回复评价、不满意原因、对话质量记录、故障上下文快照、处理状态。
|
||
|
||
## 做了什么
|
||
|
||
| 项 | 复用 | 新增 |
|
||
|---|---|---|
|
||
| 👍 / 👎 保存 | C 轮的 `reply_ratings`、`/api/reply-ratings`、`persistReplyRating` | `set_reply_rating()` 同签名替换:改成 👍 或取消时删掉快照 |
|
||
| 故障上下文快照 | `consultationHistoryWindow` / `parseSessionContextSummary`(咨询路由给模型的同一窗口)、`resolveSessionLanguageModel().contextWindow`、`usage_ledger` | `lib/reply-quality-snapshot.ts`(纯函数)、`lib/reply-quality-capture.ts`(服务端读会话 → 建快照 → `save_reply_quality_snapshot()`);表 `reply_quality_snapshots` |
|
||
| 不满意原因 | `ChatMessageActions` 行 | `components/reply-down-reason.tsx`(`MessageEntry` 自持状态,Home 不变)、`PUT /api/reply-ratings/reason`、`set_reply_quality_reason()` |
|
||
| 保留期 | 注销 worker 的 unref 定时器模式 | `reply-quality-retention-worker.ts`:每 6 小时调 `expire_reply_quality_snapshot_bodies(90)`,日志只写条数 |
|
||
| 后台 | `ResourceTable`、RBAC、`audit.admin_audit_logs`、C 轮反馈页结构 | 权限 `support.quality.read/write`(owner / operations / support 读写,auditor 只读);`/admin/reply-quality`、`/api/admin/reply-quality`(列表不带正文)、`/[id]`(查看写审计 `reply_quality.open`;改状态写 `reply_quality.update`)、`/stats`(按天 / 按模型 👍👎) |
|
||
| 隐私政策草稿 | `legal-documents.ts` | 收集、用途、保存期限三处各一句 |
|
||
|
||
## 决定
|
||
|
||
- **上下文是重建的,不是生成时存的。** 👎 时服务器按会话里存的消息,用咨询路由同一个函数重建那一轮模型读到的窗口:只用该轮之前写好的会话摘要(`throughMessageIndex < 提问位置`),预算按会话当前模型的上下文窗口。偏差:该轮之后换过模型时预算可能不同;星盘数据块与系统提示不在快照里(运行事实里有 requestId、模型、token、耗时,可对 `usage_ledger` 查)。不在生成时存,是因为同日 D 轮正在改咨询路由,另存一份会改动热路径。
|
||
- 客户端只发位置与哈希,不发文本;服务器只在「该位置是助手回答且哈希一致」时存,重新生成过的回答不存。
|
||
- 校正会话的上下文另有组装方式,快照只取提问前最近 6 条并标 `rectification_recent_turns`。校正页的 👍 / 👎 本身没持久化,这轮不接。
|
||
- 库里没有「Agent 执行故障」记录(咨询请求页只有请求生命周期),所以对话质量记录目前只有 👎 一种来源。
|
||
- 原因提交会先等本条评价保存完成,避免快照还没建就写原因(`reply_quality_not_found`)。
|
||
|
||
## 改动的既有断言
|
||
|
||
| 文件 | 原值 | 新值 | 原因 |
|
||
|---|---|---|---|
|
||
| `tests/feedback-complaints-20260930.test.tsx` | `persistReplyRating(feedbackKey, toggleChatMessageFeedback(messageFeedback[feedbackKey], requested), message.text)` 一行 | `const next = toggleChatMessageFeedback(...)` 后 `persistReplyRating(feedbackKey, next, message.text)` | 同一个切换结果还要用来打开原因面板;保存的值不变 |
|
||
|
||
## 验证(Linux,Node 22;基线 = `dbac4196` 全量)
|
||
|
||
| 项 | 结果 |
|
||
|---|---|
|
||
| `tsc --noEmit` | 0 错 |
|
||
| `npm run lint` | 0 error,126 warning(同基线) |
|
||
| `npm test` 全量 | 基线 4,447 / fail 24 → 本轮 4,460 / fail 24,cancelled 0;失败名单与基线逐条相同(无 Docker 的数据库 / 部署套件);新增 13 条,消失 0 条 |
|
||
| 新测试 `reply-quality-20260930.test.tsx` | 13 / 13 |
|
||
| `home-shell-growth-contract` | 通过(Home 的 useState / useRef 数不变) |
|
||
| `next build` | exit 0;`/` 仍 `○ Static`;新增 `/admin/reply-quality` 与三个后台接口 |
|
||
| 首屏 gzip-9 | 651,976 B → 652,921 B(+0.14%,原因面板组件进了对话转录) |
|
||
| 隐私标记 `tests/test_repo_privacy_markers.py` | 通过 |
|
||
| 真实数据库 / 真机 | 未验证,见 `BLOCKED.md`「对话质量记录」 |
|