fix: preserve recovery helper traversal
Independent Staging Quality Gate / validate (push) Successful in 12m19s
Independent Staging Quality Gate / publish (push) Successful in 2m14s

This commit is contained in:
Jesse_Chen
2026-08-16 03:30:03 +08:00
parent 29a7295667
commit 9f31de7b4a
5 changed files with 9 additions and 7 deletions
+4 -4
View File
@@ -3415,10 +3415,10 @@
- 首次发现:2026-08-15
- 最近更新:2026-08-15
- 影响面:`Create Production Recovery Point`、生产 schema migration 前置恢复门禁,以及后续 production deploy。
- 用户现象:`main``staging` 已同步且 release gate 成功,但生产恢复 workflow Run 1865 在创建备份前失败,日志为 `/opt/jyotisha-production/.state/mutation.lock: Permission denied`。首次修复后的 Run 1869 仍在 `pg_dump` 前失败,准确日志为 `sudo: a password is required`;迁移和部署因此持续阻断,生产运行版本未改变。
- 根因:生产 bootstrap 遗留的 `.state` 路径或既有 lock 仍为非 `deploy` 所有。原恢复脚本只执行 `install -d -m 700`;对已经存在的目录该命令不会恢复所有权,随后由 `deploy` 打开共享锁即被内核拒绝。首次修复又错误假设主机已为 `deploy` 配置独立的免密 `chown`,但真实生产 sudo 边界只允许已审查的 Docker 命令,因此 `sudo -n chown` 立即失败。次失败都发生在 `pg_dump` 前,没有生成可用恢复证明,也没有执行 schema migration。
- 修复:恢复脚本先对 `.state``backups` 和既有 `mutation.lock` 做类型与非符号链接校验,取得当前运行 PostgreSQL 容器的不可变本地 image ID,再复用既有 `sudo -n docker` 边界启动一次性所有权修复容器:`--pull never`、无网络、只读根文件系统、`no-new-privileges`、删除全部 capability 后只保留 `CHOWN`,并且仅绑定 `.state``backups`。helper 恢复两个目录 mount point 和既有普通 lock 到当前 `deploy` UID/GID;不递归改动历史备份、不删除或替换 lock inode,随后仍用同一个 `flock -n` fail-closed 获取共享 host lock。同步更新生产 runbook 和静态安全契约测试,明确不得再扩大主机 sudoers。
- 用户现象:`main``staging` 已同步且 release gate 成功,但生产恢复 workflow Run 1865 在创建备份前失败,日志为 `/opt/jyotisha-production/.state/mutation.lock: Permission denied`。首次修复后的 Run 1869 仍在 `pg_dump` 前失败,准确日志为 `sudo: a password is required`。切换到受限 helper 后,Run 1873 又在同一备份前阶段报告 `chown: /opt/jyotisha-production/.state/mutation.lock: Permission denied`;迁移和部署因此持续阻断,生产运行版本未改变。
- 根因:生产 bootstrap 遗留的 `.state` 路径或既有 lock 仍为非 `deploy` 所有。原恢复脚本只执行 `install -d -m 700`;对已经存在的目录该命令不会恢复所有权,随后由 `deploy` 打开共享锁即被内核拒绝。首次修复又错误假设主机已为 `deploy` 配置独立的免密 `chown`,但真实生产 sudo 边界只允许已审查的 Docker 命令,因此 `sudo -n chown` 立即失败。第二次修复的 helper 只保留 `CHOWN` capability,却按“父目录在前、lock 子文件在后”的顺序处理;当 `.state` 先被改成 `deploy:deploy 0700` 后,容器内 UID 0 已不再是目录 owner,且没有 `DAC_OVERRIDE`,所以无法继续遍历到 lock。三次失败都发生在 `pg_dump` 前,没有生成可用恢复证明,也没有执行 schema migration。
- 修复:恢复脚本先对 `.state``backups` 和既有 `mutation.lock` 做类型与非符号链接校验,取得当前运行 PostgreSQL 容器的不可变本地 image ID,再复用既有 `sudo -n docker` 边界启动一次性所有权修复容器:`--pull never`、无网络、只读根文件系统、`no-new-privileges`、删除全部 capability 后只保留 `CHOWN`,并且仅绑定 `.state``backups`。helper 按 child-first 顺序先恢复既有普通 lock,再恢复两个目录 mount point 到当前 `deploy` UID/GID,因此不需要增加 `DAC_OVERRIDE`;不递归改动历史备份、不删除或替换 lock inode,随后仍用同一个 `flock -n` fail-closed 获取共享 host lock。同步更新生产 runbook 和静态安全契约测试,明确不得再扩大主机 sudoers,并锁定子文件必须先于 mode-`0700` 父目录处理
- 验证:本地 `bash -n`、YAML parse、聚焦 workflow contract 与 `git diff --check` 必须通过;远端必须重新完成 staging quality/deploy、exact-SHA release gate、真实 production dump + disposable restore + off-site artifact,再允许 migration/deploy。
- 防复发:生产私有状态目录必须保持 `deploy:deploy 0700`,共享 lock 必须是普通非 symlink 文件且不可通过删除重建来“修复”;任何恢复流程失败都不得手填 `restore_verified=true` 或跳过恢复门禁。
- 相关记录:生产迁移 runbook、Run 1865、Run 1869
- 相关记录:生产迁移 runbook、Run 1865、Run 1869、Run 1873
- 修复版本:待提交(精确 SHA 以重新发布后的远端分支与 production health 为准)
@@ -90,7 +90,7 @@ Use distinct production credentials for PostgreSQL roles, Better Auth, Resend, b
## Database migration engineering gate
Before importing data, dispatch Gitea Actions → `Create Production Recovery Point` for the exact accepted release SHA. It validates the private production `.state` and `backups` paths, obtains the immutable image ID of the already-running PostgreSQL container, and reuses the existing passwordless Docker boundary to run that local image with `--pull never`, no network, a read-only root filesystem, `no-new-privileges`, and only `CHOWN` retained. With only `.state` and `backups` bind-mounted, it restores the two directory mount points and an existing regular shared lock to documented `deploy` ownership when bootstrap ownership has drifted. It then creates an AES-256-CBC/PBKDF2 encrypted custom-format dump on the production host, restores it into a disposable database, verifies non-sensitive table counts, removes only that disposable database, and uploads the encrypted dump plus verification manifest as a 30-day Gitea Actions artifact. It must preserve the existing lock inode, avoid recursively changing historical backups, and fail closed on symlinks, non-regular lock paths, or an unexpected image identity. The workflow prints the non-sensitive `recovery_reference` and exact UTC `recovery_created_at`; use those values only after the artifact upload succeeds.
Before importing data, dispatch Gitea Actions → `Create Production Recovery Point` for the exact accepted release SHA. It validates the private production `.state` and `backups` paths, obtains the immutable image ID of the already-running PostgreSQL container, and reuses the existing passwordless Docker boundary to run that local image with `--pull never`, no network, a read-only root filesystem, `no-new-privileges`, and only `CHOWN` retained. With only `.state` and `backups` bind-mounted, it restores an existing regular shared lock before changing its mode-`0700` parent directory, then restores the two directory mount points to documented `deploy` ownership when bootstrap ownership has drifted. The child-first order avoids adding `DAC_OVERRIDE` while retaining access to the lock path. It then creates an AES-256-CBC/PBKDF2 encrypted custom-format dump on the production host, restores it into a disposable database, verifies non-sensitive table counts, removes only that disposable database, and uploads the encrypted dump plus verification manifest as a 30-day Gitea Actions artifact. It must preserve the existing lock inode, avoid recursively changing historical backups, and fail closed on symlinks, non-regular lock paths, or an unexpected image identity. The workflow prints the non-sensitive `recovery_reference` and exact UTC `recovery_created_at`; use those values only after the artifact upload succeeds.
Then dispatch Gitea Actions → `Migrate Production Database` for the same exact accepted release SHA. The workflow requires `main == staging == deploy_sha`, the same successful staging backend gate, the same manual release gate, and the public staging `/api/health` identity for that SHA. It also requires the recovery workflow's non-sensitive reference, exact UTC creation time, and `restore_verified=true`; the recovery point must be no more than 24 hours old and must already have passed the restore verification. It verifies the current production revision, obtains the gate-attested immutable Web image, runs the schema checker, applies only pending application schema migrations, and requires the checker to converge afterward.