选项
首页首页 Skill 文档 shipping-artifacts

shipping-artifacts

phuryn/pm-skills phuryn/pm-skills

记录 AI 构建的应用程序的架构、权限、机密信息和测试覆盖率图,以便在发布前对其进行审查。

...展开全部
0
更新时间 2026-09-29

传递“工件”:让AI生成的代码可审查的文档

目的

AI 代理虽然能快速编写代码,却无法留下持久的意图记录——即系统应执行什么操作、谁被允许执行什么操作、机密信息存储在哪里,以及哪些规则已实际经过验证。 若缺乏这些记录,任何人类(乃至任何审计代理)都无法判断代码是否安全可发布。本指南定义了一套能恢复代码可审查性的核心文档。

这些文档位于/documentation/目录下,专为两类读者编写:人类审查员和下一代 AI 编码代理。它们构成了后续每次审计中“预期状态”的部分——安全或性能审查的有效性,取决于其能否将代码与这些预期意图进行对比。

该文档集的组织方式

该文档集并非固定列表——它由一小部分核心文档,以及仅在具备相应能力时才添加的条件性文档组成。

  • 核心文档——每个可审查的应用都具备这些要素,因此必须始终生成它们。
  • 条件性文档——仅当应用实际具备相应功能时才包含。若不具备,请在architecture.md中写上一行说明(例如“无定时任务——无cron.md.”),而非创建一个空文档。 可审查性源于真实的映射,而“我们不做 X”也是映射的一部分。
  • 大多数文档是由/document-app 通过代码反向工程生成的。唯一的例外是tests.md,它由 /derive-tests 根据其他文档衍生而来——它是验证地图,而非子系统的描述。

对当前状态要毫不掩饰地如实反映,但不必过度担忧。我们的任务是绘制准确的地图,而非出具“健康证明”。每份文档都简短精炼,大量采用表格和项目符号,并省略通用的理论阐述。

核心文档

每个条目包含:文件名 · 一行目的说明 · 必须涵盖的内容 · 审阅者如何使用它。

  1. architecture.md— 系统是什么以及其整体架构。

    • 必须涵盖:产品概述 + 关键假设;技术栈;身份验证/会话/声明的端到端流转方式; 信任边界(例如服务角色与客户端);一份简短的“已知风险/假设”列表(每项内容均需附有其在代码中的具体位置,而非泛泛的检查清单);一个包含所有其他文档的“相关文档”索引。
    • 审阅者使用方式:作为根文档——所有其他文档均由此处交叉引用。
  2. flows.md—— 实际行使权限并产生副作用的用户旅程。

    • 必须涵盖:每个承重流程的“参与者 + 先决条件 + 成功结果”;UI → 服务器 → 数据 → 任务 → 提供商 → 代理之间的分步序列;每个受保护步骤的授权检查(涉及哪些声明/角色/范围,针对哪些资源,以及预期的拒绝情况);信任边界的跨越(浏览器→服务器、服务器→提供商、任务→应用、代理→工具、webhook→应用);每个步骤引发的状态变化和副作用(写入操作、邮件入队、任务触发、出站调用)。
    • 供审阅者使用:静态的permissions.md矩阵无法展示的运行时视图——授权在何处 以何种顺序被强制执行,以及在何处可以跳过。
    • 反PRD规则:不涉及权限、数据完整性、外部副作用、资金、隐私或运营安全的流程不应出现在此处。这是一份安全/运维地图,而非功能规格说明书。
  3. permissions.md— 谁被允许执行什么操作。

    • 必须涵盖:角色/声明;作用域的来源(令牌 vs. 数据库);资源 × 操作 × 角色的矩阵;哪些表具有行级安全,哪些依赖于代码强制检查。
    • 审阅者用途:这是访问控制审计与代码进行对比的基准。flows.md展示了动态运行情况;此处则是静态参考。
  4. variables.md— 配置和密钥,与风险相关联。

    • 必须包含:名称 · 使用方 · 作用域(服务器/客户端) · 来源 · 轮换周期 · 风险的表格;明确确认客户端未打包任何机密信息;上线前检查清单。
    • 审核人员用途:机密/个人身份信息(PII)泄露风险面,以及事件响应期间的轮换计划。
  5. tests.md— 验证映射:哪些已记录的规则实际经过了检查,哪些仅为建议,哪些则未被任何机制检查。

    • 必须以三个清晰分隔的部分呈现,以确保地图不会显示虚假的“绿色”状态:
      • 现有覆盖范围—当前仓库中已有的测试,每项测试均与对应的规则绑定(确保地图反映实际情况,而非愿望清单)。
      • 拟议测试— 尚未编写的推荐用例,按测试类型标记(自动化单元/集成测试 · 受控生产环境测试 · 人工审查)。
      • 缺口— 完全未经过验证的已记录规则,按违反这些规则可能暴露的问题严重程度排序。
    • 每行包含:用例 → 规则 → 预期行为(包括拒绝/负面情况) → 证据来源(文档 + 代码) → 状态(现有 / 拟议 / 无)。同时注明哪些检查是持续集成(CI)所必需的,以及哪些检查是合并到主分支的门槛。
    • 审阅者用途:这是“已记录 == 已实现”原则的实践体现——它显示其他文档中声称的每条规则,目前是否已被测试实际验证、仅为提案,还是尚未经过验证。
    • 由/derive-tests(而非/document-app)生成,因为它是从其他文档和现有测试套件推导而来的,而非直接读取自子系统。

条件性文档(仅在该功能存在时包含)

  1. emails.md— 系统发送的每条通知。仅当应用程序发送事务性或自动化电子邮件时才包含。

    • 必须记录:队列 → 处理器 → 提供者的路径;模板及其接受的变量;重试/退避行为;发送失败时应检查的位置。
    • 审阅者用途:发现未经验证的模板输入以及个人身份信息(PII)泄露风险点。
  2. cron.md— 所有定时任务及其安全运行方式。仅在存在定时或后台任务时包含。

    • 必须涵盖:清单表(任务 → 计划 → 函数 → 密钥 → 限制 → 重试);每个任务如何保持幂等性;内部调用如何进行身份验证;以及查看最近运行记录的位置。
    • 审核人员用途:查找可伪造的触发条件和无限制的后台任务。
  3. seo.md— 单页应用如何处理 SEO 和社交媒体预览。仅在存在公开/可索引或面向爬虫的路由时包含此文档。

    • 必须涵盖:预览方案(静态元数据/预渲染/边缘 HTML);路由 → 需 SEO → 仅公开数据表;动态元数据的净化方式;机器人与人类的路由区分。
    • 审核人员用途:检测“仅限公共数据”规则的违规情况以及机器人路由上的元数据注入。
  4. automation.md— 嵌入式代理及其他自动化路径。仅当应用嵌入了 AI 代理、LLM 工作流、工具调用、Webhook 或外部自动化时才需包含。

    • 必须针对每项自动化/代理记录:触发条件 + 所有者 + 是否自动运行或仅在批准后运行; 其可能读取的输入以及可调用的具体工具/API(工具接口本身即为硬性防护措施);控制点所在位置(提示词)与非提示词的硬性防护措施;返回给应用的输出契约(模式、验证、错误处理);应用程序拥有的副作用与代理拥有的建议;以及各项控制措施——审批门槛、审计/时间线日志记录、速率限制、重试机制、紧急关闭开关。
    • 评审用途:使隐藏的自动化路径可见,并划清代理提出的建议与应用程序强制执行的内容之间的界限——这是现代基于 AI 构建的应用程序中风险最高的环节。

注释

  • 每份生成的文档都会在architecture.md文件的“相关文档”部分中添加对自身的引用,从而确保该文档集可被发现。
  • 跳过任何不适用的条件文档,并用一句话说明情况,而非编造内容。
  • 请勿在这些文档中包含示例和现成模板——它们描述的是本系统,而非通用方法。
  • 代理操作上下文文件(CLAUDE.md/AGENTS.md)是另一种产物——其指令源自这些文档,而非系统文档。该文件由/ship-check 在交接步骤中生成,而非在此处生成。
  • tests.md由/derive-tests 生成;其余文件由/document-app 生成。
  • 请勿包含“更新日期”行;文件的历史记录才是权威来源。
在 GitHub 上查看
---
name: shipping-artifacts
description: Documents AI-built apps with architecture, permissions, secrets, and test coverage maps to make them reviewable before shipping.
---

# Shipping Artifacts: The Docs That Make AI-Built Code Reviewable

## Purpose

AI agents write code fast, but they leave no durable record of *intent* — what the system is supposed to do, who is allowed to do what, where the secrets live, which rules are actually verified. Without that record, no human (and no auditing agent) can tell whether the code is safe to ship. This skill defines the small set of documents that restore reviewability.

These docs live in `/documentation/` and are written for two readers: a human reviewer and the next AI coding agent. They are the **intended-state** half of every later audit — a security or performance review is only as good as the intent it can compare the code against.

## How the set is organized

The set is **not** a fixed list — it is a small **core** plus **conditional** docs you add only when the capability exists.

- **Core docs** — every reviewable app has these surfaces, so always produce them.
- **Conditional docs** — include one only if the app actually has that capability. If it doesn't, write a single line in `architecture.md` ("No scheduled work — no `cron.md`.") rather than inventing an empty document. Reviewability comes from an honest map, and "we don't do X" is part of the map.
- Most docs are reverse-engineered from code by `/document-app`. The one exception is `tests.md`, which is *derived from the other docs* by `/derive-tests` — it is the verification map, not a description of a subsystem.

Be brutally honest about the current state without being paranoid. The job is an accurate map, not a clean bill of health. Each doc is short, table-and-bullet heavy, and skips generic theory.

## Core documents

Each entry: file · one-line purpose · what it must capture · how a reviewer uses it.

1. **`architecture.md`** — what the system is and how it hangs together.
   - Must capture: product overview + key assumptions; tech stack; how auth/sessions/claims flow end to end; the trust boundaries (e.g. service-role vs. client); a short **Known risks / assumptions** list (each entry backed by where it shows up in the code, not a generic checklist); a "Related Documents" index of every other doc produced.
   - Reviewer use: the root document — everything else is cross-referenced from here.

2. **`flows.md`** — the journeys where permissions and side effects are actually exercised.
   - Must capture: each load-bearing flow as actor + precondition + success outcome; the step-by-step sequence across UI → server → data → jobs → providers → agents; the **authz check at each protected step** (which claim/role/scope, on which resource, and the expected *deny* case); the **trust-boundary crossings** (browser→server, server→provider, job→app, agent→tool, webhook→app); the state changes and side effects each step causes (writes, emails queued, jobs triggered, outbound calls).
   - Reviewer use: the runtime view a static `permissions.md` matrix can't show — *where* and *in what order* authorization is enforced, and where it can be skipped.
   - **Anti-PRD rule:** a flow that doesn't touch permissions, data integrity, external side effects, money, privacy, or operational safety does not belong here. This is a security/operations map, not a feature spec.

3. **`permissions.md`** — who is allowed to do what.
   - Must capture: roles/claims; where scope is derived (token vs. DB); a resource × operation × role matrix; which tables have row-level security and which rely on code-enforced checks.
   - Reviewer use: the baseline an access-control audit compares the code against. `flows.md` shows it in motion; this is the static reference.

4. **`variables.md`** — configuration and secrets, mapped to risk.
   - Must capture: a table of Name · used-by · scope (server/client) · source · rotation · risk; explicit confirmation that no secret is bundled client-side; a pre-go-live checklist.
   - Reviewer use: the secrets/PII-leak surface and the rotation plan during incident response.

5. **`tests.md`** — the verification map: which documented rules are actually checked, which are only proposed, and which are checked by nothing.
   - Must capture, in three clearly separated sections so the map can't read falsely green:
     - **Existing coverage** — tests that are in the repo *today*, each tied to the rule it pins (so the map reflects reality, not a wish-list).
     - **Proposed tests** — recommended cases not yet written, marked by **test type** (automated unit/integration · guarded live · manual review).
     - **Gaps** — documented rules with no verification at all, ranked by what crossing them exposes.
   - Each row carries: use-case → rule → expected behavior (including the deny/negative case) → evidence source (doc + code) → status (existing / proposed / none). It also notes which checks are CI-required and gate merges to `main`.
   - Reviewer use: the operational form of "documented == implemented" — it shows whether each rule the other docs claim is actually pinned by a test today, only proposed, or unverified.
   - Produced by `/derive-tests` (not `/document-app`), because it is derived from the other docs and the existing test suite rather than read off a subsystem.

## Conditional documents (include only when the capability exists)

6. **`emails.md`** — every notification the system sends. *Include only if the app sends transactional or automated email.*
   - Must capture: the queue → processor → provider path; templates and the variables they accept; retry/backoff behavior; where to look when a send fails.
   - Reviewer use: spotting unvalidated template inputs and PII exposure boundaries.

7. **`cron.md`** — all scheduled work and how to operate it safely. *Include only if scheduled or background jobs exist.*
   - Must capture: an inventory table (job → schedule → function → secrets → limits → retry); how each job stays idempotent; how internal calls authenticate; where to see last runs.
   - Reviewer use: finding forgeable triggers and unbounded background jobs.

8. **`seo.md`** — how a single-page app handles SEO and social previews. *Include only if there are public/indexable or bot-facing routes.*
   - Must capture: the preview approach (static meta / prerender / edge HTML); a route → needs-SEO → public-data-only table; how dynamic metadata is sanitized; bot-vs-human routing.
   - Reviewer use: catching public-data-only violations and metadata injection on bot routes.

9. **`automation.md`** — embedded agents and other automation paths. *Include only if the app embeds AI agents, LLM workflows, tool-calling, webhooks, or external automation.*
   - Must capture, per automation/agent: trigger + owner + whether it runs automatically or only after approval; the inputs it may read and the **exact tools/APIs it may call** (the tool surface is itself a hard guardrail); where **steering** lives (the prompt) vs. the **non-prompt hard guardrails**; the **output contract** back to the app (schema, validation, failure handling); **app-owned side effects vs. agent-owned suggestions**; and the controls — approval gates, audit/timeline logging, rate limits, retries, kill switch.
   - Reviewer use: makes hidden automation paths visible and draws the line between what an agent *proposes* and what the app *enforces* — the highest-risk surface in modern AI-built apps.

## Notes

- Each produced doc adds a reference to itself in `architecture.md` under a "Related Documents" section, so the set stays discoverable.
- Skip any conditional document that doesn't apply, and say so in one line rather than inventing content.
- Keep examples and finished templates out of these docs — they describe *this* system, not the general method.
- The agent operating-context file (`CLAUDE.md` / `AGENTS.md`) is a *different* artifact — instructions derived from these docs, not system documentation. It is produced at the handoff step by `/ship-check`, not here.
- `tests.md` is produced by `/derive-tests`; the rest are produced by `/document-app`.
- Do not include an "updated date" line; the file's history is the source of truth.

所有文件

1 个文件

安装 shipping-artifacts

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/phuryn/pm-skills/tree/main/pm-ai-shipping/skills/shipping-artifacts # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能

相关技能

tc-tracker
更新时间 2026-08-27
nuxthub
更新时间 2026-08-23
golang-dependency-injection
更新时间 2026-06-29
altimate-data-engineering-skills
更新时间 2026-08-23
OR