santa-method
affaan-m/ECC
由两名独立的审核员对输出质量进行核查,必须两人均通过审核后方可发货。
...展开全部圣诞老人法
多智能体对抗性验证框架。列出清单,仔细核对两遍。如果存在问题,就不断修正,直到完美无缺。
核心洞见:单个智能体审查其自身产出的结果时,会延续产生该结果时所存在的偏见、知识缺口和系统性错误。而两位互不了解背景的独立审查者则能打破这种失效模式。
何时启用
在以下情况下调用此技能:
- 输出结果将被发布、部署或由终端用户使用时
- 必须遵守合规、监管或品牌限制
- 代码在未经人工审查的情况下发布到生产环境
- 内容准确性至关重要(技术文档、教育资料、面向客户的文案)
- 大规模批量生成,且抽查无法发现系统性模式
- 产生“幻觉”的风险较高(声明、统计数据、API 引用、法律术语)
请勿用于内部草稿、探索性研究或可进行确定性验证的任务(此类情况请使用构建/测试/代码检查管道)。
架构
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
阶段详情
第 1 阶段:生成清单(生成)
执行主要任务。无需更改您的常规生成工作流。“圣诞老人法”是一个生成后的验证层,而非生成策略。
# The generator runs as normal
output = generate(task_spec)
第 2 阶段:双重检查(独立双重审查)
并行启动两个审查代理。关键不变量:
- 上下文隔离——两位评审员均无法查看对方的评估结果
- 统一的评分标准——双方均采用相同的评估标准
- 相同输入——双方均收到原始规格说明及生成的输出结果
- 结构化输出——每位评审员返回的是类型化的裁决结果,而非自然语言描述
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
评分标准设计
评分标准是最重要的输入。模糊的评分标准会导致模糊的评审结果。每项标准都必须具备客观的“通过/未通过”判定条件。
| 评判标准 | 通过条件 | 未通过信号 |
|---|---|---|
| 事实准确性 | 所有论点均可通过原始资料或常识进行验证 | 虚构的统计数据、错误的版本号、不存在的API |
| 无虚构内容 | 不存在虚构的实体、引语、URL 或参考文献 | 指向不存在的页面的链接,以及未注明来源的引语 |
| 完整性 | 规范中的每一项要求均已得到处理 | 缺失章节、遗漏边界情况、覆盖不完整 |
| 合规性 | 通过了所有项目特定的约束条件 | 使用了禁用词、语气不当、不符合监管要求 |
| 内部一致性 | 输出内容无矛盾 | A部分说X,B部分却说非X |
| 技术正确性 | 代码能编译/运行,算法正确 | 语法错误、逻辑错误、复杂度声明有误 |
领域特定评分标准扩展
内容/营销:
- 符合品牌语调
- 满足SEO要求(关键词密度、元标签、页面结构)
- 未滥用竞争对手商标
- 包含行动号召(CTA)且链接正确
代码:
- 类型安全(无
any内存泄漏,正确的空值处理) - 错误处理覆盖率
- 安全性(代码中无机密信息、输入验证、防范注入攻击)
- 新路径的测试覆盖率
合规敏感性(受监管、法律、财务):
- 不作结果保证,不作无根据的声明
- 已包含必要的免责声明
- 仅使用经批准的术语
- 符合管辖区域的措辞
第三阶段:好坏判定(裁决关卡)
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
为何必须两者都通过:如果只有一位评审员发现问题,那么该问题确实存在。另一位评审员的盲点,正是“圣诞老人法”旨在消除的故障模式。
第4阶段:修改直至合格(收敛循环)
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
关键要求:每轮评审均需使用全新的代理。评审员不得带入前几轮的记忆,因为先前的上下文会产生锚定偏见。
实现模式
模式 A:Claude 代码子代理(推荐)
子代理可实现真正的上下文隔离。每个审阅者都是一个独立进程,不共享任何状态。
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
模式 B:顺序内联(备选方案)
当无法使用子代理时,可通过显式重置上下文来模拟隔离:
- 生成输出
- 新上下文:“你是评审员 1。仅根据此评分标准进行评估。找出问题。”
- 逐字记录发现的问题
- 彻底清除上下文
- 新上下文:“你是评审员2。仅根据本评分标准进行评估。找出问题。”
- 比较两份评审结果,修改,重复
子代理模式具有绝对优势——内联模拟会导致评审员之间出现上下文泄漏的风险。
模式 C:批量抽样
对于大型批次(100+项),对每项都运行完整的“圣诞老人”程序成本过高。请使用分层抽样:
- 对随机样本(占批量的10-15%,至少5个项目)运行“圣诞老人”算法
- 按类型对失败情况进行分类(幻觉、合规性、完整性等)
- 若出现系统性问题,对整个批次进行针对性修复
- 对已修复的批次进行重新抽样并重新验证
- 重复此过程,直至抽样合格
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
失效模式及缓解措施
| 失效模式 | 症状 | 缓解措施 |
|---|---|---|
| 无限循环 | 修复后,审核人员仍不断发现新问题 | 设置最大迭代次数上限(3)。上报。 |
| 走过场 | 两位审阅者全都通过了所有内容 | 对抗性提示:“你的工作是发现问题,而不是批准。” |
| 主观偏移 | 评审员标记的是风格偏好,而非错误 | 仅采用严格的评分标准,且仅包含客观的“通过/未通过”标准 |
| 修复回归问题 | 修复问题 A 却引入了问题 B | 每轮由新的评审员发现回归问题 |
| 评审员一致性偏差 | 两位评审员都漏审了同一处问题 | 通过独立性可缓解该问题,但无法彻底消除。对于关键输出,应增加第三位评审员或进行人工抽查。 |
| 成本激增 | 大型输出结果的迭代次数过多 | 采用批量抽样模式。为每个验证周期设定预算上限。 |
与其他技能的整合
| 技能 | 关系 |
|---|---|
| 验证循环 | 用于确定性检查(构建、代码检查、测试)。Santa用于语义检查(准确性、幻觉)。先运行验证循环,再运行Santa。 |
| 评估框架 | Santa方法的结果将作为评估指标的输入。通过追踪Santa运行过程中的pass@k指标,来衡量生成器随时间推移的质量变化。 |
| 持续学习 v2 | Santa的发现将转化为“直觉”。同一标准上反复失败 → 系统将学习避免该模式的行为。 |
| 策略性压缩 | 在进行代码紧凑化之前运行“Santa”。避免在验证过程中丢失代码审查上下文。 |
指标
跟踪以下指标以衡量 Santa 方法的有效性:
- 首次通过率:第一轮通过“圣诞老人”测试的输出占比(目标:>70%)
- 收敛所需平均迭代次数:达到 NICE 所需的平均轮数(目标:<1.5)
- 问题分类:失败类型的分布情况(幻觉 vs. 完整性 vs. 合规性)
- 评审员一致性:两位评审员均标记的问题占总问题的百分比,与仅由一位标记的问题相比(一致性低 = 评审标准需收紧)
- 漏检率:发布后发现的、本应被“圣诞老人”方法检测出的问题(目标:0)
成本分析
每个验证周期中,“圣诞老人法”的成本约为仅生成代币成本的2-3倍。对于大多数高风险输出而言,这非常划算:
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
对于批处理操作,采样模式将成本降低至全量验证的约15-20%,同时能捕获超过90%的系统性问题。
---
name: santa-method
description: Uses two independent review agents to verify output quality, requiring both to pass before shipping.
---
# Santa Method
Multi-agent adversarial verification framework. Make a list, check it twice. If it's naughty, fix it until it's nice.
The core insight: a single agent reviewing its own output shares the same biases, knowledge gaps, and systematic errors that produced the output. Two independent reviewers with no shared context break this failure mode.
## When to Activate
Invoke this skill when:
- Output will be published, deployed, or consumed by end users
- Compliance, regulatory, or brand constraints must be enforced
- Code ships to production without human review
- Content accuracy matters (technical docs, educational material, customer-facing copy)
- Batch generation at scale where spot-checking misses systemic patterns
- Hallucination risk is elevated (claims, statistics, API references, legal language)
Do NOT use for internal drafts, exploratory research, or tasks with deterministic verification (use build/test/lint pipelines for those).
## Architecture
```
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
```
## Phase Details
### Phase 1: Make a List (Generate)
Execute the primary task. No changes to your normal generation workflow. Santa Method is a post-generation verification layer, not a generation strategy.
```python
# The generator runs as normal
output = generate(task_spec)
```
### Phase 2: Check It Twice (Independent Dual Review)
Spawn two review agents in parallel. Critical invariants:
1. **Context isolation** — neither reviewer sees the other's assessment
2. **Identical rubric** — both receive the same evaluation criteria
3. **Same inputs** — both receive the original spec AND the generated output
4. **Structured output** — each returns a typed verdict, not prose
```python
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
```
```python
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
```
### Rubric Design
The rubric is the most important input. Vague rubrics produce vague reviews. Every criterion must have an objective pass/fail condition.
| Criterion | Pass Condition | Failure Signal |
|-----------|---------------|----------------|
| Factual accuracy | All claims verifiable against source material or common knowledge | Invented statistics, wrong version numbers, nonexistent APIs |
| Hallucination-free | No fabricated entities, quotes, URLs, or references | Links to pages that don't exist, attributed quotes with no source |
| Completeness | Every requirement in the spec is addressed | Missing sections, skipped edge cases, incomplete coverage |
| Compliance | Passes all project-specific constraints | Banned terms used, tone violations, regulatory non-compliance |
| Internal consistency | No contradictions within the output | Section A says X, section B says not-X |
| Technical correctness | Code compiles/runs, algorithms are sound | Syntax errors, logic bugs, wrong complexity claims |
#### Domain-Specific Rubric Extensions
**Content/Marketing:**
- Brand voice adherence
- SEO requirements met (keyword density, meta tags, structure)
- No competitor trademark misuse
- CTA present and correctly linked
**Code:**
- Type safety (no `any` leaks, proper null handling)
- Error handling coverage
- Security (no secrets in code, input validation, injection prevention)
- Test coverage for new paths
**Compliance-Sensitive (regulated, legal, financial):**
- No outcome guarantees or unsubstantiated claims
- Required disclaimers present
- Approved terminology only
- Jurisdiction-appropriate language
### Phase 3: Naughty or Nice (Verdict Gate)
```python
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
```
Why both must pass: if only one reviewer catches an issue, that issue is real. The other reviewer's blind spot is exactly the failure mode Santa Method exists to eliminate.
### Phase 4: Fix Until Nice (Convergence Loop)
```python
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
```
Critical: each review round uses **fresh agents**. Reviewers must not carry memory from previous rounds, as prior context creates anchoring bias.
## Implementation Patterns
### Pattern A: Claude Code Subagents (Recommended)
Subagents provide true context isolation. Each reviewer is a separate process with no shared state.
```bash
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
```
```python
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
```
### Pattern B: Sequential Inline (Fallback)
When subagents aren't available, simulate isolation with explicit context resets:
1. Generate output
2. New context: "You are Reviewer 1. Evaluate ONLY against this rubric. Find problems."
3. Record findings verbatim
4. Clear context completely
5. New context: "You are Reviewer 2. Evaluate ONLY against this rubric. Find problems."
6. Compare both reviews, fix, repeat
The subagent pattern is strictly superior — inline simulation risks context bleed between reviewers.
### Pattern C: Batch Sampling
For large batches (100+ items), full Santa on every item is cost-prohibitive. Use stratified sampling:
1. Run Santa on a random sample (10-15% of batch, minimum 5 items)
2. Categorize failures by type (hallucination, compliance, completeness, etc.)
3. If systematic patterns emerge, apply targeted fixes to the entire batch
4. Re-sample and re-verify the fixed batch
5. Continue until a clean sample passes
```python
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
```
## Failure Modes and Mitigations
| Failure Mode | Symptom | Mitigation |
|-------------|---------|------------|
| Infinite loop | Reviewers keep finding new issues after fixes | Max iteration cap (3). Escalate. |
| Rubber stamping | Both reviewers pass everything | Adversarial prompt: "Your job is to find problems, not approve." |
| Subjective drift | Reviewers flag style preferences, not errors | Tight rubric with objective pass/fail criteria only |
| Fix regression | Fixing issue A introduces issue B | Fresh reviewers each round catch regressions |
| Reviewer agreement bias | Both reviewers miss the same thing | Mitigated by independence, not eliminated. For critical output, add a third reviewer or human spot-check. |
| Cost explosion | Too many iterations on large outputs | Batch sampling pattern. Budget caps per verification cycle. |
## Integration with Other Skills
| Skill | Relationship |
|-------|-------------|
| Verification Loop | Use for deterministic checks (build, lint, test). Santa for semantic checks (accuracy, hallucinations). Run verification-loop first, Santa second. |
| Eval Harness | Santa Method results feed eval metrics. Track pass@k across Santa runs to measure generator quality over time. |
| Continuous Learning v2 | Santa findings become instincts. Repeated failures on the same criterion → learned behavior to avoid the pattern. |
| Strategic Compact | Run Santa BEFORE compacting. Don't lose review context mid-verification. |
## Metrics
Track these to measure Santa Method effectiveness:
- **First-pass rate**: % of outputs that pass Santa on round 1 (target: >70%)
- **Mean iterations to convergence**: average rounds to NICE (target: <1.5)
- **Issue taxonomy**: distribution of failure types (hallucination vs. completeness vs. compliance)
- **Reviewer agreement**: % of issues flagged by both reviewers vs. only one (low agreement = rubric needs tightening)
- **Escape rate**: issues found post-ship that Santa should have caught (target: 0)
## Cost Analysis
Santa Method costs approximately 2-3x the token cost of generation alone per verification cycle. For most high-stakes output, this is a bargain:
```
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
```
For batch operations, the sampling pattern reduces cost to ~15-20% of full verification while catching >90% of systematic issues.
所有文件
1 个文件安装 santa-method
下载技能文件并将其解压到 .claude/skills/ 目录中。
下载ZIP克隆仓库并复制技能文件到您的项目中。
git clone https://github.com/affaan-m/ECC/tree/main/skills/santa-method # Copy SKILL.md to your .claude/skills/ directory
复制





首页
