santa-method
affaan-m/ECC
2名の独立したレビュー担当者が出力品質を確認し、両者が承認して初めて出荷されます。
...すべて拡張しますサンタ・メソッド
マルチエージェント対抗検証フレームワーク。「リストを作成し、二度確認する。もし問題があれば、改善されるまで修正する。」
核心となる洞察:自身の出力をレビューする単一のエージェントは、その出力を生み出したのと同じバイアス、知識の欠如、および体系的な誤りを共有している。共通の文脈を持たない2人の独立したレビューアが、この失敗モードを打破する。
いつ活用すべきか
以下の場合にこのスキルを活用してください:
- 出力が公開、展開、またはエンドユーザーに利用される場合
- コンプライアンス、規制、またはブランド上の制約を遵守する必要がある場合
- 人間のレビューを経ずにコードが本番環境にリリースされる場合
- コンテンツの正確性が重要な場合(技術文書、教育資料、顧客向けコピーなど)
- スポットチェックでは体系的なパターンを見逃してしまうような、大規模なバッチ生成が行われる場合
- 「幻覚」のリスクが高い(主張、統計、API リファレンス、法的文言)
内部の草案、探索的な調査、または決定論的な検証が可能なタスクには使用しないでください(それらにはビルド/テスト/lintパイプラインを使用してください)。
アーキテクチャ
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
フェーズの詳細
フェーズ 1: リストの作成(生成)
主要なタスクを実行します。通常の生成ワークフローに変更はありません。「サンタ・メソッド」は生成戦略ではなく、生成後の検証レイヤーです。
# The generator runs as normal
output = generate(task_spec)
フェーズ 2: 再確認(独立した二重レビュー)
2人のレビュー担当者を並行して起動します。重要な不変条件:
- コンテキストの分離 — どちらのレビュー担当者も、もう一方の評価内容を見ることができない
- 同一の評価基準 — 両者とも同じ評価基準を受け取る
- 同一の入力 — 両者とも元の仕様書および生成された出力を受け取る
- 構造化された出力 — 各レビュー担当者は、文章ではなく型付けされた判定結果を返す
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
評価基準の設計
評価基準表は最も重要な入力要素である。曖昧な評価基準表は、曖昧なレビューを生む。すべての基準には、客観的な合格・不合格の条件が定められていなければならない。
| 評価基準 | 合格条件 | 不合格の兆候 |
|---|---|---|
| 事実の正確性 | すべての主張が、出典資料または一般常識と照らし合わせて検証可能であること | 架空の統計データ、誤ったバージョン番号、存在しないAPI |
| 妄想が含まれていないこと | 架空のエンティティ、引用、URL、参照がないこと | 存在しないページへのリンク、出典のない引用 |
| 完全性 | 仕様書のすべての要件が網羅されている | セクションの欠落、エッジケースの省略、網羅性の不備 |
| 準拠性 | プロジェクト固有の制約をすべて満たしている | 禁止用語の使用、文体の違反、規制への不適合 |
| 内部の一貫性 | 出力内容に矛盾がない | セクションAではXと記載されているが、セクションBではXではないと記載されている |
| 技術的な正確性 | コードがコンパイル・実行され、アルゴリズムが妥当であること | 構文エラー、論理バグ、複雑度に関する誤った主張 |
ドメイン固有の評価基準の拡張
コンテンツ/マーケティング:
- ブランドボイスへの準拠
- SEO要件を満たしている(キーワード密度、メタタグ、構造)
- 競合他社の商標の不正使用がないこと
- CTAが設置され、正しくリンクされている
コード:
- 型安全性(
anyリークなし、適切なnull処理) - エラー処理の網羅性
- セキュリティ(コード内の機密情報の排除、入力の妥当性チェック、インジェクションの防止)
- 新しいパスに対するテストカバレッジ
コンプライアンス関連(規制、法務、財務):
- 結果の保証や根拠のない主張は一切行わない
- 必要な免責事項が記載されていること
- 承認された用語のみを使用する
- 管轄区域に応じた適切な表現
フェーズ3:合格か不合格か(判定ゲート)
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
両方が合格しなければならない理由:もし1人の査読者のみが問題を指摘した場合、その問題は実際に存在する。もう1人の査読者の「死角」こそが、まさに「サンタ・メソッド」が排除するために存在する失敗モードそのものである。
フェーズ4:問題解消まで修正(収束ループ)
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
重要:各レビューラウンドでは、新しいエージェントを使用する。以前のラウンドの記憶をレビュー担当者が持ち越してはならない。過去の文脈はアンカリングバイアスを引き起こすためである。
実装パターン
パターンA:Claudeコードサブエージェント(推奨)
サブエージェントは真のコンテキスト分離を実現します。各レビュアーは、状態を共有しない独立したプロセスです。
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
パターンB:順次インライン(フォールバック)
サブエージェントが利用できない場合は、明示的なコンテキストリセットによって隔離をシミュレートします:
- 出力を生成
- 新しいコンテキスト:「あなたはレビュアー1です。この評価基準のみに基づいて評価を行ってください。問題点を見つけ出してください。」
- 発見事項を逐語的に記録する
- コンテキストを完全にクリアする
- 新しいコンテキスト:「あなたは査読者2です。この評価基準にのみ基づいて評価してください。問題点を見つけてください。」
- 両方のレビューを比較し、修正し、繰り返す
サブエージェント方式が圧倒的に優れています。インラインシミュレーションでは、査読者間で文脈が混在するリスクがあります。
パターンC:バッチサンプリング
大規模なバッチ(100件以上)の場合、すべての項目に対してサンタをフル実行するのはコスト的に非現実的です。層別抽出を使用します:
- 無作為抽出されたサンプル(バッチの10~15%、最低5項目)に対してサンタを実行する
- 失敗をタイプ(幻覚、コンプライアンス、完全性など)別に分類する
- 体系的なパターンが確認された場合は、バッチ全体に対して対象を絞った修正を適用する
- 修正済みのロットについて再サンプリングを行い、再検証する
- 問題のないサンプルが合格するまでこの手順を繰り返す
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
故障モードと緩和策
| 不具合モード | 症状 | 是正措置 |
|---|---|---|
| 無限ループ | 修正後もレビュー担当者が新たな不具合を次々と発見する | 反復回数の上限(3回)。エスカレーションを行う。 |
| 形式的な承認 | 両方のレビュー担当者がすべてを承認してしまう | 敵対的なプロンプト:「あなたの仕事は問題を見つけることであって、承認することではない。」 |
| 主観的な基準のずれ | 査読者がエラーではなく、スタイルの好みを指摘する | 客観的な合格・不合格基準のみを定めた厳格な評価基準 |
| 回帰の修正 | 問題Aの修正により問題Bが発生する | 各ラウンドで新しい査読者がリグレッションを検出する |
| 査読者の合意バイアス | 両方の査読者が同じ点を見落とす | 独立性によって軽減されるが、完全に排除されるわけではない。重要な出力については、3人目のレビューアを追加するか、人間による抜き打ちチェックを行う。 |
| コストの急増 | 大規模な成果物に対する反復回数が多すぎる | バッチサンプリング方式を採用する。検証サイクルごとの予算上限を設定する。 |
他のスキルとの統合
| スキル | 関係 |
|---|---|
| 検証ループ | 決定論的チェック(ビルド、リント、テスト)に使用。セマンティックチェック(精度、幻覚)にはSantaを使用。まず検証ループを実行し、次にSantaを実行する。 |
| 評価ハーネス | Santaメソッドの結果は評価メトリクスに反映される。Santaの実行におけるpass@kを追跡し、時間の経過に伴うジェネレータの品質を測定する。 |
| 継続的学習 v2 | Santaによる発見は「直感」となる。同じ基準で繰り返し失敗した場合 → そのパターンを回避する行動を学習する。 |
| 戦略的コンパクト | コンパクト化を行う前にサンタを実行する。検証の途中でレビューの文脈が失われないようにする。 |
指標
サンタ・メソッドの有効性を測定するために、以下の指標を追跡する:
- 初回合格率:第1ラウンドでサンタに合格した出力の割合(目標:70%以上)
- 収束までの平均反復回数:NICEに到達するまでの平均ラウンド数(目標:1.5未満)
- 課題の分類:失敗タイプの分布(幻覚 vs. 完全性 vs. 準拠性)
- レビューア間の一致率:両方のレビューアによって指摘された課題と、一方のみによって指摘された課題の割合(一致率が低い場合=評価基準の厳格化が必要)
- エスケープ率:リリース後に発見されたが、本来ならサンタが検出すべきだった課題の割合(目標:0)
コスト分析
Santaメソッドのコストは、検証サイクルあたり、生成のみにかかるトークンコストの約2~3倍です。リスクの高い出力のほとんどにとって、これは割安な投資と言えます:
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
バッチ処理の場合、サンプリングパターンにより、コストは完全検証の約15~20%に抑えられつつ、体系的な問題の90%以上を検出できます。
---
name: santa-method
description: Uses two independent review agents to verify output quality, requiring both to pass before shipping.
---
# Santa Method
Multi-agent adversarial verification framework. Make a list, check it twice. If it's naughty, fix it until it's nice.
The core insight: a single agent reviewing its own output shares the same biases, knowledge gaps, and systematic errors that produced the output. Two independent reviewers with no shared context break this failure mode.
## When to Activate
Invoke this skill when:
- Output will be published, deployed, or consumed by end users
- Compliance, regulatory, or brand constraints must be enforced
- Code ships to production without human review
- Content accuracy matters (technical docs, educational material, customer-facing copy)
- Batch generation at scale where spot-checking misses systemic patterns
- Hallucination risk is elevated (claims, statistics, API references, legal language)
Do NOT use for internal drafts, exploratory research, or tasks with deterministic verification (use build/test/lint pipelines for those).
## Architecture
```
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
```
## Phase Details
### Phase 1: Make a List (Generate)
Execute the primary task. No changes to your normal generation workflow. Santa Method is a post-generation verification layer, not a generation strategy.
```python
# The generator runs as normal
output = generate(task_spec)
```
### Phase 2: Check It Twice (Independent Dual Review)
Spawn two review agents in parallel. Critical invariants:
1. **Context isolation** — neither reviewer sees the other's assessment
2. **Identical rubric** — both receive the same evaluation criteria
3. **Same inputs** — both receive the original spec AND the generated output
4. **Structured output** — each returns a typed verdict, not prose
```python
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
```
```python
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
```
### Rubric Design
The rubric is the most important input. Vague rubrics produce vague reviews. Every criterion must have an objective pass/fail condition.
| Criterion | Pass Condition | Failure Signal |
|-----------|---------------|----------------|
| Factual accuracy | All claims verifiable against source material or common knowledge | Invented statistics, wrong version numbers, nonexistent APIs |
| Hallucination-free | No fabricated entities, quotes, URLs, or references | Links to pages that don't exist, attributed quotes with no source |
| Completeness | Every requirement in the spec is addressed | Missing sections, skipped edge cases, incomplete coverage |
| Compliance | Passes all project-specific constraints | Banned terms used, tone violations, regulatory non-compliance |
| Internal consistency | No contradictions within the output | Section A says X, section B says not-X |
| Technical correctness | Code compiles/runs, algorithms are sound | Syntax errors, logic bugs, wrong complexity claims |
#### Domain-Specific Rubric Extensions
**Content/Marketing:**
- Brand voice adherence
- SEO requirements met (keyword density, meta tags, structure)
- No competitor trademark misuse
- CTA present and correctly linked
**Code:**
- Type safety (no `any` leaks, proper null handling)
- Error handling coverage
- Security (no secrets in code, input validation, injection prevention)
- Test coverage for new paths
**Compliance-Sensitive (regulated, legal, financial):**
- No outcome guarantees or unsubstantiated claims
- Required disclaimers present
- Approved terminology only
- Jurisdiction-appropriate language
### Phase 3: Naughty or Nice (Verdict Gate)
```python
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
```
Why both must pass: if only one reviewer catches an issue, that issue is real. The other reviewer's blind spot is exactly the failure mode Santa Method exists to eliminate.
### Phase 4: Fix Until Nice (Convergence Loop)
```python
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
```
Critical: each review round uses **fresh agents**. Reviewers must not carry memory from previous rounds, as prior context creates anchoring bias.
## Implementation Patterns
### Pattern A: Claude Code Subagents (Recommended)
Subagents provide true context isolation. Each reviewer is a separate process with no shared state.
```bash
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
```
```python
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
```
### Pattern B: Sequential Inline (Fallback)
When subagents aren't available, simulate isolation with explicit context resets:
1. Generate output
2. New context: "You are Reviewer 1. Evaluate ONLY against this rubric. Find problems."
3. Record findings verbatim
4. Clear context completely
5. New context: "You are Reviewer 2. Evaluate ONLY against this rubric. Find problems."
6. Compare both reviews, fix, repeat
The subagent pattern is strictly superior — inline simulation risks context bleed between reviewers.
### Pattern C: Batch Sampling
For large batches (100+ items), full Santa on every item is cost-prohibitive. Use stratified sampling:
1. Run Santa on a random sample (10-15% of batch, minimum 5 items)
2. Categorize failures by type (hallucination, compliance, completeness, etc.)
3. If systematic patterns emerge, apply targeted fixes to the entire batch
4. Re-sample and re-verify the fixed batch
5. Continue until a clean sample passes
```python
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
```
## Failure Modes and Mitigations
| Failure Mode | Symptom | Mitigation |
|-------------|---------|------------|
| Infinite loop | Reviewers keep finding new issues after fixes | Max iteration cap (3). Escalate. |
| Rubber stamping | Both reviewers pass everything | Adversarial prompt: "Your job is to find problems, not approve." |
| Subjective drift | Reviewers flag style preferences, not errors | Tight rubric with objective pass/fail criteria only |
| Fix regression | Fixing issue A introduces issue B | Fresh reviewers each round catch regressions |
| Reviewer agreement bias | Both reviewers miss the same thing | Mitigated by independence, not eliminated. For critical output, add a third reviewer or human spot-check. |
| Cost explosion | Too many iterations on large outputs | Batch sampling pattern. Budget caps per verification cycle. |
## Integration with Other Skills
| Skill | Relationship |
|-------|-------------|
| Verification Loop | Use for deterministic checks (build, lint, test). Santa for semantic checks (accuracy, hallucinations). Run verification-loop first, Santa second. |
| Eval Harness | Santa Method results feed eval metrics. Track pass@k across Santa runs to measure generator quality over time. |
| Continuous Learning v2 | Santa findings become instincts. Repeated failures on the same criterion → learned behavior to avoid the pattern. |
| Strategic Compact | Run Santa BEFORE compacting. Don't lose review context mid-verification. |
## Metrics
Track these to measure Santa Method effectiveness:
- **First-pass rate**: % of outputs that pass Santa on round 1 (target: >70%)
- **Mean iterations to convergence**: average rounds to NICE (target: <1.5)
- **Issue taxonomy**: distribution of failure types (hallucination vs. completeness vs. compliance)
- **Reviewer agreement**: % of issues flagged by both reviewers vs. only one (low agreement = rubric needs tightening)
- **Escape rate**: issues found post-ship that Santa should have caught (target: 0)
## Cost Analysis
Santa Method costs approximately 2-3x the token cost of generation alone per verification cycle. For most high-stakes output, this is a bargain:
```
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
```
For batch operations, the sampling pattern reduces cost to ~15-20% of full verification while catching >90% of systematic issues.
すべてのファイル
1件のファイルsanta-methodをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/affaan-m/ECC/tree/main/skills/santa-method # Copy SKILL.md to your .claude/skills/ directory
コピー





家
