santa-method
affaan-m/ECC
두 명의 독립적인 검토 담당자가 출력 품질을 확인하며, 두 명 모두의 승인을 받아야만 출하할 수 있습니다.
...모든 것을 확장하십시오산타 방식
다중 에이전트 적대적 검증 프레임워크. 목록을 작성하고 두 번 확인하세요. 문제가 있다면, 해결될 때까지 수정하세요.
핵심 통찰: 자신의 산출물을 검토하는 단일 에이전트는 그 산출물을 생성한 것과 동일한 편향, 지식의 공백, 체계적인 오류를 공유하게 된다. 공통된 맥락이 없는 두 명의 독립적인 검토자는 이러한 실패 모드를 해소한다.
활성화 시점
다음과 같은 경우 이 기술을 활용하세요:
- 출력 결과가 공개되거나 배포되거나 최종 사용자가 이용할 경우
- 규정 준수, 규제 또는 브랜드 제약 사항을 반드시 준수해야 할 때
- 사람의 검토 없이 코드가 프로덕션 환경에 배포되는 경우
- 콘텐츠의 정확성이 중요한 경우(기술 문서, 교육 자료, 고객 대상 문구)
- 무작위 검사로만으로는 체계적인 패턴을 포착하기 어려운 대규모 일괄 생성 시
- 허위 정보 생성 위험이 높은 경우(주장, 통계, API 참조, 법률 용어)
내부 초안, 탐색적 연구 또는 결정론적 검증이 필요한 작업에는 사용하지 마십시오(이러한 용도에는 빌드/테스트/린트 파이프라인을 사용하십시오).
아키텍처
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
단계별 세부 사항
1단계: 목록 작성 (생성)
주요 작업을 실행하십시오. 평소의 생성 워크플로에는 변경 사항이 없습니다. 산타 방식(Santa Method)은 생성 전략이 아닌, 생성 후 검증 단계입니다.
# The generator runs as normal
output = generate(task_spec)
2단계: 두 번 확인하기(독립적인 이중 검토)
두 명의 검토 담당자를 병렬로 배치합니다. 핵심 불변 조건:
- 컨텍스트 격리 — 두 검토자 모두 상대방의 평가 내용을 볼 수 없음
- 동일한 평가 기준 — 두 검토자 모두 동일한 평가 기준을 받음
- 동일한 입력 — 두 검토자 모두 원본 사양서와 생성된 출력을 받음
- 구조화된 출력 — 각 검토자는 서술문이 아닌 유형화된 판정을 반환
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
평가 기준표 설계
평가 기준표는 가장 중요한 입력 요소입니다. 모호한 평가 기준표는 모호한 검토 결과를 낳습니다. 모든 기준에는 객관적인 합격/불합격 조건이 반드시 명시되어야 합니다.
| 평가 기준 | 합격 조건 | 불합격 신호 |
|---|---|---|
| 사실적 정확성 | 모든 주장이 출처 자료나 상식에 비추어 검증 가능해야 함 | 허위 통계, 잘못된 버전 번호, 존재하지 않는 API |
| 허구적 내용이 없음 | 허구적인 실체, 인용문, URL 또는 참고 문헌 없음 | 존재하지 않는 페이지로의 링크, 출처가 명시된 인용문이지만 출처가 없음 |
| 완전성 | 사양서의 모든 요구 사항이 다루어짐 | 누락된 섹션, 생략된 경계 사례, 불완전한 커버리지 |
| 규정 준수 | 모든 프로젝트별 제약 조건을 충족함 | 금지된 용어 사용, 어조 위반, 규정 미준수 |
| 내부 일관성 | 출력 내용 내 모순 없음 | A 절에서는 X라고 명시되어 있고, B 절에서는 X가 아니라고 명시됨 |
| 기술적 정확성 | 코드가 컴파일/실행되며, 알고리즘이 타당함 | 구문 오류, 논리 오류, 잘못된 복잡도 주장 |
도메인별 평가 기준 확장
콘텐츠/마케팅:
- 브랜드 목소리 준수
- SEO 요건 충족 (키워드 밀도, 메타 태그, 구조)
- 경쟁사의 상표권 오용 없음
- CTA가 포함되어 있고 올바르게 링크됨
코드:
- 타입 안전성 (
any누수 없음, 적절한 null 처리) - 오류 처리 범위
- 보안 (코드에 기밀 정보 없음, 입력 유효성 검사, 인젝션 방지)
- 새로운 경로에 대한 테스트 커버리지
규정 준수 관련 사항 (규제, 법률, 재무):
- 결과에 대한 보증이나 근거 없는 주장 없음
- 필요한 면책 조항이 명시되어 있음
- 승인된 용어만 사용
- 관할권에 적합한 언어 사용
3단계: 합격 또는 불합격 (판결 관문)
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
두 명 모두 통과해야 하는 이유: 단 한 명의 검토자만 문제를 발견하더라도, 그 문제는 실재하는 것입니다. 다른 검토자의 사각지대는 바로 ‘산타 방법’이 제거하고자 하는 실패 모드 그 자체입니다.
4단계: ‘착해질 때까지 수정’(수렴 루프)
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
중요: 각 검토 라운드에서는 새로운 검토자를 배정해야 합니다. 이전 라운드의 맥락이 앵커링 편향을 유발하므로, 검토자는 이전 라운드의 정보를 기억해서는 안 됩니다.
구현 패턴
패턴 A: 클로드 코드 서브에이전트 (권장)
서브에이전트는 진정한 맥락 격리를 제공합니다. 각 검토자는 상태를 공유하지 않는 별도의 프로세스입니다.
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
패턴 B: 순차적 인라인 방식 (대체 방안)
서브에이전트를 사용할 수 없는 경우, 명시적인 컨텍스트 재설정을 통해 격리를 시뮬레이션합니다:
- 출력 생성
- 새로운 컨텍스트: "당신은 검토자 1입니다. 오직 이 평가 기준에 대해서만 평가하십시오. 문제점을 찾아내십시오."
- 발견 사항을 그대로 기록
- 컨텍스트를 완전히 초기화합니다
- 새로운 컨텍스트: "당신은 심사위원 2입니다. 오직 이 평가 기준에 대해서만 평가하십시오. 문제점을 찾아내십시오."
- 두 검토 결과를 비교하고, 수정하고, 반복하십시오
서브에이전트 패턴이 훨씬 우수합니다 — 인라인 시뮬레이션은 검토자 간 맥락이 혼합될 위험이 있습니다.
패턴 C: 일괄 표본 추출
대규모 배치(100개 이상 항목)의 경우, 모든 항목에 대해 ‘산타’를 완전히 실행하는 것은 비용 면에서 비현실적입니다. 계층화 표본 추출을 사용하십시오:
- 무작위 표본(배치의 10~15%, 최소 5개 항목)에 대해 ‘산타’를 실행하십시오
- 오류 유형(환각, 규정 준수, 완전성 등)별로 분류
- 체계적인 패턴이 나타나면 전체 배치에 대해 표적화된 수정 조치를 적용합니다
- 수정된 배치에 대해 재표본 추출 및 재검증 수행
- 결함이 없는 표본이 검증을 통과할 때까지 이 과정을 반복하십시오
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
고장 모드 및 완화 조치
| 불량 유형 | 증상 | 완화 조치 |
|---|---|---|
| 무한 루프 | 수정 후에도 검토자가 계속해서 새로운 문제를 발견함 | 반복 횟수 상한(3). 상급자에게 보고. |
| 형식적인 승인 | 두 검토자 모두 모든 항목을 통과시킴 | 도발적인 프롬프트: "당신의 임무는 문제를 찾는 것이지, 승인하는 것이 아닙니다." |
| 주관적인 기준 편차 | 검토자들이 오류가 아닌 스타일 선호도를 지적함 | 객관적인 합격/불합격 기준만을 담은 엄격한 평가 기준 |
| 회귀 현상 수정 | 문제 A를 수정하면 문제 B가 발생합니다 | 매 라운드마다 새로운 심사위원들이 회귀 현상을 포착 |
| 검토자 간 일치 편향 | 두 심사위원 모두 같은 부분을 놓침 | 독립성을 통해 완화될 수는 있으나 완전히 제거되지는 않습니다. 중요한 결과물의 경우, 세 번째 검토자를 추가하거나 사람이 직접 무작위 검사를 실시하십시오. |
| 비용 급증 | 대규모 산출물에 대한 반복 작업이 지나치게 잦음 | 일괄 샘플링 방식. 검증 주기당 예산 상한선 설정. |
다른 기술과의 통합
| 스킬 | 관계 |
|---|---|
| 검증 루프 | 결정론적 검사(빌드, 린트, 테스트)에 사용. 의미론적 검사(정확도, 허위 출력)에는 Santa를 사용. 먼저 검증 루프를 실행하고, 그 다음 Santa를 실행합니다. |
| 평가 하네스 | Santa 메서드의 결과가 평가 지표에 반영됩니다. Santa 실행 전반에 걸쳐 pass@k를 추적하여 시간 경과에 따른 생성기 품질을 측정합니다. |
| 지속적 학습 v2 | Santa의 발견 사항은 본능이 됩니다. 동일한 기준에서 반복적으로 실패할 경우 → 해당 패턴을 피하기 위한 학습된 행동이 형성됩니다. |
| 전략적 압축 | 컴팩트 처리 전에 산타를 실행합니다. 검증 도중 검토 맥락을 잃지 않도록 하십시오. |
지표
산타 방법의 효과를 측정하기 위해 다음 항목을 추적하십시오:
- 1차 통과율: 1차 라운드에서 산타를 통과한 출력의 비율 (목표: >70%)
- 수렴까지의 평균 반복 횟수: NICE에 도달하기까지의 평균 라운드 수 (목표: <1.5)
- 이슈 분류: 실패 유형의 분포(환각 vs. 완전성 vs. 규정 준수)
- 검토자 간 일치도: 두 검토자 모두에게 표시된 이슈와 한 명에게만 표시된 이슈의 비율 (일치도가 낮을 경우 = 평가 기준이 더 엄격해져야 함)
- 누락률: 출시 후 발견된 문제 중 산타가 포착했어야 할 문제의 비율 (목표: 0)
비용 분석
산타 방식은 검증 주기당 생성 토큰 비용의 약 2~3배가 소요됩니다. 대부분의 중요한 결과물에 대해서는 이는 매우 경제적인 비용입니다:
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
배치 작업의 경우, 샘플링 패턴을 통해 비용을 전체 검증 비용의 약 15~20% 수준으로 낮추면서도 체계적인 결함의 90% 이상을 포착합니다.
---
name: santa-method
description: Uses two independent review agents to verify output quality, requiring both to pass before shipping.
---
# Santa Method
Multi-agent adversarial verification framework. Make a list, check it twice. If it's naughty, fix it until it's nice.
The core insight: a single agent reviewing its own output shares the same biases, knowledge gaps, and systematic errors that produced the output. Two independent reviewers with no shared context break this failure mode.
## When to Activate
Invoke this skill when:
- Output will be published, deployed, or consumed by end users
- Compliance, regulatory, or brand constraints must be enforced
- Code ships to production without human review
- Content accuracy matters (technical docs, educational material, customer-facing copy)
- Batch generation at scale where spot-checking misses systemic patterns
- Hallucination risk is elevated (claims, statistics, API references, legal language)
Do NOT use for internal drafts, exploratory research, or tasks with deterministic verification (use build/test/lint pipelines for those).
## Architecture
```
┌─────────────┐
│ GENERATOR │ Phase 1: Make a List
│ (Agent A) │ Produce the deliverable
└──────┬───────┘
│ output
▼
┌──────────────────────────────┐
│ DUAL INDEPENDENT REVIEW │ Phase 2: Check It Twice
│ │
│ ┌───────────┐ ┌───────────┐ │ Two agents, same rubric,
│ │ Reviewer B │ │ Reviewer C │ │ no shared context
│ └─────┬─────┘ └─────┬─────┘ │
│ │ │ │
└────────┼──────────────┼────────┘
│ │
▼ ▼
┌──────────────────────────────┐
│ VERDICT GATE │ Phase 3: Naughty or Nice
│ │
│ B passes AND C passes → NICE │ Both must pass.
│ Otherwise → NAUGHTY │ No exceptions.
└──────┬──────────────┬─────────┘
│ │
NICE NAUGHTY
│ │
▼ ▼
[ SHIP ] ┌─────────────┐
│ FIX CYCLE │ Phase 4: Fix Until Nice
│ │
│ iteration++ │ Collect all flags.
│ if i > MAX: │ Fix all issues.
│ escalate │ Re-run both reviewers.
│ else: │ Loop until convergence.
│ goto Ph.2 │
└──────────────┘
```
## Phase Details
### Phase 1: Make a List (Generate)
Execute the primary task. No changes to your normal generation workflow. Santa Method is a post-generation verification layer, not a generation strategy.
```python
# The generator runs as normal
output = generate(task_spec)
```
### Phase 2: Check It Twice (Independent Dual Review)
Spawn two review agents in parallel. Critical invariants:
1. **Context isolation** — neither reviewer sees the other's assessment
2. **Identical rubric** — both receive the same evaluation criteria
3. **Same inputs** — both receive the original spec AND the generated output
4. **Structured output** — each returns a typed verdict, not prose
```python
REVIEWER_PROMPT = """
You are an independent quality reviewer. You have NOT seen any other review of this output.
## Task Specification
{task_spec}
## Output Under Review
{output}
## Evaluation Rubric
{rubric}
## Instructions
Evaluate the output against EACH rubric criterion. For each:
- PASS: criterion fully met, no issues
- FAIL: specific issue found (cite the exact problem)
Return your assessment as structured JSON:
{
"verdict": "PASS" | "FAIL",
"checks": [
{"criterion": "...", "result": "PASS|FAIL", "detail": "..."}
],
"critical_issues": ["..."], // blockers that must be fixed
"suggestions": ["..."] // non-blocking improvements
}
Be rigorous. Your job is to find problems, not to approve.
"""
```
```python
# Spawn reviewers in parallel (Claude Code subagents)
review_b = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer B")
review_c = Agent(prompt=REVIEWER_PROMPT.format(...), description="Santa Reviewer C")
# Both run concurrently — neither sees the other
```
### Rubric Design
The rubric is the most important input. Vague rubrics produce vague reviews. Every criterion must have an objective pass/fail condition.
| Criterion | Pass Condition | Failure Signal |
|-----------|---------------|----------------|
| Factual accuracy | All claims verifiable against source material or common knowledge | Invented statistics, wrong version numbers, nonexistent APIs |
| Hallucination-free | No fabricated entities, quotes, URLs, or references | Links to pages that don't exist, attributed quotes with no source |
| Completeness | Every requirement in the spec is addressed | Missing sections, skipped edge cases, incomplete coverage |
| Compliance | Passes all project-specific constraints | Banned terms used, tone violations, regulatory non-compliance |
| Internal consistency | No contradictions within the output | Section A says X, section B says not-X |
| Technical correctness | Code compiles/runs, algorithms are sound | Syntax errors, logic bugs, wrong complexity claims |
#### Domain-Specific Rubric Extensions
**Content/Marketing:**
- Brand voice adherence
- SEO requirements met (keyword density, meta tags, structure)
- No competitor trademark misuse
- CTA present and correctly linked
**Code:**
- Type safety (no `any` leaks, proper null handling)
- Error handling coverage
- Security (no secrets in code, input validation, injection prevention)
- Test coverage for new paths
**Compliance-Sensitive (regulated, legal, financial):**
- No outcome guarantees or unsubstantiated claims
- Required disclaimers present
- Approved terminology only
- Jurisdiction-appropriate language
### Phase 3: Naughty or Nice (Verdict Gate)
```python
def santa_verdict(review_b, review_c):
"""Both reviewers must pass. No partial credit."""
if review_b.verdict == "PASS" and review_c.verdict == "PASS":
return "NICE" # Ship it
# Merge flags from both reviewers, deduplicate
all_issues = dedupe(review_b.critical_issues + review_c.critical_issues)
all_suggestions = dedupe(review_b.suggestions + review_c.suggestions)
return "NAUGHTY", all_issues, all_suggestions
```
Why both must pass: if only one reviewer catches an issue, that issue is real. The other reviewer's blind spot is exactly the failure mode Santa Method exists to eliminate.
### Phase 4: Fix Until Nice (Convergence Loop)
```python
MAX_ITERATIONS = 3
for iteration in range(MAX_ITERATIONS):
verdict, issues, suggestions = santa_verdict(review_b, review_c)
if verdict == "NICE":
log_santa_result(output, iteration, "passed")
return ship(output)
# Fix all critical issues (suggestions are optional)
output = fix_agent.execute(
output=output,
issues=issues,
instruction="Fix ONLY the flagged issues. Do not refactor or add unrequested changes."
)
# Re-run BOTH reviewers on fixed output (fresh agents, no memory of previous round)
review_b = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
review_c = Agent(prompt=REVIEWER_PROMPT.format(output=output, ...))
# Exhausted iterations — escalate
log_santa_result(output, MAX_ITERATIONS, "escalated")
escalate_to_human(output, issues)
```
Critical: each review round uses **fresh agents**. Reviewers must not carry memory from previous rounds, as prior context creates anchoring bias.
## Implementation Patterns
### Pattern A: Claude Code Subagents (Recommended)
Subagents provide true context isolation. Each reviewer is a separate process with no shared state.
```bash
# In a Claude Code session, use the Agent tool to spawn reviewers
# Both agents run in parallel for speed
```
```python
# Pseudocode for Agent tool invocation
reviewer_b = Agent(
description="Santa Review B",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
reviewer_c = Agent(
description="Santa Review C",
prompt=f"Review this output for quality...\n\nRUBRIC:\n{rubric}\n\nOUTPUT:\n{output}"
)
```
### Pattern B: Sequential Inline (Fallback)
When subagents aren't available, simulate isolation with explicit context resets:
1. Generate output
2. New context: "You are Reviewer 1. Evaluate ONLY against this rubric. Find problems."
3. Record findings verbatim
4. Clear context completely
5. New context: "You are Reviewer 2. Evaluate ONLY against this rubric. Find problems."
6. Compare both reviews, fix, repeat
The subagent pattern is strictly superior — inline simulation risks context bleed between reviewers.
### Pattern C: Batch Sampling
For large batches (100+ items), full Santa on every item is cost-prohibitive. Use stratified sampling:
1. Run Santa on a random sample (10-15% of batch, minimum 5 items)
2. Categorize failures by type (hallucination, compliance, completeness, etc.)
3. If systematic patterns emerge, apply targeted fixes to the entire batch
4. Re-sample and re-verify the fixed batch
5. Continue until a clean sample passes
```python
import random
def santa_batch(items, rubric, sample_rate=0.15):
sample = random.sample(items, max(5, int(len(items) * sample_rate)))
for item in sample:
result = santa_full(item, rubric)
if result.verdict == "NAUGHTY":
pattern = classify_failure(result.issues)
items = batch_fix(items, pattern) # Fix all items matching pattern
return santa_batch(items, rubric) # Re-sample
return items # Clean sample → ship batch
```
## Failure Modes and Mitigations
| Failure Mode | Symptom | Mitigation |
|-------------|---------|------------|
| Infinite loop | Reviewers keep finding new issues after fixes | Max iteration cap (3). Escalate. |
| Rubber stamping | Both reviewers pass everything | Adversarial prompt: "Your job is to find problems, not approve." |
| Subjective drift | Reviewers flag style preferences, not errors | Tight rubric with objective pass/fail criteria only |
| Fix regression | Fixing issue A introduces issue B | Fresh reviewers each round catch regressions |
| Reviewer agreement bias | Both reviewers miss the same thing | Mitigated by independence, not eliminated. For critical output, add a third reviewer or human spot-check. |
| Cost explosion | Too many iterations on large outputs | Batch sampling pattern. Budget caps per verification cycle. |
## Integration with Other Skills
| Skill | Relationship |
|-------|-------------|
| Verification Loop | Use for deterministic checks (build, lint, test). Santa for semantic checks (accuracy, hallucinations). Run verification-loop first, Santa second. |
| Eval Harness | Santa Method results feed eval metrics. Track pass@k across Santa runs to measure generator quality over time. |
| Continuous Learning v2 | Santa findings become instincts. Repeated failures on the same criterion → learned behavior to avoid the pattern. |
| Strategic Compact | Run Santa BEFORE compacting. Don't lose review context mid-verification. |
## Metrics
Track these to measure Santa Method effectiveness:
- **First-pass rate**: % of outputs that pass Santa on round 1 (target: >70%)
- **Mean iterations to convergence**: average rounds to NICE (target: <1.5)
- **Issue taxonomy**: distribution of failure types (hallucination vs. completeness vs. compliance)
- **Reviewer agreement**: % of issues flagged by both reviewers vs. only one (low agreement = rubric needs tightening)
- **Escape rate**: issues found post-ship that Santa should have caught (target: 0)
## Cost Analysis
Santa Method costs approximately 2-3x the token cost of generation alone per verification cycle. For most high-stakes output, this is a bargain:
```
Cost of Santa = (generation tokens) + 2×(review tokens per round) × (avg rounds)
Cost of NOT Santa = (reputation damage) + (correction effort) + (trust erosion)
```
For batch operations, the sampling pattern reduces cost to ~15-20% of full verification while catching >90% of systematic issues.
모든 파일
1개 파일santa-method 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/affaan-m/ECC/tree/main/skills/santa-method # Copy SKILL.md to your .claude/skills/ directory
복사





집
