옵션

ab-test-analysis

phuryn/pm-skills phuryn/pm-skills

통계적 유의성, 표본 크기 검증, 신뢰 구간을 바탕으로 A/B 테스트 결과를 분석하고, 캠페인 실행·확대·중단에 대한 권장 사항을 제시합니다.

...모든 것을 확장하십시오
0
업데이트 된 시간 2026년 9월 29일

A/B 테스트 분석

통계적 엄밀성을 바탕으로 A/B 테스트 결과를 평가하고, 그 결과를 명확한 제품 결정으로 전환합니다.

배경

$ARGUMENTS에 대한 A/B 테스트 결과를 분석하고 있습니다.

사용자가 데이터 파일(CSV, Excel 또는 분석 데이터 내보내기 파일)을 제공하면, 이를 직접 읽어 분석하십시오. 필요한 경우 통계 계산을 위한 Python 스크립트를 생성하십시오.

지침

  1. 실험 내용 파악:

    • 가설은 무엇이었나요?
    • 무엇이 변경되었나요(변형)?
    • 주요 지표는 무엇인가요? 보조 지표가 있나요?
    • 테스트는 얼마나 오래 진행되었나요?
    • 트래픽 분배 비율은 어떻게 되나요?
  2. 테스트 설정을 검증하십시오:

    • 표본 크기: 예상 효과 크기에 비해 표본 크기가 충분히 큰가요?
      • 다음 공식을 사용하십시오: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
      • 검정력이 부족한 경우(검정력 <80%) 표시하십시오
    • 기간: 테스트가 최소 1~2개의 완전한 비즈니스 주기에 걸쳐 진행되었나요?
    • 무작위 배정: 표본 비율 불일치(SRM)의 증거가 있는가?
    • 신기함/선점 효과: 초기 행동 변화를 상쇄할 만큼 충분한 시간이 있었는가?
  3. 통계적 유의성 계산:

    • 대조군 및 변형군의 전환율
    • 상대적 상승률: (변형군 - 대조군) / 대조군 × 100
    • p-값: 양측 z-검정 또는 카이-제곱 검정 사용
    • 신뢰 구간: 차이에 대한 95% 신뢰 구간
    • 통계적 유의성: p < 0.05인가?
    • 실무적 유의성: 리프트가 비즈니스 측면에서 의미가 있는가?

    사용자가 원시 데이터를 제공하면, 이를 계산하기 위해 파이썬 스크립트를 생성하여 실행하십시오.

  4. 가드레일 지표 확인:

    • 가드레일 지표(매출, 참여도, 페이지 로딩 시간) 중 저하된 항목이 있는가?
    • 가드레일 지표가 악화된 상태에서 주요 지표가 개선되었다고 해도 진정한 성공이라고 볼 수 없습니다
  5. 결과 해석:

    결과 권장 사항
    상당한 긍정적 상승세, 가드레일 문제 없음 출시 — 100%로 확대 적용
    상당한 긍정적 효과, 가드레일 관련 우려 사항 있음 조사 — 배포 전 장단점 파악
    크지 않으나 긍정적인 추세 테스트 기간 연장 — 더 많은 데이터 또는 더 큰 효과가 필요
    유의미하지 않음, 정체 시험 중단 — 유의미한 차이가 감지되지 않음
    통계적으로 유의미한 부정적 상승 효과 출시하지 마세요 — 대조군으로 되돌리고 원인을 분석하세요
  6. 분석 요약 제공:

    ## A/B Test Results: [Test Name]
    
    **Hypothesis**: [What we expected]
    **Duration**: [X days] | **Sample**: [N control / M variant]
    
    | Metric | Control | Variant | Lift | p-value | Significant? |
    |---|---|---|---|---|---|
    | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
    | [Guardrail] | ... | ... | ... | ... | ... |
    
    **Recommendation**: [Ship / Extend / Stop / Investigate]
    **Reasoning**: [Why]
    **Next steps**: [What to do]
    

단계별로 생각하십시오. 마크다운 형식으로 저장하십시오. 원시 데이터가 제공된 경우 계산을 위한 Python 스크립트를 생성하십시오.

추가 자료

  • A/B 테스트 입문 + 예시
  • 제품 아이디어 테스트: 궁극의 검증 실험 라이브러리
  • 올바른 지표를 추적하고 계신가요?
GitHub에서 보기
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---

## A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

### Context

You are analyzing A/B test results for **$ARGUMENTS**.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

### Instructions

1. **Understand the experiment**:
   - What was the hypothesis?
   - What was changed (the variant)?
   - What is the primary metric? Any guardrail metrics?
   - How long did the test run?
   - What is the traffic split?

2. **Validate the test setup**:
   - **Sample size**: Is the sample large enough for the expected effect size?
     - Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
     - Flag if the test is underpowered (<80% power)
   - **Duration**: Did the test run for at least 1-2 full business cycles?
   - **Randomization**: Any evidence of sample ratio mismatch (SRM)?
   - **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?

3. **Calculate statistical significance**:
   - **Conversion rate** for control and variant
   - **Relative lift**: (variant - control) / control × 100
   - **p-value**: Using a two-tailed z-test or chi-squared test
   - **Confidence interval**: 95% CI for the difference
   - **Statistical significance**: Is p < 0.05?
   - **Practical significance**: Is the lift meaningful for the business?

   If the user provides raw data, generate and run a Python script to calculate these.

4. **Check guardrail metrics**:
   - Did any guardrail metrics (revenue, engagement, page load time) degrade?
   - A winning primary metric with degraded guardrails may not be a true win

5. **Interpret results**:

   | Outcome | Recommendation |
   |---|---|
   | Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
   | Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
   | Not significant, positive trend | **Extend the test** — need more data or larger effect |
   | Not significant, flat | **Stop the test** — no meaningful difference detected |
   | Significant negative lift | **Don't ship** — revert to control, analyze why |

6. **Provide the analysis summary**:
   ```
   ## A/B Test Results: [Test Name]

   **Hypothesis**: [What we expected]
   **Duration**: [X days] | **Sample**: [N control / M variant]

   | Metric | Control | Variant | Lift | p-value | Significant? |
   |---|---|---|---|---|---|
   | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
   | [Guardrail] | ... | ... | ... | ... | ... |

   **Recommendation**: [Ship / Extend / Stop / Investigate]
   **Reasoning**: [Why]
   **Next steps**: [What to do]
   ```

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.

---

### Further Reading

- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)

모든 파일

1개 파일

ab-test-analysis 설치

스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.

ZIP 다운로드

저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.

git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory

복사 복사
빠른 설정: 스킬 폴더를 .claude/skills/로 복사하세요. Claude가 해당 스킬을 자동으로 감지하여 사용할 것입니다.
저장소 phuryn/pm-skills

관련 스킬

microservices-patterns
업데이트 된 시간 2026년 6월 29일
jpa-patterns
업데이트 된 시간 2026년 6월 30일
fabric-lakehouse
업데이트 된 시간 2026년 6월 30일
prisma-expert
업데이트 된 시간 2026년 6월 29일
OR