ab-test-analysis
phuryn/pm-skills
통계적 유의성, 표본 크기 검증, 신뢰 구간을 바탕으로 A/B 테스트 결과를 분석하고, 캠페인 실행·확대·중단에 대한 권장 사항을 제시합니다.
...모든 것을 확장하십시오A/B 테스트 분석
통계적 엄밀성을 바탕으로 A/B 테스트 결과를 평가하고, 그 결과를 명확한 제품 결정으로 전환합니다.
배경
$ARGUMENTS에 대한 A/B 테스트 결과를 분석하고 있습니다.
사용자가 데이터 파일(CSV, Excel 또는 분석 데이터 내보내기 파일)을 제공하면, 이를 직접 읽어 분석하십시오. 필요한 경우 통계 계산을 위한 Python 스크립트를 생성하십시오.
지침
실험 내용 파악:
- 가설은 무엇이었나요?
- 무엇이 변경되었나요(변형)?
- 주요 지표는 무엇인가요? 보조 지표가 있나요?
- 테스트는 얼마나 오래 진행되었나요?
- 트래픽 분배 비율은 어떻게 되나요?
테스트 설정을 검증하십시오:
- 표본 크기: 예상 효과 크기에 비해 표본 크기가 충분히 큰가요?
- 다음 공식을 사용하십시오: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- 검정력이 부족한 경우(검정력 <80%) 표시하십시오
- 기간: 테스트가 최소 1~2개의 완전한 비즈니스 주기에 걸쳐 진행되었나요?
- 무작위 배정: 표본 비율 불일치(SRM)의 증거가 있는가?
- 신기함/선점 효과: 초기 행동 변화를 상쇄할 만큼 충분한 시간이 있었는가?
- 표본 크기: 예상 효과 크기에 비해 표본 크기가 충분히 큰가요?
통계적 유의성 계산:
- 대조군 및 변형군의 전환율
- 상대적 상승률: (변형군 - 대조군) / 대조군 × 100
- p-값: 양측 z-검정 또는 카이-제곱 검정 사용
- 신뢰 구간: 차이에 대한 95% 신뢰 구간
- 통계적 유의성: p < 0.05인가?
- 실무적 유의성: 리프트가 비즈니스 측면에서 의미가 있는가?
사용자가 원시 데이터를 제공하면, 이를 계산하기 위해 파이썬 스크립트를 생성하여 실행하십시오.
가드레일 지표 확인:
- 가드레일 지표(매출, 참여도, 페이지 로딩 시간) 중 저하된 항목이 있는가?
- 가드레일 지표가 악화된 상태에서 주요 지표가 개선되었다고 해도 진정한 성공이라고 볼 수 없습니다
결과 해석:
결과 권장 사항 상당한 긍정적 상승세, 가드레일 문제 없음 출시 — 100%로 확대 적용 상당한 긍정적 효과, 가드레일 관련 우려 사항 있음 조사 — 배포 전 장단점 파악 크지 않으나 긍정적인 추세 테스트 기간 연장 — 더 많은 데이터 또는 더 큰 효과가 필요 유의미하지 않음, 정체 시험 중단 — 유의미한 차이가 감지되지 않음 통계적으로 유의미한 부정적 상승 효과 출시하지 마세요 — 대조군으로 되돌리고 원인을 분석하세요 분석 요약 제공:
## A/B Test Results: [Test Name] **Hypothesis**: [What we expected] **Duration**: [X days] | **Sample**: [N control / M variant] | Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | [Guardrail] | ... | ... | ... | ... | ... | **Recommendation**: [Ship / Extend / Stop / Investigate] **Reasoning**: [Why] **Next steps**: [What to do]
단계별로 생각하십시오. 마크다운 형식으로 저장하십시오. 원시 데이터가 제공된 경우 계산을 위한 Python 스크립트를 생성하십시오.
추가 자료
- A/B 테스트 입문 + 예시
- 제품 아이디어 테스트: 궁극의 검증 실험 라이브러리
- 올바른 지표를 추적하고 계신가요?
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---
## A/B Test Analysis
Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.
### Context
You are analyzing A/B test results for **$ARGUMENTS**.
If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.
### Instructions
1. **Understand the experiment**:
- What was the hypothesis?
- What was changed (the variant)?
- What is the primary metric? Any guardrail metrics?
- How long did the test run?
- What is the traffic split?
2. **Validate the test setup**:
- **Sample size**: Is the sample large enough for the expected effect size?
- Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- Flag if the test is underpowered (<80% power)
- **Duration**: Did the test run for at least 1-2 full business cycles?
- **Randomization**: Any evidence of sample ratio mismatch (SRM)?
- **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?
3. **Calculate statistical significance**:
- **Conversion rate** for control and variant
- **Relative lift**: (variant - control) / control × 100
- **p-value**: Using a two-tailed z-test or chi-squared test
- **Confidence interval**: 95% CI for the difference
- **Statistical significance**: Is p < 0.05?
- **Practical significance**: Is the lift meaningful for the business?
If the user provides raw data, generate and run a Python script to calculate these.
4. **Check guardrail metrics**:
- Did any guardrail metrics (revenue, engagement, page load time) degrade?
- A winning primary metric with degraded guardrails may not be a true win
5. **Interpret results**:
| Outcome | Recommendation |
|---|---|
| Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
| Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
| Not significant, positive trend | **Extend the test** — need more data or larger effect |
| Not significant, flat | **Stop the test** — no meaningful difference detected |
| Significant negative lift | **Don't ship** — revert to control, analyze why |
6. **Provide the analysis summary**:
```
## A/B Test Results: [Test Name]
**Hypothesis**: [What we expected]
**Duration**: [X days] | **Sample**: [N control / M variant]
| Metric | Control | Variant | Lift | p-value | Significant? |
|---|---|---|---|---|---|
| [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
| [Guardrail] | ... | ... | ... | ... | ... |
**Recommendation**: [Ship / Extend / Stop / Investigate]
**Reasoning**: [Why]
**Next steps**: [What to do]
```
Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.
---
### Further Reading
- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)
모든 파일
1개 파일ab-test-analysis 설치
스킬 파일을 다운로드하여 .claude/skills/ 디렉터리에 압축을 풀어주세요.
ZIP 다운로드저장소를 클론하고 스킬 파일을 프로젝트에 복사하세요.
git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory
복사





집
