ab-test-analysis
phuryn/pm-skills
A/Bテストの結果を、統計的有意性、標本サイズの妥当性検証、信頼区間、および実施・延長・中止に関する推奨事項に基づいて分析します。
...すべて拡張しますA/Bテストの分析
統計的に厳密な手法を用いてA/Bテストの結果を評価し、その結果を明確な製品方針へと反映させます。
背景
$ARGUMENTSのA/Bテスト結果を分析しています。
ユーザーからデータファイル(CSV、Excel、または分析ツールのエクスポートデータ)が提供された場合は、それらを直接読み込んで分析してください。必要に応じて、統計計算用のPythonスクリプトを生成してください。
手順
実験内容を理解する:
- 仮説は何でしたか?
- 何が変更されたか(バリアント)?
- 主要な指標は何ですか? ガードレール指標はありますか?
- テストはどのくらいの期間実施されたか?
- トラフィックの割り当てはどのようになっていますか?
テスト設定の妥当性を検証する:
- 標本サイズ:予想される効果サイズに対して、標本サイズは十分に大きいですか?
- 次の式を使用してください:n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- 検定の検出力が不足している場合(検出力が80%未満)はフラグを立てる
- 期間:テストは少なくとも1~2回の完全なビジネスサイクルにわたって実施されましたか?
- 無作為化:標本比率の不一致(SRM)を示す証拠はあるか?
- 新規性/プライマシー効果:初期の行動変化が解消されるのに十分な時間が確保されていたか?
- 標本サイズ:予想される効果サイズに対して、標本サイズは十分に大きいですか?
統計的有意性の算出:
- 対照群およびバリエーションのコンバージョン率
- 相対リフト:(バリアント - 対照群) / 対照群 × 100
- p値:両側z検定またはカイ二乗検定を使用
- 信頼区間:差の95%信頼区間
- 統計的有意性:p < 0.05 か?
- 実用上の有意性:リフトはビジネスにとって有意義か?
ユーザーが生データを提供した場合は、これらを計算するためのPythonスクリプトを生成して実行する。
ガードレール指標を確認する:
- ガードレール指標(収益、エンゲージメント、ページ読み込み時間)に悪化は見られたか?
- 主要指標の数値は良好でも、ガードレール指標が低下している場合は、真の成功とは言えない可能性がある
結果の解釈:
結果 推奨事項 著しいプラス効果が見られ、ガードレールの問題もない リリースする — 100%に展開する 大幅なパフォーマンス向上、ガードレールに関する懸念あり 調査 — リリース前にトレードオフを把握する 顕著ではないが、好ましい傾向 テストを延長 — さらなるデータまたはより大きな効果が必要 有意差なし、横ばい 試験を中止 — 有意な差は検出されなかった 有意なマイナス効果 リリースしない — コントロール版に戻し、原因を分析する 分析の概要を提示する:
## A/B Test Results: [Test Name] **Hypothesis**: [What we expected] **Duration**: [X days] | **Sample**: [N control / M variant] | Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | [Guardrail] | ... | ... | ... | ... | ... | **Recommendation**: [Ship / Extend / Stop / Investigate] **Reasoning**: [Why] **Next steps**: [What to do]
段階的に検討する。Markdown形式で保存する。生データが提供されている場合は、計算用のPythonスクリプトを生成する。
参考資料
- A/Bテスト入門+事例
- プロダクトアイデアの検証:究極の検証実験ライブラリ
- 適切な指標を追跡できていますか?
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---
## A/B Test Analysis
Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.
### Context
You are analyzing A/B test results for **$ARGUMENTS**.
If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.
### Instructions
1. **Understand the experiment**:
- What was the hypothesis?
- What was changed (the variant)?
- What is the primary metric? Any guardrail metrics?
- How long did the test run?
- What is the traffic split?
2. **Validate the test setup**:
- **Sample size**: Is the sample large enough for the expected effect size?
- Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- Flag if the test is underpowered (<80% power)
- **Duration**: Did the test run for at least 1-2 full business cycles?
- **Randomization**: Any evidence of sample ratio mismatch (SRM)?
- **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?
3. **Calculate statistical significance**:
- **Conversion rate** for control and variant
- **Relative lift**: (variant - control) / control × 100
- **p-value**: Using a two-tailed z-test or chi-squared test
- **Confidence interval**: 95% CI for the difference
- **Statistical significance**: Is p < 0.05?
- **Practical significance**: Is the lift meaningful for the business?
If the user provides raw data, generate and run a Python script to calculate these.
4. **Check guardrail metrics**:
- Did any guardrail metrics (revenue, engagement, page load time) degrade?
- A winning primary metric with degraded guardrails may not be a true win
5. **Interpret results**:
| Outcome | Recommendation |
|---|---|
| Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
| Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
| Not significant, positive trend | **Extend the test** — need more data or larger effect |
| Not significant, flat | **Stop the test** — no meaningful difference detected |
| Significant negative lift | **Don't ship** — revert to control, analyze why |
6. **Provide the analysis summary**:
```
## A/B Test Results: [Test Name]
**Hypothesis**: [What we expected]
**Duration**: [X days] | **Sample**: [N control / M variant]
| Metric | Control | Variant | Lift | p-value | Significant? |
|---|---|---|---|---|---|
| [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
| [Guardrail] | ... | ... | ... | ... | ... |
**Recommendation**: [Ship / Extend / Stop / Investigate]
**Reasoning**: [Why]
**Next steps**: [What to do]
```
Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.
---
### Further Reading
- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)
すべてのファイル
1件のファイルab-test-analysisをインストール
スキルファイルをダウンロードし、.claude/skills/ ディレクトリに解凍してください。
ZIPをダウンロードリポジトリをクローンし、スキルファイルをプロジェクトにコピーしてください。
git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory
コピー





家
