选项
首页首页 Skill 数据库管理 ab-test-analysis

ab-test-analysis

phuryn/pm-skills phuryn/pm-skills

分析具有统计学意义的A/B测试结果,包括样本量验证、置信区间,以及上线/延长/停止的建议。

...展开全部
0
更新时间 2026-09-29

A/B 测试分析

以严谨的统计方法评估A/B测试结果,并将分析结果转化为清晰的产品决策。

背景

您正在分析 $ARGUMENTS 的 A/B 测试结果。

如果用户提供了数据文件(CSV、Excel 或分析导出文件),请直接读取并分析这些文件。必要时,生成用于统计计算的 Python 脚本。

操作指南

  1. 理解实验:

    • 假设是什么?
    • 进行了哪些更改(即变体)?
    • 主要指标是什么?是否有任何防护指标?
    • 测试持续了多长时间?
    • 流量分配比例是多少?
  2. 验证测试设置:

    • 样本量:样本量是否足以反映预期的效应量?
      • 使用公式:n = (Z²α/2 × 2 × p × (1-p)) / MDE²
      • 若检验功效不足(<80%),请标注
    • 持续时间:测试是否至少运行了1-2个完整的业务周期?
    • 随机化:是否有样本比例失配(SRM)的迹象?
    • 新颖性/首因效应:是否有足够的时间消除最初的行为变化?
  3. 计算统计显著性:

    • 对照组和变体组的转化率
    • 相对提升率:(变体组 - 对照组) / 对照组 × 100
    • p值:采用双尾z检验或卡方检验
    • 置信区间:差异的95%置信区间
    • 统计学显著性:p 是否小于 0.05?
    • 实际意义:该提升率对业务有意义吗?

    如果用户提供了原始数据,请生成并运行一个 Python 脚本来计算这些指标。

  4. 检查防护指标:

    • 是否有任何防护指标(收入、用户参与度、页面加载时间)出现恶化?
    • 若某项主要指标表现优异但防护指标出现恶化,则可能并非真正的成功
  5. 解读结果:

    结果 建议
    显著的积极提升,无警戒线问题 上线——全面推广
    显著的积极提升,但存在安全防护措施方面的顾虑 调查——在发布前权衡利弊
    影响不大,呈积极趋势 延长测试——需要更多数据或更显著的效果
    无显著变化,持平 停止测试——未检测到有意义的差异
    显著的负向提升 不要发布——恢复到对照组,分析原因
  6. 提供分析摘要:

    ## A/B Test Results: [Test Name]
    
    **Hypothesis**: [What we expected]
    **Duration**: [X days] | **Sample**: [N control / M variant]
    
    | Metric | Control | Variant | Lift | p-value | Significant? |
    |---|---|---|---|---|---|
    | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
    | [Guardrail] | ... | ... | ... | ... | ... |
    
    **Recommendation**: [Ship / Extend / Stop / Investigate]
    **Reasoning**: [Why]
    **Next steps**: [What to do]
    

循序渐进地思考。保存为 Markdown 格式。若提供原始数据,请生成用于计算的 Python 脚本。

进一步阅读

  • A/B 测试入门 + 示例
  • 产品创意的测试:终极验证实验库
  • 您是否在追踪正确的指标?
在 GitHub 上查看
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---

## A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

### Context

You are analyzing A/B test results for **$ARGUMENTS**.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

### Instructions

1. **Understand the experiment**:
   - What was the hypothesis?
   - What was changed (the variant)?
   - What is the primary metric? Any guardrail metrics?
   - How long did the test run?
   - What is the traffic split?

2. **Validate the test setup**:
   - **Sample size**: Is the sample large enough for the expected effect size?
     - Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
     - Flag if the test is underpowered (<80% power)
   - **Duration**: Did the test run for at least 1-2 full business cycles?
   - **Randomization**: Any evidence of sample ratio mismatch (SRM)?
   - **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?

3. **Calculate statistical significance**:
   - **Conversion rate** for control and variant
   - **Relative lift**: (variant - control) / control × 100
   - **p-value**: Using a two-tailed z-test or chi-squared test
   - **Confidence interval**: 95% CI for the difference
   - **Statistical significance**: Is p < 0.05?
   - **Practical significance**: Is the lift meaningful for the business?

   If the user provides raw data, generate and run a Python script to calculate these.

4. **Check guardrail metrics**:
   - Did any guardrail metrics (revenue, engagement, page load time) degrade?
   - A winning primary metric with degraded guardrails may not be a true win

5. **Interpret results**:

   | Outcome | Recommendation |
   |---|---|
   | Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
   | Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
   | Not significant, positive trend | **Extend the test** — need more data or larger effect |
   | Not significant, flat | **Stop the test** — no meaningful difference detected |
   | Significant negative lift | **Don't ship** — revert to control, analyze why |

6. **Provide the analysis summary**:
   ```
   ## A/B Test Results: [Test Name]

   **Hypothesis**: [What we expected]
   **Duration**: [X days] | **Sample**: [N control / M variant]

   | Metric | Control | Variant | Lift | p-value | Significant? |
   |---|---|---|---|---|---|
   | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
   | [Guardrail] | ... | ... | ... | ... | ... |

   **Recommendation**: [Ship / Extend / Stop / Investigate]
   **Reasoning**: [Why]
   **Next steps**: [What to do]
   ```

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.

---

### Further Reading

- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)

所有文件

1 个文件

安装 ab-test-analysis

下载技能文件并将其解压到 .claude/skills/ 目录中。

下载ZIP

克隆仓库并复制技能文件到您的项目中。

git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory

复制 复制
快速设置: 将技能文件夹复制到 .claude/skills/ Claude 会自动检测并使用该技能

相关技能

microservices-patterns
更新时间 2026-06-29
jpa-patterns
更新时间 2026-06-30
fabric-lakehouse
更新时间 2026-06-30
prisma-expert
更新时间 2026-06-29
OR