選項
首頁首頁 Skill 資料庫管理 ab-test-analysis

ab-test-analysis

phuryn/pm-skills phuryn/pm-skills

分析 A/B 測試結果,包含統計顯著性、樣本數驗證、信賴區間,以及「上線/延長/停止」的建議。

...展開全部
0
更新時間 2026-09-29

A/B 測試分析

以嚴謹的統計方法評估 A/B 測試結果,並將發現轉化為明確的產品決策。

背景

您正在分析 $ARGUMENTS 的 A/B 測試結果。

若使用者提供資料檔案(CSV、Excel 或分析工具匯出檔案),請直接讀取並進行分析。必要時,請生成 Python 腳本以進行統計運算。

操作說明

  1. 了解實驗:

    • 假設是什麼?
    • 變更了哪些內容(變體)?
    • 主要指標是什麼?是否有任何防護指標?
    • 測試持續了多久?
    • 流量分配比例為何?
  2. 驗證測試設定:

    • 樣本大小:樣本量是否足以反映預期的效應大小?
      • 使用公式:n = (Z²α/2 × 2 × p × (1-p)) / MDE²
      • 若檢定效力不足(<80% 效力)請標記
    • 持續時間:測試是否至少運行了 1 至 2 個完整的商業週期?
    • 隨機化:是否有樣本比例不匹配(SRM)的跡象?
    • 新奇效應/首因效應:是否有足夠時間讓初期行為變化消退?
  3. 計算統計顯著性:

    • 對照組與變體組的轉換率
    • 相對提升率:(變體組 - 對照組) / 對照組 × 100
    • p 值:採用雙尾 z 檢定或卡方檢定
    • 信賴區間:差異的 95% 信賴區間
    • 統計顯著性:p 是否小於 0.05?
    • 實務意義:此提升幅度對企業而言是否具有實質意義?

    若使用者提供原始資料,請編寫並執行 Python 腳本以計算這些數值。

  4. 檢查防護指標:

    • 是否有任何防護指標(營收、參與度、頁面載入時間)出現惡化?
    • 若主要指標表現優異,但護欄指標卻出現惡化,這可能並非真正的成功
  5. 解讀結果:

    結果 建議
    顯著的正向提升,無警戒線問題 上線 — 全面推行至 100%
    顯著正向效益,但存在防護欄疑慮 進行調查 — 在正式上線前釐清利弊權衡
    影響不顯著,但呈現正向趨勢 延長測試 — 需要更多數據或更顯著的效果
    無顯著性,呈現平穩態勢 停止測試——未檢測到有意義的差異
    顯著的負向提升 不要上線 — 恢復為對照組,分析原因
  6. 提供分析摘要:

    ## A/B Test Results: [Test Name]
    
    **Hypothesis**: [What we expected]
    **Duration**: [X days] | **Sample**: [N control / M variant]
    
    | Metric | Control | Variant | Lift | p-value | Significant? |
    |---|---|---|---|---|---|
    | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
    | [Guardrail] | ... | ... | ... | ... | ... |
    
    **Recommendation**: [Ship / Extend / Stop / Investigate]
    **Reasoning**: [Why]
    **Next steps**: [What to do]
    

請逐步思考。以 Markdown 格式儲存。若提供原始資料,請生成用於計算的 Python 腳本。

延伸閱讀

  • A/B 測試入門 + 範例
  • 測試產品構想:終極驗證實驗資料庫
  • 您是否正在追蹤正確的指標?
在 GitHub 上查看
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---

## A/B Test Analysis

Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.

### Context

You are analyzing A/B test results for **$ARGUMENTS**.

If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.

### Instructions

1. **Understand the experiment**:
   - What was the hypothesis?
   - What was changed (the variant)?
   - What is the primary metric? Any guardrail metrics?
   - How long did the test run?
   - What is the traffic split?

2. **Validate the test setup**:
   - **Sample size**: Is the sample large enough for the expected effect size?
     - Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
     - Flag if the test is underpowered (<80% power)
   - **Duration**: Did the test run for at least 1-2 full business cycles?
   - **Randomization**: Any evidence of sample ratio mismatch (SRM)?
   - **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?

3. **Calculate statistical significance**:
   - **Conversion rate** for control and variant
   - **Relative lift**: (variant - control) / control × 100
   - **p-value**: Using a two-tailed z-test or chi-squared test
   - **Confidence interval**: 95% CI for the difference
   - **Statistical significance**: Is p < 0.05?
   - **Practical significance**: Is the lift meaningful for the business?

   If the user provides raw data, generate and run a Python script to calculate these.

4. **Check guardrail metrics**:
   - Did any guardrail metrics (revenue, engagement, page load time) degrade?
   - A winning primary metric with degraded guardrails may not be a true win

5. **Interpret results**:

   | Outcome | Recommendation |
   |---|---|
   | Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
   | Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
   | Not significant, positive trend | **Extend the test** — need more data or larger effect |
   | Not significant, flat | **Stop the test** — no meaningful difference detected |
   | Significant negative lift | **Don't ship** — revert to control, analyze why |

6. **Provide the analysis summary**:
   ```
   ## A/B Test Results: [Test Name]

   **Hypothesis**: [What we expected]
   **Duration**: [X days] | **Sample**: [N control / M variant]

   | Metric | Control | Variant | Lift | p-value | Significant? |
   |---|---|---|---|---|---|
   | [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
   | [Guardrail] | ... | ... | ... | ... | ... |

   **Recommendation**: [Ship / Extend / Stop / Investigate]
   **Reasoning**: [Why]
   **Next steps**: [What to do]
   ```

Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.

---

### Further Reading

- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)

所有檔案

1 個檔案

安裝 ab-test-analysis

請下載並將技能檔案解壓縮至您的 .claude/skills/ 目錄中。

下載 ZIP

複製儲存庫並將技能檔案複製到您的專案中。

git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory

複製 複製
快速設定: 將技能資料夾複製到 .claude/skills/ Claude 會自動偵測並使用該技能
儲存庫 phuryn/pm-skills

相關技能

microservices-patterns
更新時間 2026-06-29
jpa-patterns
更新時間 2026-06-30
fabric-lakehouse
更新時間 2026-06-30
prisma-expert
更新時間 2026-06-29
OR