ab-test-analysis
phuryn/pm-skills
Analise os resultados dos testes A/B com significância estatística, validação do tamanho da amostra, intervalos de confiança e recomendações de lançamento/prorrogação/encerramento.
...Expandir tudoAnálise de Testes A/B
Avalie os resultados dos testes A/B com rigor estatístico e traduza as descobertas em decisões claras de produto.
Contexto
Você está analisando os resultados do teste A/B para $ARGUMENTS.
Se o usuário fornecer arquivos de dados (CSV, Excel ou exportações de análise), leia e analise-os diretamente. Gere scripts em Python para cálculos estatísticos quando necessário.
Instruções
Compreenda o experimento:
- Qual era a hipótese?
- O que foi alterado (a variante)?
- Qual é a métrica principal? Existem métricas de segurança?
- Por quanto tempo o teste foi executado?
- Qual é a divisão do tráfego?
Valide a configuração do teste:
- Tamanho da amostra: A amostra é grande o suficiente para o tamanho do efeito esperado?
- Use a fórmula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- Sinalize se o teste tiver poder estatístico insuficiente (
- Duração: O teste foi executado por pelo menos 1-2 ciclos comerciais completos?
- Randomização: Há evidências de incompatibilidade na proporção da amostra (SRM)?
- Efeitos de novidade/primazia: Houve tempo suficiente para eliminar as mudanças iniciais de comportamento?
- Tamanho da amostra: A amostra é grande o suficiente para o tamanho do efeito esperado?
Calcule a significância estatística:
- Taxa de conversão para o controle e a variante
- Elevação relativa: (variante - controle) / controle × 100
- Valor-p: Usando um teste z de duas caudas ou teste qui-quadrado
- Intervalo de confiança: IC de 95% para a diferença
- Significância estatística: O valor-p é
- Significância prática: A elevação é relevante para o negócio?
Se o usuário fornecer dados brutos, gere e execute um script em Python para calcular esses valores.
Verifique as métricas de segurança:
- Alguma métrica de segurança (receita, engajamento, tempo de carregamento da página) piorou?
- Uma métrica principal vencedora com métricas de segurança degradadas pode não ser uma vitória real
Interprete os resultados:
Resultado Recomendação Elevação positiva significativa, sem problemas nas métricas de segurança **Implementar** — liberar para 100% Elevação positiva significativa, preocupações com as métricas de segurança **Investigar** — entender os trade-offs antes de implementar Não significativo, tendência positiva **Prolongar o teste** — são necessários mais dados ou um efeito maior Não significativo, plano **Parar o teste** — nenhuma diferença significativa detectada Elevação negativa significativa **Não implementar** — reverter para o controle, analisar o motivo Forneça o resumo da análise:
## Resultados do Teste A/B: [Nome do Teste] **Hipótese**: [O que esperávamos] **Duração**: [X dias] | **Amostra**: [N controle / M variante] | Métrica | Controle | Variante | Elevação | Valor-p | Significativo? | |---|---|---|---|---|---| | [Principal] | X% | Y% | +Z% | 0.0X | Sim/Não | | [Segurança] | ... | ... | ... | ... | ... | **Recomendação**: [Implementar / Prolongar / Parar / Investigar] **Raciocínio**: [Por quê] **Próximos passos**: [O que fazer]
Pense passo a passo. Salve como markdown. Gere scripts em Python para cálculos se dados brutos forem fornecidos.
Leitura Adicional
- Testes A/B 101 + Exemplos
- Testando Ideias de Produto: A Biblioteca Definitiva de Experimentos de Validação
- Você está rastreando as métricas certas?
---
name: ab-test-analysis
description: Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
---
## A/B Test Analysis
Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.
### Context
You are analyzing A/B test results for **$ARGUMENTS**.
If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.
### Instructions
1. **Understand the experiment**:
- What was the hypothesis?
- What was changed (the variant)?
- What is the primary metric? Any guardrail metrics?
- How long did the test run?
- What is the traffic split?
2. **Validate the test setup**:
- **Sample size**: Is the sample large enough for the expected effect size?
- Use the formula: n = (Z²α/2 × 2 × p × (1-p)) / MDE²
- Flag if the test is underpowered (<80% power)
- **Duration**: Did the test run for at least 1-2 full business cycles?
- **Randomization**: Any evidence of sample ratio mismatch (SRM)?
- **Novelty/primacy effects**: Was there enough time to wash out initial behavior changes?
3. **Calculate statistical significance**:
- **Conversion rate** for control and variant
- **Relative lift**: (variant - control) / control × 100
- **p-value**: Using a two-tailed z-test or chi-squared test
- **Confidence interval**: 95% CI for the difference
- **Statistical significance**: Is p < 0.05?
- **Practical significance**: Is the lift meaningful for the business?
If the user provides raw data, generate and run a Python script to calculate these.
4. **Check guardrail metrics**:
- Did any guardrail metrics (revenue, engagement, page load time) degrade?
- A winning primary metric with degraded guardrails may not be a true win
5. **Interpret results**:
| Outcome | Recommendation |
|---|---|
| Significant positive lift, no guardrail issues | **Ship it** — roll out to 100% |
| Significant positive lift, guardrail concerns | **Investigate** — understand trade-offs before shipping |
| Not significant, positive trend | **Extend the test** — need more data or larger effect |
| Not significant, flat | **Stop the test** — no meaningful difference detected |
| Significant negative lift | **Don't ship** — revert to control, analyze why |
6. **Provide the analysis summary**:
```
## A/B Test Results: [Test Name]
**Hypothesis**: [What we expected]
**Duration**: [X days] | **Sample**: [N control / M variant]
| Metric | Control | Variant | Lift | p-value | Significant? |
|---|---|---|---|---|---|
| [Primary] | X% | Y% | +Z% | 0.0X | Yes/No |
| [Guardrail] | ... | ... | ... | ... | ... |
**Recommendation**: [Ship / Extend / Stop / Investigate]
**Reasoning**: [Why]
**Next steps**: [What to do]
```
Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.
---
### Further Reading
- [A/B Testing 101 + Examples](https://www.productcompass.pm/p/ab-testing-101-for-pms)
- [Testing Product Ideas: The Ultimate Validation Experiments Library](https://www.productcompass.pm/p/the-ultimate-experiments-library)
- [Are You Tracking the Right Metrics?](https://www.productcompass.pm/p/are-you-tracking-the-right-metrics)
Todos os arquivos
1 arquivosInstalar ab-test-analysis
Baixe e extraia os arquivos de habilidade para o seu diretório .claude/skills/.
Baixar ZIPClone o repositório e copie os arquivos da habilidade para o seu projeto.
git clone https://github.com/phuryn/pm-skills/tree/main/pm-data-analytics/skills/ab-test-analysis # Copy SKILL.md to your .claude/skills/ directory
Copiar





Lar
