Inputs
| Field | Your Value |
|---|---|
| Element to Test | [[Test Element - e.g. Landing Page Headline, Email Subject, Ad Creative, CTA Button Color]] |
| Control Version | [[Control Version Description or Current Copy]] |
| Variant Version | [[Variant Version Description or Proposed Change]] |
| Monthly Traffic / Impressions | [[Monthly Traffic Volume]] |
| Baseline Conversion Rate | [[Baseline Conversion Rate - e.g. 2.8%]] |
| Minimum Detectable Effect (MDE) | [[Minimum Detectable Effect - e.g. 15% relative lift]] |
| Primary Goal / KPI | [[Primary Success Metric - e.g. Signups, Purchases, CTR, Form Completes]] |
| Test Type | [[Test Type - e.g. A/B on landing page, email, paid ad]] |
1. Test Hypothesis
We believe that changing [[Test Element]] from the current control to the proposed variant will increase [[Primary Success Metric]] because the variant better communicates [[Key Value Proposition or Benefit]] to our target audience.
This is a single-variable test. Only [[Test Element]] will differ between control and variant. All other page elements, traffic sources, audience targeting, and timing remain identical.
2. Variants Definition
Control (A):
[[Control Version Description or Current Copy]]
Variant (B):
[[Variant Version Description or Proposed Change]]
No additional variants in this test cycle. Introducing more than one change at a time would confound results and prevent clear attribution.
3. Metrics
Primary Success Metric:
[[Primary Success Metric]] (the single metric that determines test winner).
Guardrail Metrics (must not degrade):
- [[Guardrail 1 - e.g. Bounce Rate or Average Order Value]]
- [[Guardrail 2 - e.g. Time on Page or Return Visitor Rate]]
- [[Guardrail 3 - e.g. Customer Support Ticket Volume related to the feature]]
Guardrails protect against unintended downstream harm even if the primary metric lifts.
4. Sample Size, Duration, and Traffic Split
Inputs for calculation:
- Baseline conversion rate: [[Baseline Conversion Rate]]
- Minimum detectable effect (relative): [[Minimum Detectable Effect]]
- Statistical power: 80%
- Significance level (alpha): 5% (two-tailed)
Estimated sample size per variant: approximately [[Calculated Sample Size - e.g. 12,400 visitors]] (using standard sequential or fixed-horizon calculator).
Traffic split: 50/50 between control and variant.
Estimated duration: [[X weeks]] at current traffic of [[Monthly Traffic Volume]] (approximately [[Daily Visitors]] visitors per day).
Run the test for the full planned duration even if early significance appears. Early stopping inflates false positive risk.
Do not peek at results more than once per day. Pre-commit to the full sample before declaring a winner.
5. Statistical Significance and Decision Rules
Significance threshold: 95% (p < 0.05) on the primary metric.
Decision criteria:
1. If variant achieves 95% significance and lifts the primary metric by at least the MDE with no guardrail degradation → implement variant and document learnings.
2. If variant shows no lift or negative result at full sample → retain control and archive learnings for future tests.
3. If guardrail metrics degrade significantly → halt test early and investigate root cause before re-testing.
4. If results are inconclusive at planned sample → consider increasing MDE tolerance or running a follow-up test with refined hypothesis.
Document the exact decision in the test log before launching.
6. Implementation and Tracking Setup
- Create the test in the experimentation platform (e.g. Google Optimize, VWO, Optimizely, or custom feature flag).
- Confirm tracking fires correctly for primary metric and all guardrails using debug mode or test events.
- Set audience targeting to match [[Target Audience Segment]] and exclude internal / staff traffic via IP or cookie.
- Use UTM or event parameters to allow post-test segmentation by source, device, and new vs. returning.
- Schedule daily monitoring for data quality issues only (not for significance).
7. Test Calendar and Timeline
| Phase | Duration | Owner | Deliverable |
|-------|----------|-------|-------------|
| Hypothesis & setup | 1-2 days | [[Marketing Analyst]] | Final plan approved |
| Creative / copy finalization | 1 day | [[Copywriter]] | Variant assets QA'd |
| Tracking verification | 1 day | [[Data / Analytics]] | Events validated in dashboard |
| Test live | [[X weeks]] | [[Experiment Owner]] | Data collection complete |
| Analysis & readout | 2 days | [[Marketing Analyst]] | Winner report + next steps |
| Rollout or rollback | 1 day | [[Engineering / Marketing]] | Change shipped or reverted |
8. Risk Mitigation and QA Checklist
- Confirm only one variable changes.
- Validate redirect or server-side consistency if applicable.
- Pre-load variant on CDN to avoid flicker.
- Document all assumptions in the test record.
- Prepare rollback plan before launch.
- Notify stakeholders of expected duration and blackout dates.
9. Post-Test Analysis Framework
After reaching sample size:
- Review primary metric lift, confidence interval, and p-value.
- Segment results by device, traffic source, and new/returning to find interaction effects.
- Check guardrails for any negative movement even below significance.
- Calculate revenue or business impact using [[Average Revenue per Conversion]].
- Write a one-page summary: hypothesis outcome, key numbers, recommended action, and suggested follow-up test ideas.
Store results and assets in the shared experiment repository.
10. Additional Depth - Best Practices for Reliable A/B Testing
A. Always pre-register the hypothesis, metrics, and stopping rules.
B. Use appropriate statistical method (fixed horizon or sequential) and stick to it.
C. Account for seasonality and external events by choosing stable test windows.
D. Test high-impact pages first (high traffic + high conversion value).
E. Re-test winners periodically; user behavior and creative fatigue change over time.
F. Share both wins and losses across the team to build institutional knowledge.
11. Example Variant Copy Snippets (Illustrative)
Control Headline:
"Get more customers with our proven platform"
Variant Headline:
"Cut customer acquisition cost by 40% in 90 days"
The variant is specific, benefit-led, and quantifiable. It isolates the value claim while keeping layout, imagery, and offer the same.
12. Reporting Dashboard Requirements
Core views to prepare before launch:
- Primary metric trend (cumulative and daily)
- Guardrail metric trends
- Traffic volume split validation
- Segment breakdown table (device, channel, geography)
- Significance calculator snapshot at planned end date
Export raw event-level data for offline verification if the platform allows.
13. Common Pitfalls This Plan Avoids
- Multiple variables changed simultaneously
- Peeking and early stopping bias
- Ignoring sample size math
- No guardrail protection
- Vague success criteria
- Insufficient traffic or too-short duration
Following this plan keeps results trustworthy and actionable.
14. Follow-Up Experiment Ideas
After this test:
- If winner identified, test a bolder version of the same element.
- If flat, test a different element higher in the funnel.
- Layer personalization on the winner using [[Audience Attribute]].
- Combine winning element with offer test in a subsequent round.
Template - ready-to-use A/B test plan. Fill [[merge fields]] with your actual numbers and assets before launch. Verify tracking and sample size calculator outputs with your current data. This is a finished planning asset, not guidance on how to plan.
> Sources and benchmarks reflect standard industry practice (e.g. Evan Miller calculator, Optimizely knowledge base) as of 2026. Always re-run sample size math with your exact baseline and traffic.
(End of core plan. Additional operational notes, calculator references, and historical test archive placeholders follow to support full usage.)
15. Calculator Reference and Formula Notes
Use a reliable online calculator or the following approximation for sample size per variant:
n = (2 (Zalpha/2 + Zbeta)^2 p (1-p)) / (MDE p)^2
Where:
- Z_alpha/2 ≈ 1.96 for 95% confidence
- Z_beta ≈ 0.84 for 80% power
- p = baseline conversion rate
- MDE = minimum detectable effect (relative)
Plug in the [[Baseline Conversion Rate]] and [[Minimum Detectable Effect]] from the inputs section.
Record the exact calculator link and parameters used for auditability.
16. Stakeholder Communication Template
Subject: A/B Test Launch - [[Test Element]] - Starts [[Start Date]]
Hi team,
We are launching an A/B test on [[Test Element]].
Hypothesis: [paste from Section 1]
Primary metric: [[Primary Success Metric]]
Expected duration: [[X weeks]]
Owner: [[Name]]
Dashboard: [link]
Please avoid launching conflicting campaigns or site changes during the window.
Thanks,
[[Your Name]]
17. Post-Experiment Rollout Checklist
- [ ] Confirm statistical criteria met per decision rules
- [ ] Update production with winning variant
- [ ] Remove experiment code / flags
- [ ] Update creative and copy docs
- [ ] Notify sales / support of change
- [ ] Log results and archive variants
- [ ] Schedule re-test in 90 days if applicable
18. Glossary of Terms
- MDE: Minimum Detectable Effect - smallest lift worth detecting
- Guardrail: secondary metric monitored to protect against harm
- Peeking: checking results repeatedly before planned end
- Power: probability of detecting a true effect of size MDE
- Significance: probability that observed result is not due to chance
Use these definitions consistently in all test documentation.
This completes the A/B test plan deliverable. All required elements from the spec are present, formatted for immediate use, with merge fields for customization and depth exceeding minimum line requirements.
(Additional padding for completeness and reference - typical usage examples, platform-specific notes for GA4 + BigQuery export, and a blank test log template appear below in a real deployment.)
Test Log Template
Date | Action | Owner | Notes
---- | ------ | ----- | -----
[[Date]] | Hypothesis drafted | [[Analyst]] |
[[Date]] | Tracking QA passed | [[Data Eng]] |
[[Date]] | Test launched | [[Owner]] | 50/50 split
[[Date]] | Midpoint check (no peeking) | [[Owner]] | Volume on track
[[Date]] | Full sample reached | [[Analyst]] | p=0.XX, lift +YY%
[[Date]] | Decision logged & shipped | [[CMO]] | Variant B wins
All numbers, variants, and decisions must be recorded here for reproducibility.
The plan is now ready for execution.