A/B Testing: Hypothesis, Sample, and Stopping Rules
An A/B test is a controlled comparison of Variant A and Variant B. Write the decision, primary metric, sample, and stopping rule before anyone sees results.
A/B testing compares Variant A with Variant B under a written plan so a team can keep, ship, or reject a change. Split eligible units at random. Measure one primary outcome. Ignore early charts until the stopping rule is met.
Use it inside conversion rate optimization when a specific page, offer, or ad decision is on the line. Pair the primary metric with click-through rate or conversion rate only when that number is the decision the test is meant to change.
Eight fields for a test plan written before launch.
Experiment brief
Fill this before launch. A complete brief makes the eventual ship, extend, or revert decision interpretable.
| Field | What to record | Failure mode |
|---|---|---|
| Hypothesis | If we change X for this audience, primary metric Y will move because Z | Vague “make it better” goals |
| Unit | User, session, cookie, account, or geo | Mixing units double-counts people |
| Randomization | How assignment happens and what stays stable | Self-selection, leftover targeting, or shared devices |
| Sample | Traffic needed to detect the smallest useful effect | Ending at a lucky day |
| Primary metric | One number the decision will follow | Ranking several metrics after the fact |
| Guardrails | Revenue, speed, complaint rate, or lead quality that must not collapse | A “win” that damages a downstream number |
| Stopping rule | Planned sample or end date, plus what happens if the interval includes zero | Peeking, then stopping when the chart looks good |
| Interpretation | Ship, keep testing, or revert, including an inconclusive interval | Treating a noisy gap as proof |
Google Ads custom experiments are one product path: they split campaign traffic and budget between an original and a trial campaign. A 50/50 split controls auction eligibility. It does not guarantee equal impressions or spend. Google asks teams to allow 7 to 14 days for the treatment arm to stabilize and not to treat learning-period noise as a verdict. That is platform setup, not a universal calendar for website tests.
Which design matches the question?
| Design | What changes | Use when |
|---|---|---|
| A/B test | One focused change between variants | You need to isolate a specific decision |
| Multivariate test | Several elements and combinations | Traffic is large enough to estimate interactions |
| Holdout | Exposure versus no exposure | You need incremental impact, not a creative winner |
A holdout is incrementality. An A/B test of two headlines is not.
Worked hypothetical result
Illustrative numbers for a checkout-button test. Not a benchmark and not an Anderson Collaborative case.
| Arm | Eligible users | Orders | Conversion rate |
|---|---|---|---|
| Control (A) | 4,000 | 160 | 4.00% |
| Variant (B) | 4,000 | 192 | 4.80% |
Observed gap: 0.80 percentage points, or 20% relative to control.
A two-proportion standard error for these rates is about 0.46 percentage points. A rough 95% interval is about -0.10 to +1.70 percentage points. That interval includes no difference. The 20% relative lift is therefore still compatible with noise at this sample.
Business value is a separate question. If a true 0.80 point gain held on 50,000 monthly eligible users, that would be about 400 extra orders. If the true gain were near zero, shipping B would add implementation cost for no extra orders. The interval does not yet tell you which world you are in, so the stopping rule should say “extend or keep A,” not “B won.”
Do not stop because day four looked decisive. Do not stack extra metrics until one of them is significant. If headline, button, and price all changed, the result estimates that bundle; it cannot identify which component caused the change.
Practical check
Write the ship / hold / revert rule before launch. Compare variants only after the planned sample. When several elements changed together, make only a bundle-level decision. Check whether a guardrail such as speed, refunds, or qualified leads moved the wrong way.
FAQs
What is A/B testing in marketing?
A/B testing randomly splits eligible traffic between Variant A and Variant B, then compares one prechosen primary metric. The result is usable only when the unit, sample, and stopping rule were set before launch.
How long should an A/B test run?
There is no universal calendar length. Duration follows the sample needed for the smallest effect worth acting on, plus enough time to cover weekday and weekend mixes. Stop at the planned sample or date, not because an early chart looks favorable.
What should you write before launching an A/B test?
Write the hypothesis, randomization unit, sample, primary metric, guardrail metrics, and stopping rule. If several elements change together, the test estimates the bundle’s effect; it cannot isolate which element caused the difference.
How is an A/B test different from a holdout?
An A/B test compares two versions of an experience. A holdout withholds the activity from a control group to estimate incremental impact. Google Ads custom experiments split campaign traffic. Conversion Lift is a separate incrementality product.
PUT THIS KNOWLEDGE TO WORK
NEED MORE HELP?
Talk with our team about applying A/B Testing to your marketing.
Get a free marketing audit call