Skip to content
Anderson Collaborative

Home › Knowledge Base ›A/B Testing: Hypothesis, Sample, and Stopping Rules

A/B Testing: Hypothesis, Sample, and Stopping Rules

An A/B test is a controlled comparison of Variant A and Variant B. Write the decision, primary metric, sample, and stopping rule before anyone sees results.

Updated September 20, 2026· 6 min read

A/B testing compares Variant A with Variant B under a written plan so a team can keep, ship, or reject a change. Split eligible units at random. Measure one primary outcome. Ignore early charts until the stopping rule is met.

Use it inside conversion rate optimization when a specific page, offer, or ad decision is on the line. Pair the primary metric with click-through rate or conversion rate only when that number is the decision the test is meant to change.

Eight-field A/B experiment brief covering hypothesis, unit, randomization, sample, primary metric, guardrails, stopping rule, and interpretation

Eight fields for a test plan written before launch.

Experiment brief

Fill this before launch. A complete brief makes the eventual ship, extend, or revert decision interpretable.

FieldWhat to recordFailure mode
HypothesisIf we change X for this audience, primary metric Y will move because ZVague “make it better” goals
UnitUser, session, cookie, account, or geoMixing units double-counts people
RandomizationHow assignment happens and what stays stableSelf-selection, leftover targeting, or shared devices
SampleTraffic needed to detect the smallest useful effectEnding at a lucky day
Primary metricOne number the decision will followRanking several metrics after the fact
GuardrailsRevenue, speed, complaint rate, or lead quality that must not collapseA “win” that damages a downstream number
Stopping rulePlanned sample or end date, plus what happens if the interval includes zeroPeeking, then stopping when the chart looks good
InterpretationShip, keep testing, or revert, including an inconclusive intervalTreating a noisy gap as proof

Google Ads custom experiments are one product path: they split campaign traffic and budget between an original and a trial campaign. A 50/50 split controls auction eligibility. It does not guarantee equal impressions or spend. Google asks teams to allow 7 to 14 days for the treatment arm to stabilize and not to treat learning-period noise as a verdict. That is platform setup, not a universal calendar for website tests.

Which design matches the question?

DesignWhat changesUse when
A/B testOne focused change between variantsYou need to isolate a specific decision
Multivariate testSeveral elements and combinationsTraffic is large enough to estimate interactions
HoldoutExposure versus no exposureYou need incremental impact, not a creative winner

A holdout is incrementality. An A/B test of two headlines is not.

Worked hypothetical result

Illustrative numbers for a checkout-button test. Not a benchmark and not an Anderson Collaborative case.

ArmEligible usersOrdersConversion rate
Control (A)4,0001604.00%
Variant (B)4,0001924.80%

Observed gap: 0.80 percentage points, or 20% relative to control.

A two-proportion standard error for these rates is about 0.46 percentage points. A rough 95% interval is about -0.10 to +1.70 percentage points. That interval includes no difference. The 20% relative lift is therefore still compatible with noise at this sample.

Business value is a separate question. If a true 0.80 point gain held on 50,000 monthly eligible users, that would be about 400 extra orders. If the true gain were near zero, shipping B would add implementation cost for no extra orders. The interval does not yet tell you which world you are in, so the stopping rule should say “extend or keep A,” not “B won.”

Do not stop because day four looked decisive. Do not stack extra metrics until one of them is significant. If headline, button, and price all changed, the result estimates that bundle; it cannot identify which component caused the change.

Practical check

Write the ship / hold / revert rule before launch. Compare variants only after the planned sample. When several elements changed together, make only a bundle-level decision. Check whether a guardrail such as speed, refunds, or qualified leads moved the wrong way.

FAQs

What is A/B testing in marketing?

A/B testing randomly splits eligible traffic between Variant A and Variant B, then compares one prechosen primary metric. The result is usable only when the unit, sample, and stopping rule were set before launch.

How long should an A/B test run?

There is no universal calendar length. Duration follows the sample needed for the smallest effect worth acting on, plus enough time to cover weekday and weekend mixes. Stop at the planned sample or date, not because an early chart looks favorable.

What should you write before launching an A/B test?

Write the hypothesis, randomization unit, sample, primary metric, guardrail metrics, and stopping rule. If several elements change together, the test estimates the bundle’s effect; it cannot isolate which element caused the difference.

How is an A/B test different from a holdout?

An A/B test compares two versions of an experience. A holdout withholds the activity from a control group to estimate incremental impact. Google Ads custom experiments split campaign traffic. Conversion Lift is a separate incrementality product.

PUT THIS KNOWLEDGE TO WORK

NEED MORE HELP?

Talk with our team about applying A/B Testing to your marketing.

Get a free marketing audit call