How Long Should You Run an A/B Test?

You should run an A/B test for at least two full weeks, and keep it running until it reaches the sample size you calculated before starting. For most stores that means two to four weeks. Calling a winner early because the graph looks good is the most common way to get a wrong answer.

Here is how to set the length before you start, and the traps waiting on both ends.

arrested development two weeks meme

Key Points

RuleThe numberWhy
Minimum length2 full weeksWeekday and weekend shoppers convert differently
Sample sizeSet before you startStops you from calling winners early
Maximum length6 to 8 weeksDeleted cookies and seasonality muddy the data
Peeking earlyDo not act on itCan push false positives from 5% to 26%

Why Two Full Weeks Is the Floor

Two full weeks is the floor for an A/B test because your store runs on a weekly cycle. Weekday browsers, weekend buyers, and payday spikes all convert at different rates. One full week captures that cycle once, and two weeks capture it twice.

A test that runs Tuesday to Friday only samples one kind of shopper. The people comparing options on a Wednesday lunch break behave differently from the ones buying on a Sunday night, and those customer behavior signals are exactly what your test is trying to average out.

So always end a test on the same weekday you started it. Run whole weeks, never parts of them.

Two weeks is a floor, not a target. If your sample size math says four weeks, four weeks it is. The floor only exists so a lucky weekend never gets to call itself a trend.

Also, pick your start date with the calendar open. A test that needs four clean weeks should not start two weeks before your biggest sale of the quarter, because promo traffic behaves like a different species and neither half of the data will match the other.

Why Stopping a Test Early Backfires

Stopping a test early turns your significance number into fiction. Evan Miller showed that checking daily and stopping the moment one side looks significant can push a 5% false positive rate to 26%. That is a 1 in 4 chance of shipping a change that does nothing.

The math in How Not To Run an A/B Test is worth reading once in your life. Significance testing assumes you picked the sample size ahead of time. Peek and stop early, and that assumption dies quietly. Spotify’s experimentation team found even casual repeated checking can roughly double false positives, which is why they built sequential testing to make monitoring safe.

Decide the sample size before the test starts, then let it finish. Some of the best A/B testing tools now run sequential statistics for you, so you can watch the dashboard without wrecking the result.

Looking is fine. Acting is the problem. Check for bugs and broken tracking as often as you like, and save the winner-picking for the finish line you drew on day one.

How to Calculate Your Test Length

Calculate your test length by dividing the sample you need by your daily traffic. A store converting at 3% needs roughly 50,000 visitors per variation to spot a 10% lift. At 2,000 visitors a day, that two-way test runs close to eight weeks.

Plug your own numbers into a sample size calculator before you build anything. Ten minutes of math tells you whether the test you are excited about is a two-week project or a five-month one.

If the answer comes back in months, that is its own answer. You do not have a testing problem, you have a traffic problem, and A/B testing with low traffic calls for a different playbook. Test length is math you do before the test, not a feeling you get during it.

Two inputs move that math the most. A higher baseline conversion rate shortens the test, and chasing a bigger lift shortens it a lot. Aiming at a 20% lift instead of 10% cuts the needed sample to roughly a quarter, which is why bold changes are the cheap ones to test.

When to Run Longer, and When to Pull the Plug

Run longer when something temporary might be inflating the result, like a novelty effect or a holiday sitting in the middle of your data. Pull the plug at six to eight weeks, or immediately if the variation is clearly broken and costing real sales.

Novelty cuts both ways. Returning visitors poke at anything new, which flatters the variation for a week or two before the effect fades. Ron Kohavi’s rules of thumb for experimenters add a sobering one: most real wins are small, so a huge early lift is more likely a fluke than a breakthrough.

At the other end, past eight weeks, deleted cookies mix your groups together and the season quietly changes under the test. I have killed tests I was excited about because Black Friday walked into the middle of them. Log what you learned in your CRO audit notes and move on.

The one exception to patience is damage control. A variation that breaks checkout on one browser or tanks revenue from day one does not deserve its full run. Kill it, fix it, and relaunch it as a fresh test rather than patching it mid-flight.

Want Tests That Actually Settle Arguments?

Test length is not a judgment call. Two full weeks minimum, the sample size you calculated up front, and a hard stop before cookie churn and seasonality blur the answer. Everything else is peeking with extra steps.

Book a call and I’ll look at your traffic and conversion numbers, then tell you honestly whether testing is worth running yet and exactly how long your first test should take.

Sources:

Similar Posts