A/B testing: why most tests prove nothing
An A/B test compares two versions of the same page to determine which one produces the better result. Its difficulty is not technical but statistical, and it comes at a cost: most tests run in small and mid-sized businesses never gather enough traffic to tell a real effect from a fluctuation, yet they are declared conclusive, then billed.
- The number of conversions you need is calculated before the test launches, never after.
- Stopping a test the moment one variant pulls ahead is the most common way to get it wrong.
- A test must run over whole weeks to absorb day-of-week and cycle variations.
- Below a certain volume, applying best practices beats testing. It also costs far less.
- A poorly designed test does not give you zero information, it gives you false information, and costly decisions follow from it.
On this page
A/B test
An A/B test is an experiment in which two versions of the same page or message are shown simultaneously to randomly assigned groups of visitors, in order to compare their conversion rate. Its validity rests on three conditions: random assignment, a volume large enough that the observed gap is not due to chance, and a duration covering a full cycle of behaviour.
The question of volume, which decides everything
For an executive team, the real stake in a testing program is not the conversion gained, it is the risk of paying for false conclusions. An underpowered test produces an apparent winner, the team applies it, and the company decides on noise. The calculation that prevents this takes ten minutes and happens before you spend a dollar.
If you have ever been shown a "conclusive" test followed by no improvement in sales, this is almost always what happened. A test is run like a quality control: without a sufficient sample size, the result is not debatable, it is simply invalid.
This is the point where most testing efforts fail, and it is settled before you begin.
Detecting a gap takes more observations the smaller the gap is. Going from 2% to 4% conversion can be detected with little traffic, because the effect is enormous. Going from 2% to 2.2% requires more than a hundred thousand visitors per variant, because the effect is the same order of magnitude as the noise.
Yet realistic gains almost always sit on the side of small gaps. This mismatch is what explains why so many tests are launched without ever being able to conclude, and why the budget they consume would have been better spent elsewhere.
| Starting rate | Effect to detect | Order of magnitude needed per variant |
|---|---|---|
| 2% | Doubling, toward 4% | A few thousand visitors |
| 2% | 25% lift, toward 2.5% | A few tens of thousands |
| 2% | 10% lift, toward 2.2% | More than a hundred thousand |
| 10% | 10% lift, toward 11% | A few tens of thousands |
The calculation is done with a sample size calculator, before launching. If the number of visitors required exceeds what the page receives in two months, the test will not happen. Better to know that before configuring the tool.
What to check before authorizing a testing program
A credible test requires two whole weeks at the minimum and a number of conversions calculated in advance. Below that threshold, it costs more than time: it produces false conclusions on which the company then commits decisions. The Baymard Institute measures up to 35% conversion lift attainable through funnel design alone, with no test at all. That is almost always the best starting point.
- How many conversions per week does this page produce, and how many would it take to conclude?
- Was the sample size calculation done before launch, and by whom?
- What are we measuring the win on, the click or the sale?
- How many tests launched last year led to a change we kept?
- Are there known fixes to apply before testing anything at all?
A solid answer gives a number of conversions needed, calculated in advance, and is willing to say that a test is impossible. A weak answer offers to test button colours, or promises a percentage gain before anything has been measured.
Deciding whether your volume allows testing, or whether it is better to fix without testing, is settled in one meeting. A 90-minute consultation settles it on your real numbers, with a written summary that circulates through your organization.
The method mistakes that invalidate a test
Stopping the moment one variant pulls ahead
This is the most frequent and most costly mistake. Early in a test, gaps are enormous and purely random. Checking the results every day and stopping at the first favourable gap guarantees you conclude on noise.
Testing less than a full week
Buying behaviour varies sharply by day. A test launched on a Tuesday and stopped on a Friday compares two different populations. The rule is to cover whole weeks, two at the minimum.
Testing several elements at once
Changing the headline, the button and the image at the same time produces a result with no explanation. You will know that version B wins, never why, so nothing will be transferable elsewhere.
Measuring the click rather than the sale
A variant can increase clicks on a button and reduce sales, if it draws less qualified visitors. The success metric must be the final commercial action.
Running several tests on the same path
Two simultaneous tests on successive pages contaminate each other. The results become uninterpretable and no one notices.
A test declared conclusive at 95% confidence is wrong one time in twenty. Across twenty tests run in a year, one result is false on average, and nothing sets it apart from the others. That is a reason to retest important changes, not to give up on testing, whose cost stays below that of a redesign decided blind.
What to test first
The order follows expected revenue, not ease of implementation. Testing a detail on a low-traffic page costs weeks for no gain.
The order matters, because available volume is limited and each test consumes a share of it.
| Rank | Element | Expected effect size |
|---|---|---|
| 1 | The main value proposition and the headline | Strong |
| 2 | The number of fields in a form | Strong |
| 3 | The presence and placement of proof | Medium to strong |
| 4 | The structure of prices and options | Medium to strong |
| 5 | The wording of the call to action | Low to medium |
| 6 | The colour of a button | Negligible |
The last row deserves to be said plainly. Button-colour tests have circulated for fifteen years as the canonical example, and they produce effects so small that no SMB site has the volume to detect them. Time spent on this kind of test is time taken from the first four rows.
What stays in-house is the definition of what counts as success. A conversion is not the same thing depending on the product margin, the length of the sales cycle or the customer's value over three years, and only the company holds those figures. A test optimized on the wrong event loses money with full statistical rigour. What can be delegated is the rest: power calculation, instrumentation, segment splitting, reading the results and the reasoned refusal of impossible tests. It is this last point that sets a good provider apart.
When you lack the traffic to test
This is the situation of most businesses, and there are alternatives that produce more than inconclusive tests.
| Method | What it gives you | Volume required |
|---|---|---|
| Applying documented best practices | Known gains, without local validation | None |
| Usability tests with five people | The major obstacles, explained | None |
| Session recordings | Where people get stuck, without knowing why | Low |
| Interviews with recent customers | The real objection, in their words | None |
| Before-and-after test, with no variant | A clue, never proof | Low |
The second row is by far the most profitable below a certain volume. Watching five people try to complete a task on your site almost always reveals obstacles no A/B test would have identified, because a test tells you that one version loses, never why.
For reference, the Baymard Institute estimates that a large eCommerce site can gain on average 35% in conversion rate through the sole application of documented design principles, with no prior test.
An A/B test tells you which of the two versions wins. It never tells you why. It is a decision tool, not an understanding tool.
Falia analysis gridBefore launching a test
The elements to test first are described in what makes a landing page convert. The decision mechanisms that explain the results are detailed in the drivers of an online decision. The checkout case is covered in cart abandonment. Execution lives on the CRO agency page.
The general order of conversion fixes is set out in what conversion rate optimization really fixes.
Improving the conversion rate is at the heart of the Improve your site's conversion goal.
Already running a marketing team? See how we plug in as reinforcement on conversion rate optimization.
Frequently asked questions about A/B testing
What is an A/B test?
It is an experiment in which two versions of the same page are shown simultaneously to randomly assigned groups of visitors, in order to compare their conversion rate. Its validity rests on random assignment, a sufficient volume and a duration covering a full cycle of behaviour.
How many visitors do you need for an A/B test?
It depends entirely on the size of the effect you want to detect. Doubling a 2% rate can be detected with a few thousand visitors. Taking it to 2.2% requires more than a hundred thousand per variant. The calculation is done before launching, with a sample size calculator.
How long should a test run?
At least two whole weeks, to absorb behaviour variations by day of the week. A test launched on a Tuesday and stopped on a Friday compares two different populations, whatever the volume reached.
Can you stop a test as soon as one version wins?
No, and it is the most frequent mistake. Early in a test, gaps are large and purely random. Checking the results every day and stopping at the first favourable gap guarantees you conclude on statistical noise.
What if I do not have enough traffic?
Apply documented design principles rather than test, and watch five people use your site. The Baymard Institute estimates that a large eCommerce site can gain on average 35% in conversion through the sole application of these principles, with no prior test.
Does the colour of a button change anything?
Very little. The effect is so small that no SMB site has the volume needed to detect it. Time spent on this kind of test is taken from the elements that truly matter: the value proposition, the number of fields and the placement of proof.
- Baymard Institute, research on checkout usability, accessed July 2026.
- Jakob Nielsen, Nielsen Norman Group, Why You Only Need to Test with 5 Users, accessed July 2026.
- Google Analytics Help, Google Optimize sunset, accessed July 2026.

Gabriel almost always takes your first call and carries out your audit. He builds the strategy starting from your growth goal: where to put your budget, which market to test and how to connect each lead to a real sale in your CRM. He mainly leads engagements for three goals: Optimize the profitability of your digital campaigns, Develop a new market, and Generate demand and growth. With Geneviève, he also works on organic search (SEO), AI visibility (GEO) and conversion rate optimization (CRO). The sales a Google Ads or Meta Ads campaign brings in depend on the page that receives the click. He writes mainly about marketing strategy, paid advertising and measurement.
About Falia →