Insights · campaign testing

Your Klaviyo AB test sample size was never big enough to crown that winner

The winner banner is built from a few dozen opens in a two hour window. Here is how to run a test worth trusting.

Two women sitting in the sun against a drystone wall, each holding a coloured Thermos flask
Alex Gregoriades, operations director at Engage Commerce
Alex GregoriadesOperations director
31 August 2026 1,419 words7 minute read
In short · six parts

Klaviyo puts a green banner across the top of the campaign report and everybody stops thinking. Behind that banner sit a couple of hundred sleepy inboxes and a timer that ran out. Before you build a strategy on the verdict, count the opens it was actually built on.

A man lies along a grey sofa scrolling his phone, with a kitchen counter, a laptop and shelves of glassware behind him

The green banner is a verdict on two hours, not on your list

The banner says variation B won, and a confidence score sits next to it. That score is the part that convinces people, because a number makes it look like maths happened. Behind it, a tenth of the list was split three ways, a couple of hours passed, and the platform picked a side.

Then the winning subject line went out to everybody else, and somebody wrote always lead with urgency in a Slack message that'll be quoted in a strategy deck six months from now. A green banner isn't evidence, it's a timer running out. The verdict happened quickly, and quickly isn't the same as right.

Count the opens per variant before you trust a winner

Take a 5,000 person list, because it's a completely normal size and the numbers only get worse below it. A test group of a tenth of the list is 500 people, and split across three subject lines that's roughly 167 recipients per variant. If one in four of them opens, your winner is being crowned by somewhere between 35 and 45 opens.

That's barely a sample. It's a focus group you'd never trust for any other decision in the business. A gap of three or four points between two subject lines sits comfortably inside the noise at that size, and the losing line could easily be the better one on a bigger send. Klaviyo's confidence score is honestly calculating win probability from the data you gave it, and that data was never enough to ask the question.

So before the next test, convert your test group into raw opens or clicks per variant and run that number through a proper significance calculator. A dramatic swing might be genuinely detectable on your list. The small difference most subject line tests actually produce needs a sample far beyond what a rushed split can give, and knowing which case you're in is a bit more discipline before you hit send.

The test closes before your list has finished opening

Sample size isn't the only problem, and the next one does more damage. Klaviyo's default test duration is a matter of hours, not because that's statistically sound but because the person running the test wants an answer before lunch. Your list doesn't open email on a schedule that respects the timer.

Work openers open in the first hour. Evening scrollers and people who check email once a day open twelve, twenty-four, sometimes forty-eight hours later. Close the window at hour two and the verdict was built by the most reactive slice of your list, then applied to everybody who behaves nothing like them.

Pull the time-to-open curve from your last five or six campaigns, reading opens over time instead of the final total. If half your opens land after hour six, a two hour test was never going to catch a representative picture. It caught whoever was already holding their phone.

Three Klaviyo campaign rows from one February, each naming a sale send and its audience segments beside a sent pill, the send date and performance columns

The fix costs almost nothing. Set the duration to where your own curve flattens, which for a lot of DTC lists means twelve to twenty-four hours instead of two, and let the winner go out the next morning. You lose a few hours of send timing on one campaign. You get back a verdict that includes the people who were asleep when it went out.

Judge it on clicks or revenue, because Apple opens email for you

Even a perfectly sized, perfectly timed test hands you a false winner if you judge it on opens. Apple's Mail Privacy Protection has been pre-fetching and auto-opening email for a meaningful share of iPhone inboxes since 2021, and those recorded opens say nothing about which subject line was better. A machine fired them.

You'd never crown a landing page test on a metric both variants can hit passively, and the same logic applies inside the inbox. So switch the winning metric to clicks, or on a campaign to placed order rate, which Klaviyo lets you set in the test itself. On a flow it only offers opens or clicks, so somebody reads the placed orders by hand. It just isn't the default, so most accounts never touch it.

What to check before you let the dashboard decide

Five things, run through before Klaviyo runs the numbers for you.

  • Raw sends per variant once the list is actually split, not the headline test group share
  • Whether that number, through a proper significance calculator, could detect a lift you would act on
  • Where your time-to-open curve flattens, and whether the test duration is set to match it
  • Whether the winning metric is opens, clicks or placed orders, given how much of the list reads on Apple Mail
  • Whether the gap between variants clears the margin of error, not just the other variant's number

If a test fails more than one of those, treat the result as a hunch worth trying again, not a decision worth building a strategy on. The check takes ten minutes, and it decides what the banner is allowed to mean.

Fewer tests on bigger splits beat more tests on slivers

Look, the instinct after a shaky result is to run more tests, faster, on smaller slices, and that makes it worse. The boring option fixes it. Run fewer subject line tests, with two variants instead of three because every extra variant shrinks the usable sample, sized to detect a difference that would genuinely change what you send next.

For smaller lists where a proper test group still won't clear a useful sample, and plenty won't, stop treating each send as its own verdict. Run the same subject line structure across several campaigns and compare cumulative performance instead. That builds a sample from real volume over time, instead of letting one rushed afternoon decide your tone of voice for the quarter.

A Klaviyo growth overview on the message type breakdown tab, daily campaign and flow bars running across a month beside a conversion summary

Klaviyo isn't lying to you. It's reporting exactly what happened to 167 people in the first two hours, and you're the one who decided that counts as the truth about the other 4,833. The accounts behind the results we publish run fewer tests than you'd expect, and every one of them was sized before it was sent.

Alex Gregoriades, operations director at Engage Commerce

Alex Gregoriades

Operations director at Engage Commerce, where he runs the accounts the writing comes out of.

What is a good sample size for a Klaviyo AB test?

Work backwards from the decision instead of from a fixed share of the list. Convert your test group into expected opens or clicks per variant, then check it against a significance calculator for the smallest lift you'd act on. If the number can't detect that lift, widen the split or cut a variant.

How long should a Klaviyo AB test run?

Until your own time-to-open curve has flattened, which you can read from opens over time on recent campaign reports. For a lot of DTC lists that means holding the test open overnight instead of accepting the default window of a couple of hours. Time sensitive sends are the honest exception.

Should I pick a winner on opens or clicks?

Placed orders where the volume carries it, clicks where it can't. Apple Mail pre-fetches images and fires opens no human chose, so open based winners partly measure automation instead of interest. Opens are only worth judging on if almost none of your list reads email on an iPhone, and that's rare.

Can I AB test on a small email list?

Yes, but not one send at a time. A small list rarely produces enough opens per variant for a single campaign to prove anything, so repeat the same subject line structure across several campaigns and compare the pattern cumulatively. Volume over time replaces the sample size one send can't supply.

Not got your answer?Chat to us
End matter

Keep testing, stop crowning

The answer to a shaky test isn't to stop testing, it's to stop crowning. Build the sample size, the time window and the winning metric into the test before the banner gets a say, and accept that some campaigns are simply too small to teach you anything on their own. A winner you waited a day for beats a hunch you shipped by lunch.

Alex GregoriadesOperations director · Engage Commerce

Would you rather this was just handled?

Bring your Klaviyo account and the thing annoying you most. We will tell you what we would fix first, on the call, before you spend anything.

Engage CommerceTheo Tziapouras, founder of Engage Commerce

Book a call with our founder.

We're all about relationships built on trust, mutual respect and a shared vision for success. If that sounds like your vibe, let's make some waves together 🌊