Insights · testing and measurement

What a failed A/B test told us about who actually buys

A losing average can bury a winning segment. Cut every test by list membership and flow engagement before you call it.

A smiling woman with curly hair in a black jumper sits in a teal armchair by a bright window, working on a laptop
Theo Tziapouras, founder and strategy at Engage Commerce
Theo TziapourasFounder and strategy
31 August 2026 1,319 words7 minute read
In short · six parts

Variant B lost and the panel said archive it. Then we cut the result by email engagement, and the bundle had been winning all along with the people one flow had already warmed up. The test was fine. The audience was wrong.

A Klaviyo top performing flows report, each row naming a flow and its trigger beside a live status pill, a deliveries column and a placed order revenue column

The bundle lost to control, until we asked who was in the room

Variant B loses. The bundle offer sits behind control in the panel, and the test's run long enough to call. Rulebook says archive the hypothesis, and that's where most tests stop.

Before we killed this one we split the result by traffic source, then by list membership, then by whether the visitor had touched a recent send. Against people who'd opened one specific flow in the weeks before landing, the bundle beat control by a wide margin. Against everyone else it dragged the blended average down far enough to bury that win. One test, two audiences, two verdicts, flattened into a single losing number.

That flattening isn't a quirk of one tool. Every testing platform does it by default, because the top line is the number it was built to lead with. The average was honest. It was also useless.

Segment the loss before you accept it

A blended result averages behaviours that may have nothing to do with each other. When a test loses on the top line, look at who was in the sample before you accept what that top line says. A losing variant with one segment well above control isn't a bad idea. It's an idea pointed at the wrong default audience.

Most conversion rate optimisation platforms put the blended lift front and centre, and bury the segment views a few clicks deeper if they exist at all. Miss that screen and you'll keep killing good ideas for being aimed at the wrong group.

  • New visitors versus returning visitors
  • Subscribed to the list versus never captured
  • Opened or clicked a flow recently versus no email contact at all
  • Touched a recent campaign versus untouched by anything you sent

Work down those cuts before you retire anything, and stop when one of them splits the verdict. Every segment agrees it lost? Retire it with a clear conscience. If they disagree, you've just found the real result.

Cross reference the test against email engagement, not the page alone

CRO tools measure the page. Email platforms measure the inbox. The two rarely sit in one view, and that gap is exactly where most teams stop looking, so we pulled the session level test data into the same table as Klaviyo flow engagement and went looking for overlap.

The winning segment came down to something no onsite test could see on its own. The people the bundle converted had opened a flow that already explained why those products belonged together, days before they reached the page. Context built in the inbox, cashed in on the site.

Onsite testing tells you what happened at the moment of decision, and nothing about what led up to it. Join the two data sets yourself. No dashboard is going to do it for you.

Turn the winning slice into a campaign, not a footnote

Once you know which segment a test actually won with, keep what worked instead of filing the lot as a loss. Here the bundle stayed live for the flow engaged segment only, as an onsite personalisation rule with a suppression on everyone else. The losing majority never saw it again.

The better move came second. The flow gained a step built to carry people toward the bundle while their attention was still warm, and a one off site experiment turned into a standing piece of lifecycle work. The test stopped being a failure. It became a targeting brief.

A woman in glasses and a light denim shirt perches on a desk edge, reading her phone with a cream mug in one hand

An offer that works for warm leads is too vague to build anything with. The version worth writing into the brief names the flow, the point in its sequence, and the offer that converts the group it produces. Specific enough to build a campaign around, and specific enough to test again.

Make the check a standing step, not a special investigation

None of this gets found while CRO and email sit on separate scorecards, in separate meetings, owned by people who never compare notes. The test data existed. The flow engagement data existed. Nobody had built the habit of putting them in one place before calling a result.

  1. Pull the result by segment before you read the headline lift
  2. Cross reference every flat or losing segment against recent flow and campaign engagement for the same people
  3. Treat a segment that beat control as the finding, not as noise inside a failed test
  4. Build the winning segment into a campaign or a new flow step
  5. Retire the idea only if it loses inside its best segment too

Run that list before any test gets marked won or lost. Every time. You won't know a result is surprising until after you've looked, so the habit has to fire on the boring ones too. It's how we read tests across our client accounts, and it's the cheapest audit you'll ever run.

Where slicing a test stops being honest

Here's the hard truth: this only holds if the sample survives the slicing. Cut a modest test into narrow segments and you can talk yourself into a winning pocket that's really noise, because a small enough slice will always tell you a story. The numbers underneath it won't hold.

So the discipline is patience. Decent traffic, or tests left running until each segment you care about is big enough to trust on its own. Treat any segment verdict as provisional until the volume turns up, and rerun the winner inside its own segment before you spend real budget on it.

Theo Tziapouras, founder and strategy at Engage Commerce

Theo Tziapouras

Founder and strategy at Engage Commerce, the ecommerce agency for 7 and 8 figure DTC brands.

Should I segment A/B test results before calling a winner?

Yes, and before calling a loser too. A blended average can hide one segment beating control underneath several that didn't, so pull the result by new versus returning, by list membership and by recent email engagement before you act on the top line. If every segment agrees, call it. If they disagree, the disagreement is the finding.

Why did my A/B test fail when the idea seemed strong?

Often the idea was pointed at the wrong default audience, not wrong in itself. An offer that needs context, a bundle or a premium variant, tends to convert people who already understand the product and confuse people who don't. Check whether the variant won inside your warmest segment before you archive it.

How do I connect CRO test data with Klaviyo engagement?

Export the test's visitor level results with whatever identifier your testing tool holds, usually an email or a profile ID for logged in and subscribed visitors, then join it against flow and campaign engagement in Klaviyo over the same window. A spreadsheet is enough. You're looking for whether the variant's converters cluster inside one engaged group.

When is a winning segment in a failed test just noise?

When the slice is too small to trust. Narrow segments throw up flattering pockets by chance, and the more cuts you make the likelier one looks like a win. Treat a segment result as provisional, rerun the test inside that segment alone, and build on it only once the pattern repeats at a size you'd accept for a normal test.

Not got your answer?Chat to us
End matter

The most useful failed test is the one you read properly

A failed A/B test is only a verdict on the average, and averages don't buy anything. Segments do. The brands that get value out of testing aren't the ones running the most experiments, they're the ones reading every result against who was actually in it. Somewhere in your archive of losing tests there's probably a winner aimed at the wrong room.

Theo TziapourasFounder and strategy · Engage Commerce

Would you rather this was just handled?

Bring your Klaviyo account and the thing annoying you most. We will tell you what we would fix first, on the call, before you spend anything.

Engage CommerceTheo Tziapouras, founder of Engage Commerce

Book a call with our founder.

We're all about relationships built on trust, mutual respect and a shared vision for success. If that sounds like your vibe, let's make some waves together 🌊