← All posts
Experiments

Holdouts: the experiment you should run before any A/B test

Imagine optimizing the boarding music on a flight to the wrong city. That's an A/B test without a holdout: you'll learn, with statistical rigor, which subject line gets more opens — while never asking whether the email program is producing any sales that wouldn't have happened anyway.

The two questions, in order

Question 1 (holdout): does this activity create lift at all? Take your audience, randomly exclude 10–20%, run the activity for everyone else, compare outcomes. If the excluded group buys at the same rate as the exposed group, the activity is decoration. This is the existence question, and it outranks everything.

Question 2 (A/B): given that it works, which version works best? Now subject lines, creatives, and bids are worth optimizing — you've established there's something real to optimize.

Teams skip Question 1 for two human reasons: nobody wants to discover their channel is decoration, and deliberately not-marketing to people feels like leaving money on the table. But the held-out 15% is the cheapest insurance you'll ever buy — it's the difference between “email drove $80K” and “email drove $80K of orders that customers were placing anyway.”

Running one without a data team

  • Size: 10–20% held out is enough for most SMB volumes; below ~200 conversions a month, run it longer rather than bigger.
  • Duration: a full purchase cycle minimum. A 7-day holdout on a 30-day consideration product measures noise.
  • Discipline: decide the read date in advance and don't peek-and-stop. Early lift that vanishes is the most common false positive.
  • Rotation: rotate who's held out so no customer is permanently unmarketed-to.

One more habit separates serious teams: write the prediction down first. “We expect the exposed group to convert 15% higher.” A prediction you record before the data arrives is the only kind that can teach you anything — everything else is narration after the fact.

Where Mayaa fits

In Mayaa, holdouts are the default, not the upsell: every paid campaign launches with a control group, every agent records its prediction before acting, and outcomes are scored against those predictions. Agents that predict well earn autonomy; activities that don't beat their holdout get killed — by the numbers, not by whoever argued loudest in the meeting.

Want this run for you? The team that wrote these posts is the team that does the work.

Start free trial
— Keep reading