A holdout test that can’t fool you.
How to test whether server-side conversions really cut your costs: two separate sides, a card locked before day one, and enough conversions to see the answer.
Every server-side tracking tool says it makes your ads cheaper. Some publish uplift figures averaged across their customers. None of that tells you whether it works on your account.
A holdout test can. Run half your spend with server-side conversions and half without, and compare what each half actually bought. It sounds simple. There are about five ways to get it wrong, and most of them make the tool look better than it is.
The one rule: score it on your numbers
Never judge the test on the ad platform's own reporting. The side you feed server-side conversions to will always report more conversions, because you've sent it more. That's the feed working, not the ads working.
Score both sides on the same thing, counted the same way, outside the platform: real cost per real outcome, from your payments, your database or your CRM.
Keep the two sides properly apart
This is where most tests quietly fail. A server feed doesn't teach one campaign. It teaches the pixel, or the dataset, behind every campaign that uses it. Run both sides on one pixel and the control side learns from the test side's data, and the test means nothing.
- Side A gets the pixel plus the server feed, on its own dataset.
- Side B gets the pixel only, on a separate dataset. It's the control: exactly what you run today.
- Split by region, with regions that don't overlap, so one person can't be in both.
- Same budget share, same creative, same offer on both sides, for the whole test.
Regions are never identical. So record both sides for a couple of weeks before you switch the feed on. What you're testing is whether the ratio between them moves, not whether one region beats another.
Lock the card before day one
Write four things down, date them, and don't change them once the test starts.
- The metric. Real cost per real outcome, scored on your backend. One metric, not a menu you choose from afterwards.
- The length. At least four weeks. The first two don't count, while the platform is still learning from the new data.
- The threshold. What counts as a win, in numbers, before you've seen any.
- The timestamp. Locked and dated, and shared with whoever has to believe the result.
The card is what stops everyone, including you, from moving the goalposts after seeing the results.
How many conversions you need
With few conversions, chance swamps everything. A test can only prove a lift bigger than the noise. The table shows roughly the smallest lift you can prove, for a given number of conversions on each side.
| Conversions per side | Smallest lift you can prove |
|---|---|
| 10 | ~250% |
| 25 | ~121% |
| 50 | ~75% |
| 100 | ~49% |
| 200 | ~32% |
| 400 | ~22% |
| 800 | ~15% |
| 1600 | ~10% |
The maths, for anyone who wants to check it: treat each side's conversions as a count, compare them as a ratio, and ask for 80% power at 5% significance. The smallest detectable lift is then about e^(2.8 × √(2 ÷ n)) − 1, where n is the conversions on each side. It assumes both sides get equal spend and the conversions are independent. Real accounts are messier, so treat the table as a floor.
Two things follow. A small account can only prove a big lift, so run longer. Or measure an earlier step with more volume, such as leads or add-to-baskets, alongside the sale.
Five ways to fool yourself
- Scoring on platform numbers. The fed side always looks better there. Use your backend.
- One shared pixel. The control learns from the test. Separate datasets, separate regions.
- Peeking and stopping early. Check the test daily and stop the first time it looks good, and you'll "find" wins that are noise. Run the full length on the card.
- Changing things mid-test. A new creative, a budget shift or a sale on one side breaks the comparison. If you must change something, change it on both sides on the same day.
- Counting the learning phase. The first weeks measure the platform adjusting, not the steady state.
Publish it either way
A proof page that only shows wins isn't proof. If you're an agency running this for a client, agree up front that the result gets reported whatever it says. A null result is still worth having: it tells you where not to spend.
Where Clickfall fits. This is the test every Clickfall founding client gets, on their own account, with the card locked before day one and the result published, good or bad. The receipt for every send, which lets you audit side A, is in build.
Sources
- Meta Business Help Center: Differences between event counts in Ads Manager, Ads Reporting and Events Manager
- Meta for Developers: Conversions API
- Meta Business Help Center: About attribution models and attribution settings
Every source was checked on 3 Oct 2026. Spotted something out of date? Tell us and we'll correct it.