Creative testing
Anyone who answers that with a single number is guessing. The honest answer is a division sum, and you already have both numbers it needs.
Ask this in any marketing forum and you will get confident, specific answers. Ten. Twenty. Fifty a month. None of those people know your budget, your cost per result, or what you are optimizing for — so none of those numbers mean anything to you.
What is true is narrower and more useful. Creative winners are rare, so more independent attempts improve your odds. Each attempt needs enough spend behind it to produce a readable result. Those two facts pull in opposite directions, and where they meet is your number.
Everything below is derived from these. There are no benchmark statistics in this article, because the honest ones don't exist.
Nobody picks the winning hook in advance — not you, not your agency, not the creator. That is what makes this a testing problem rather than a writing problem, and it is the entire argument for volume.
A variation that spent twelve dollars did not lose. It did not report. Below some floor the difference between two ads is noise, and acting on noise is worse than not testing — you kill winners and scale flukes.
Meta's published guidance is that an ad set needs roughly 50 optimization events a week to exit the learning phase — the opening period where delivery is unstable while the system works out who to show the ad to. Under that, results move for reasons unrelated to your creative.
Your testing budget divided by the minimum readable spend per variation is the maximum number of variations you can honestly test. Run more than that and you have not tested more things — you have tested nothing properly, and paid for the privilege.
The only hard part is the second number, and it depends on what you are judging on. Judging a hook on whether people keep watching is cheap. Judging an ad on cost per purchase is expensive, because purchases are expensive.
Work it. Say you have $2,000 a month for creative testing and your cost per purchase is $35. To judge a variation on cost per purchase you want it near that 50-conversions-a-week mark: 50 x $35 = $1,750 a week, for one ad set. Your whole month buys roughly one properly powered conversion test. At that budget you cannot test hooks on purchases.
Now judge the same variations on whether people watch. You need a few thousand impressions each for a retention curve to settle. Look up your own CPM — the cost of a thousand impressions, which is in your reporting and specific to you. At $12, fifty dollars buys around 4,000 impressions per variation. Twenty variations is $1,000, and half the month is still free for the survivors.
Same budget, three different questions. The cheaper the question, the more variations you can afford to ask it of.
| What you judge on | Volume one variation needs | Spend per variation | Variations $2,000 buys |
|---|---|---|---|
| Did anyone keep watching? (early) | A few thousand impressions | Your CPM x 3-5 | Fifteen to twenty |
| Did anyone click through? (middle) | A few hundred clicks | Your CPC x ~200 | Four to eight |
| Did anyone buy, and at what price? (late) | ~50 conversions in a week | 50 x your cost per result | One, sometimes none |
None of these are benchmarks. Every figure is your own account's numbers multiplied by a volume requirement, and the 50-conversions-a-week line is Meta's published learning phase guidance rather than a rule of thumb. Substitute your real CPM, CPC and cost per result and the table gives you your answer instead of somebody else's.
Forty variations, because you have forty videos
$50 each
Judged on cost per purchase at a $35 CPA. Most get one conversion or none. Every ranking is noise, and the ad you scale is the one that got lucky.
Forty screened, four tested properly
$50, then $375
Forty hooks judged on watch-through, where $50 is genuinely enough. The four survivors then get real conversion budget. Two decisions, each on data that could support it.
A test only tells you something if you can attribute the difference to a cause. Five videos that differ in hook, creator, pacing, offer and caption all at once produce a winner you cannot reproduce. Work down this list, one layer at a time.
The two failure modes are killing on day one and letting a loser run for a month. Both are common.
Delivery is at its most unstable in the first 24 hours and what you see is close to meaningless. Editing an ad set restarts the learning phase and throws away the data you just paid for.
Each variation should have its few thousand impressions by now. Look at hold rate and nothing else, and kill the bottom of the pack. You are deciding which ads deserve real budget, not which are profitable.
Move the freed-up budget into the three or four that held attention. Cost per result only starts to become readable once the spend behind each one is enough to generate the events.
A week or two of stable delivery on a meaningful budget is when cost per result means something. Scale, iterate or kill here — then start the next round of hooks so you are never waiting on one test.
These are two different questions and people conflate them constantly. Early signal answers did this creative earn attention. Late signal answers did that attention turn into money. You need both in that order, because the second is far more expensive to measure.
For early signal use hold rate — the share of viewers still watching at a given second — and the three-second or thumbstop rate, which is hold rate at the very start. Both are readable within hours because impressions are cheap. A hook that loses most of its audience in two seconds has failed, and no amount of extra budget rescues it.
For late signal use cost per result on your actual objective: purchase, install, subscription, qualified lead. That is the only number that decides whether an ad scales, and it needs the volume described above — which is why you screen on attention first. The trap in between is click-through rate: it arrives early and feels like a result, but a hook that generates curiosity clicks and no purchases looks excellent for three days and then quietly costs you money. Diagnostic, not verdict.
Kill on early signal once a variation has had its planned impressions and its hold rate sits clearly below the pack — fast, cheap, low regret. Kill on late signal only once it has spent enough to produce a real number and that number is materially worse than your target. Never kill something that has not spent enough to report; that is not a decision, it is a coin flip with extra steps. If the budget cannot support the round, run fewer variations — or get the cost per variation down, which is what how much UGC costs is really about.
Two different things get called scaling. The first is putting more budget behind the winning ad. The second is producing more variations of the winning angle. Most of the durable growth comes from the second.
More budget is the obvious move and it works until it doesn't. Raising spend on one ad set pushes it into progressively less responsive audiences, and large sudden increases can re-trigger the learning phase and cost you the stability you just bought. Raise gradually, watch cost per result rather than spend, and accept that every creative has a ceiling.
The more reliable move is to treat a winner as a discovered angle, not a finished asset. If a hook about a specific objection won, make four more videos that open on that objection with different creators and different proof. You are no longer guessing, so that round will beat your first blind one — and you have the next variations ready before the current ad tires. Reserve roughly a third of your creative budget for it. The winner is the brief for the next test, which is why ads for Meta get produced in batches around a proven angle.
Take your monthly creative testing spend, decide you are screening on hold rate first, work out what a few thousand impressions costs at your own CPM, and divide. For most small advertisers that lands between ten and twenty hooks in a first round — but that is the output of the sum, not the input. If your budget gives you six, run six properly rather than twenty badly.
Separate ad sets guarantee each variation a budget and a clean read, but they split your conversion volume and make the learning phase harder to exit. Several ads in one ad set let the algorithm allocate, which is efficient, but losing variations may barely spend and never get a fair read. At small budgets one ad set with a handful of ads is usually the more honest option, because it concentrates enough events in one place to be readable.
Enough for a round of hooks, plus roughly a third again for iterating on whatever wins. If your arithmetic says twelve variations a round and you run two rounds a month, that is twenty-four to thirty videos. At Yuugc's roughly $50 a video that is a defined line item; at real-creator rates of $150-$500 it is a different budget entirely, which is usually what forces the number down. See pricing for how volume affects the quote.
That is a normal first-round outcome and it is information. Check the offer, the landing page and the audience, in that order. If ten distinct hooks all failed to produce a viable cost per result, the problem is usually not the creative. Testing finds the best expression of an offer that works; it cannot rescue one that doesn't.
Tell us how many variations your round needs and we will quote it. Scripts written from your brief, one revision, full commercial rights, back in 48 hours.
Get a quote