Two different kinds of evidence
Ad platforms run experiments on real audiences. Meta's A/B testing, for example, shows each version to a separate segment so nobody sees both, then compares performance on a cost-per-result basis. Google Ads custom experiments split a campaign's traffic between the original and a variant.
A predicted attention map uses no audience. It describes which parts of the creative are likely to stand out. It is useful before an experiment, to decide which variants deserve a slot, and never as a replacement for one.
References: Meta: About A/B testing · Google Ads: Set up a custom experiment
Start with a hypothesis, not a pile of creatives
Meta recommends choosing a hypothesis before selecting a variable to test. The same discipline helps a creative review. Decide what the ad must communicate first, usually the product, the offer or the brand, and which change you believe will help.
One variable per comparison keeps the result readable. If the crop, headline and colour all change at once, neither the heatmap nor the live test can tell you which change mattered.
References: Meta: About A/B testing
Screen the variants
Upload each creative at its delivery size and composition. Review the original first, then the prediction. Name the elements under the strongest regions and compare them with your intended focus.
- Check that the product or offer sits in or near a strong region.
- Look for background detail, faces or props that compete with the message.
- Make sure required text stays legible at the placement's real size.
- Compare variants in one run so the same model and settings apply.
- Keep the variants whose maps support their hypothesis and rework the rest.
References: Ad creative review · Compare variants
Read scores with care
A higher attention score is a summary of the prediction, not an expected click-through rate or return on ad spend. A cluttered creative can score well because many regions are salient. Inspect where the emphasis sits before trusting a ranking.
Close scores usually mean the model sees little difference. That is useful information: the live test may also struggle to separate them, so consider a bolder variant.
References: Metric definitions
Hand over to the live test
Run the surviving variants in your ad platform's experiment tool with comparable budgets and audiences. Meta advises against informal testing by switching ad sets on and off, because it can produce overlapping audiences and unreliable results.
Record what each variant emphasised in the attention review next to its live result. Over time, that shows whether the review is filtering out weak options for your account, which no general claim can establish for you.
References: Meta: About A/B testing