Creative-test retrospective template
How to Run a 30-Minute Creative-Test Postmortem Without Inventing Learnings
A useful postmortem is a decision meeting, not a general creative critique. The team reconstructs the test, checks whether the result is interpretable, separates evidence from speculation, and assigns one next action.
Direct answer
What should your team review after an ad creative test ends?
Review the original hypothesis and isolated variable, the success metric chosen before launch, delivery comparability, campaign or site changes, overlapping tests, the observed result and its uncertainty, and what the evidence does not establish. Then classify the test as valid, valid with caveats, invalid, or undocumented; record observations, supported conclusions, hypotheses, and unknowns separately; and route the result to scale, repeat, iterate, retire, or investigate. “Invalid test” and “no clear winner” are acceptable outcomes.
01
What is the 30-minute postmortem agenda?
Use a fixed agenda: reconstruct the test for five minutes, validate execution for five, read the result for five, classify the evidence for seven, choose a route for five, and assign ownership for three.
Keep the meeting limited to people needed to interpret the setup and make the next decision. Complete the basic test record before the meeting where possible; otherwise, the discussion can be consumed by finding dates, assets, and campaign settings.
Google Ads recommends defining a clear hypothesis, testing one variable at a time, and selecting one or two success metrics before an experiment begins. TikTok’s split-testing documentation likewise describes selecting one variable while holding other variables constant. Recovering those original decisions is therefore the first task—not debating whether people personally liked the ad.
- 0–5 minutes — Reconstruct the hypothesis, control, treatment, intended variable, audience, placements, primary metric, threshold, and run dates.
- 5–10 minutes — Check delivery, tracking, concurrent activity, campaign changes, and commercial or site changes.
- 10–15 minutes — Read control and treatment performance against the predetermined primary metric.
- 15–22 minutes — Separate observations, supported conclusions, hypotheses, and unknowns.
- 22–27 minutes — Choose scale, repeat, iterate, retire, or investigate.
- 27–30 minutes — Assign the action, owner, due date, next variable, and required creative work.
02
How do you reconstruct what the test was meant to prove?
Write the original test as: “Changing X will improve Y because Z.” Identify X as the intended variable and Y as the predetermined primary metric.
Record the test ID, platform, campaign, dates, audience, placements, control, treatment, variable, hypothesis, primary metric, and decision threshold. Examples of variables include the opening hook, offer, spokesperson, proof device, format, or call to action.
If the ads differed in several ways—such as speaker, script, setting, duration, and edit—label the exercise a concept comparison. You may decide which complete concept performed better in that setting, but you cannot attribute the difference to one component. For future tests, the Ecommerce Creative Testing System explains how to connect test design with creative planning.
03
How do you decide whether the test is valid?
Classify execution as valid, valid with caveats, invalid, or unknown because documentation is incomplete. Do this before writing a creative learning.
Ask whether every variation launched on time and received enough comparable delivery. Check for materially imbalanced spend or traffic, rejection or downtime, tracking failures, and changes to targeting, placements, bid strategy, optimization event, attribution settings, offer, or landing page.
Also check whether another experiment competed for the same traffic and whether promotions, seasonality, pricing, inventory, or site performance changed. Google warns that changing a base campaign during an experiment makes attribution harder and that concurrent experiments can interfere with one another. TikTok explains that its native split tests use mutually exclusive audience groups; a manually assembled comparison may not have that protection.
Do not rescue a contaminated test because one result looks promising. Mark it invalid or valid with caveats, explain why, and route it to investigation or repetition. An honest invalid result prevents weak evidence from becoming a rule in future briefs.
- Did all variants launch and run as planned?
- Was delivery sufficient and reasonably comparable?
- Did targeting, bidding, optimization, attribution, offer, price, page, or inventory change?
- Was there rejection, downtime, or a tracking problem?
- Did another test or campaign overlap?
- Was the test stopped according to the planned rule rather than an appealing early result?
04
How should the team read the result?
Read the predetermined primary metric first, followed by the absolute difference, relative difference, platform confidence or significance status, and delivery context. Treat secondary metrics as diagnostic unless they were part of the decision rule.
Use a neutral result statement: “Treatment B produced [result] versus Control A on the predetermined [metric] from [dates]. The platform reported [confidence or significance status]. Delivery was [comparable or not comparable].” Also record spend, impressions, reach, conversions, and runtime when relevant to interpretation.
Google’s experiment reporting distinguishes estimated differences, confidence intervals, statistically significant outcomes, and cases in which there is not enough information to choose a winner. “No clear winner” is therefore a usable result, not a reason to switch to a secondary metric after seeing the data.
A treatment that loses on the predetermined conversion metric should not be declared the winner solely because it had the best attention metric. The secondary result may generate a new hypothesis—for example, that the opening earns attention but the body or offer fails to convert.
05
How do you separate evidence from interpretation?
Require every statement to be recorded as one of four knowledge types: observation, supported conclusion, hypothesis, or unknown.
An observation directly describes the result: “The testimonial-hook treatment had a lower CPA than the product-demo control during the test.” A supported conclusion stays within the tested audience, offer, placements, period, variable, and metric: “For this audience and offer, the testimonial-hook treatment outperformed the control on CPA during the test window.”
A hypothesis proposes an explanation that needs another test: “Showing social proof earlier may have reduced purchase anxiety.” An unknown names what the test could not answer: “We do not know whether the effect came from the speaker, testimonial format, claim, or their combination.” Do not merge these categories into one vague “learning.”
06
Which conclusions overreach the evidence?
Broad claims about customers, formats, or causation usually overreach when the test covered one comparison in one operating context.
Overreach: “UGC beats studio creative.” Safer: “The UGC-style treatment beat the studio control on CPA in this test. Because speaker, script, setting, and edit also changed, production style was not isolated.”
Overreach: “Customers want social proof.” Safer: “The testimonial-opening version outperformed the product-demo control. Earlier social proof is one possible explanation to test next.”
Overreach: “Short videos convert better.” Safer: “The 15-second treatment outperformed the 30-second control, but duration was not the only difference. A duration-only test is required.”
Overreach: “The hook caused a 25% CPA improvement.” Safer: “The hook treatment’s observed CPA was 25% lower during the test. A causal claim requires the hook to be isolated and the experimental evidence to support that interpretation.”
Overreach: “This creative is a winner.” Safer: “This creative won this comparison on the predetermined metric for this audience, offer, placements, and test window.”
07
How should the team choose the next route?
Choose exactly one primary route: scale, repeat, iterate, retire, or investigate. The route should follow from validity and evidence strength, not enthusiasm for the asset.
Scale when a valid result supports broader use; specify the audience, rollout scope, and monitoring guardrails. Repeat when the comparison is promising but uncertain. Iterate when the result supports a direction but reveals a specific variable to isolate next. Retire when the treatment underperformed clearly and offers no strategic reason for another test. Investigate when delivery, overlap, tracking, or another confound prevents interpretation.
Avoid “make more ads like the winner.” Name what should be preserved and what should change. For example: “Preserve the testimonial opening; test customer versus expert spokesperson while holding the claim, script structure, duration, and offer constant.”
08
What should the completed postmortem record contain?
The record should preserve the setup, validity check, results, evidence classification, decision, ownership, and unresolved questions in one place.
Copy this structure into your operating document or use the Resources/Ad Testing Tracker Template. Keeping standardized records helps future teams understand why a decision was made instead of reverse-engineering it from asset names or campaign screenshots.
TEST ID / REVIEW DATE / PARTICIPANTS 1. ORIGINAL TEST Hypothesis: Control: Treatment: Intended variable: Audience and placements: Primary success metric: Decision threshold: Run dates: 2. EXECUTION CHECK Comparable delivery: Yes / No / Unknown Tracking intact: Yes / No / Unknown Unplanned campaign changes: Overlapping activity: Offer, site, price, or inventory changes: Other contamination: Validity: Valid / Valid with caveats / Invalid / Unknown 3. RESULTS Control result: Treatment result: Absolute difference: Relative difference: Platform confidence or significance status: Secondary diagnostic metrics: Delivery limitations: 4. INTERPRETATION Observations: Supported conclusion: Hypotheses: Unknowns: What this test does not establish: 5. DECISION Route: Scale / Repeat / Iterate / Retire / Investigate Action: Owner: Due date: Next variable: Assets or brief to update:
09
Where does ATIYO fit into the workflow?
Media performance remains in the ad platform. ATIYO preserves the creative context and learnings around the test: what changed, why it was tested, what the evidence supports, what remains unknown, and what the team should do next.
ATIYO can organize the roadmap, original brief, brand context, assets, iterations, decision, and reusable learning so that the result can inform later work. It does not connect to ad accounts, buy media, calculate ROAS, or know performance unless a user records it.
After the postmortem, connect the supported conclusion and unknowns to the next roadmap item rather than leaving them in meeting notes. The Help/Roadmaps/Tracking Winners And Capturing Learnings guide covers that handoff.
Frequently asked questions
Questions about this workflow
What if the test has no clear winner?
Record “no clear winner,” preserve the delivery and uncertainty details, and choose repeat, iterate, or investigate. Do not select a winner using an unplanned secondary metric.
Can we call a multi-variable comparison a creative test?
Yes, but describe it as a concept comparison. It can identify which complete execution performed better in that context, not which individual component caused the difference.
Should TikTok Top Ads examples count as evidence?
No. TikTok’s Top Ads Dashboard can support inspiration and follow-up hypotheses, but another advertiser’s ad does not establish what will work for your audience, offer, or campaign.
Who should own the postmortem?
Assign one facilitator who can recover the test setup, enforce the evidence categories, and record the decision. The action itself should have a named owner and due date.
Primary and official sources
Sources used in this guide
External product facts were checked against the organizations’ own documentation. Features can change; confirm current details before making a purchase or campaign decision.
- Google Ads: Test with confidence with the Experiments page Guidance on hypotheses, isolated variables, predetermined success metrics, campaign changes, and experiment records.
- Google Ads: Monitor your experiments Explains experiment differences, confidence reporting, statistically significant results, and outcomes without enough information to identify a winner.
- Google Ads: Experiments FAQs Guidance concerning interference between concurrent experiments.
- TikTok Ads Manager: About Split Testing Describes controlled variables, mutually exclusive audience groups, and winner determination.
- TikTok Ads Manager: Split Testing Variables Documents the selection of one test variable, including creative variables.
- TikTok: How to use the Top Ads Dashboard Describes tools for finding and examining high-performing ads; useful for inspiration rather than proof for a brand’s own audience.
Move the plan out of scattered sheets
Run the roadmap, briefs, assets, and learnings in ATIYO.
ATIYO keeps the brand context and production decisions connected. It does not buy media, connect to ad accounts, or invent performance results.