Creative test integrity guide
Can you trust a creative test after the offer or landing page changes?
Sometimes—but do not accept the full-period result until you check how the change affected each test arm. An interrupted test can be retained, qualified, invalidated, or rerun depending on timing, exposure, measurement integrity, and whether the creative result stayed consistent before and after the change.
Direct answer
What is the direct answer?
Treat the new price, bundle, discount, or landing page as a second experimental factor. First determine whether both creative arms received the change at effectively the same time. Then separate the run into pre-change and post-change windows, check allocation and tracking, and compare the creative difference within each window. You can usually retain the learning when both arms changed together, the intended creative variable remained untouched, measurement stayed intact, and the same creative led by a reasonably consistent margin in both windows. Qualify the result when only one clean window is informative. Invalidate contaminated intervals when exposure, routing, or measurement differed by arm. Rerun when the winner reverses, no adequate clean window remains, or you need an answer for the current offer. Google recommends isolating variables because simultaneous changes make causation difficult to identify. It also advises avoiding changes to base campaigns during experiments because they complicate interpretation (Google Ads).
01
Why can a simultaneous change still contaminate the conclusion?
Giving both arms the same new offer preserves symmetry, but it does not prove that the offer had the same effect on both creatives.
A test arm is one version being compared, such as Creative A versus Creative B. Contamination occurs when another change makes it unclear whether the observed difference came from the intended creative variable or from something else.
If both arms move from full price to a 20% discount at the same time, the offer can be treated as a time block: compare A with B under full price, then compare A with B under the discount. Blocking is an experimental-design method for comparing treatments while holding a known nuisance condition constant within each block (NIST).
The remaining risk is interaction, meaning the effect of one factor depends on the setting of another. A value-led creative could become more persuasive under a discount while a quality-led creative loses its advantage. Experimental design distinguishes these interaction effects from the independent effects of each factor (NIST). A pooled full-run result can conceal that reversal.
02
Which four questions determine whether the test is salvageable?
Follow the branches in order. Do not skip directly to whether the headline winner stayed ahead.
Use the actual deployment boundary, including any known delay before the new condition reached shoppers. Document rather than conceal an uncertain transition interval.
A landing-page change is substantive because the page is the destination after the click. Google identifies usefulness, relevance, navigation, and alignment with expectations as contributors to landing-page experience (Google Ads). A revised product page can therefore change downstream behavior even when the ads remain identical.
- 1. Did the change reach both arms at effectively the same time? If no, invalidate the interval with unequal exposure. Analyze a remaining clean interval only if it is adequate; otherwise rerun.
- 2. Did the change alter the tested creative variable, destination assignment, audience eligibility, tracking, or traffic allocation? If yes, the creative-only effect is no longer identifiable for that interval. Invalidate it or rerun.
- 3. Can you separate pre-change and post-change data using a defensible timestamp? If no, treat the result as directional at most. Rerun before recording a durable learning.
- 4. Does the creative comparison point in the same direction in both windows? If yes with adequate evidence, retain it. If yes but noisy, qualify it. If the winner reverses, record an offer-dependent observation and rerun under the commercial condition you intend to keep.
03
When should you retain, qualify, invalidate, or rerun?
Choose the narrowest classification supported by the evidence, not the label that preserves the most optimistic result.
Retain: Both arms received the change together; creative assets, assignment, and tracking stayed intact; and the direction and plausible magnitude of the creative difference were consistent in both windows. Record the conclusion as robust across the two observed conditions—not across every possible offer.
Qualify: Both arms changed together, but only one window has enough evidence, or the estimated gap varies materially without a clear reversal. Use contextual language such as “Variant B led under the revised bundle,” rather than “Variant B is the winner.”
Invalidate: One arm changed first, destinations or measurement differed by arm, the intended creative variable was edited, or the ad-to-page combination changed unevenly. Preserve the incident as an audit record, but do not record a creative winner from the contaminated interval.
Rerun: The winner flips, neither clean window is adequate, assignment integrity is uncertain, or the business needs a conclusion for the current offer rather than the retired one. Invalidation describes evidence you should not use; rerunning describes the next action. A test may require both.
04
How should you reanalyze the interrupted run?
Compare the creative effect within each commercial condition rather than comparing each arm’s raw performance before and after the change.
Keep the original primary metric and population definition. Google recommends selecting success metrics before beginning an experiment (Google Ads); switching metrics after seeing disruption can turn analysis into result-shopping.
Create old-condition and new-condition windows. Within each window, calculate the same A-versus-B comparison. Check delivery, exclusions, destination URLs, tracking events, and traffic split separately for each period. An unexplained sample-ratio mismatch—an observed allocation that differs unexpectedly from the planned allocation—can indicate assignment, execution, logging, or analysis problems. Microsoft’s experimentation research warns against ignoring such mismatches (Microsoft Research).
If you have appropriate analytical support, estimate creative, offer period, and creative-by-offer interaction effects. Otherwise, a disciplined side-by-side comparison is still more honest than pooling incompatible periods. Do not keep slicing dates until a preferred winner appears.
- Record the deployment timestamp and any excluded transition period.
- Freeze the original metric, audience, attribution rule, and exclusions.
- Compare A versus B within the old condition.
- Compare A versus B within the new condition.
- Check allocation, delivery, URLs, and measurement in both windows.
- Assign one integrity label and state exactly which condition and metric it covers.
05
What do common D2C interruption scenarios mean?
The correct decision depends on where the change entered the customer journey and whether it affected the arms equally.
Both creatives improve after a discount, and their relative gap stays similar: Retain the creative learning across the full-price and discounted conditions if allocation, tracking, and uncertainty checks are satisfactory.
The winner flips after a bundle change: Do not average the windows into one winner. The evidence suggests that the creative response depends on the bundle. Qualify each observation by condition and rerun under the bundle you expect to keep.
Only Variant B points to the revised product page: Invalidate that interval as a creative-only test. Creative and destination effects are inseparable, consistent with Google’s warning that changing multiple variables obscures what caused the result (Google Ads).
Click-through rate stays stable but purchase conversion changes after a page revision: The click-level result may remain usable if the creative delivery and click measurement were unaffected. The purchase conclusion is contaminated by the downstream change. Limit the learning to the metric and funnel stage actually supported.
The change occurs near the end: A short contaminated tail does not necessarily erase a substantial clean pre-change window. Analyze the clean window separately, but do not claim that its conclusion generalizes to the new offer.
06
What should you document before closing the test?
Save enough context for another person to reconstruct which assets, offer, destination, and measurement rules produced the result.
Record the exact change, timestamp, affected arms, asset versions, destination URLs, old and new commercial conditions, transition exclusion, original metric, traffic allocation, window-level results, and final classification. Link asset and destination histories through an explicit version-control process, then add the contamination event to your testing tracker.
ATIYO can organize the roadmap, brief, brand context, asset versions, iterations, and the qualified learning when your team records that information. It does not connect to ad accounts, calculate ROAS, or independently know performance.
Media performance remains in the ad platform. ATIYO preserves creative context and learnings, including which offer, landing page, and asset versions were live when a result was produced.
Bring qualified, invalidated, and rerun decisions into the weekly creative decision meeting. The goal is to prevent a conditional observation from being reused later as a universal rule.
Frequently asked questions
Questions about this workflow
Can I trust the result if both test arms changed at the same time?
Possibly. Equal timing removes one major source of unequal exposure, but the offer may interact with the creative. Compare the A-versus-B result separately before and after the change.
Should I restart immediately after any price change?
Not automatically. Preserve the current records first. If a clean pre-change or post-change window remains, it may support a qualified conclusion. Rerun when no adequate clean interval exists or when you need a decision for the current price.
Can I keep the CTR learning but discard the conversion learning?
Yes, when the interruption affected the downstream offer or page while creative delivery and click measurement remained intact. State explicitly that the conclusion applies to click-through behavior, not purchases.
What if I cannot identify the exact change timestamp?
Do not invent a boundary. Exclude a defensible transition interval if possible; otherwise treat the result as directional and rerun before saving a durable learning.
Is a qualified result the same as a failed test?
No. A qualified result can be useful when its condition is stated precisely. It becomes misleading only when an offer-specific or window-specific observation is promoted into a general creative rule.
Primary and official sources
Sources used in this guide
External product facts were checked against the organizations’ own documentation. Features can change; confirm current details before making a purchase or campaign decision.
- Google Ads Help — Test with confidence with the Experiments page Guidance on isolating variables, choosing metrics before testing, avoiding base-campaign changes, and retaining experiment records.
- Google Ads Help — Landing page Definition of a landing page and factors associated with landing-page experience.
- NIST — Randomized block designs Methodological basis for comparing treatments within blocks that hold a nuisance condition constant.
- NIST — What is experimental design? Background on experimental factors, main effects, and interactions.
- Microsoft Research — Diagnosing Sample Ratio Mismatch Research on allocation mismatches as indicators of experiment-integrity problems.
Move the plan out of scattered sheets
Run the roadmap, briefs, assets, and learnings in ATIYO.
ATIYO keeps the brand context and production decisions connected. It does not buy media, connect to ad accounts, or invent performance results.