ATIYO

Product testing workflow

How to Build a Library of Failed and Inconclusive Product Tests

A useful product-test library preserves every run—not only polished footage and favorable outcomes. It records the original question, planned method, actual conditions, result, anomalies, status, claim limits, and retest decision so creators can find contradictory evidence before writing the next brief.

By ATIYO editorial system Source and product-claim checks completed

Direct answer

How should you organize scattered product-test results?

Create one permanent record for every test run and classify it as passed, failed, invalid, inconclusive, or stopped. Keep the planned method separate from what actually happened, preserve negative and uncertain outcomes beside successful ones, and link every conclusion to its footage and raw observations. Most importantly, assess execution validity separately from product performance: a valid test can reveal a product failure, while a broken test supports no product conclusion. Finish each record with what may be said, what cannot be said, and whether the test should be repeated, corrected, redesigned, or retired.

01

What should each product-test status mean?

Statuses should describe both whether the method worked and what the product did. Treat the following as internal workflow definitions, not universal scientific or regulatory classifications.

Passed means the planned method was completed under the required conditions, the predefined acceptance rule was met, and no recorded anomaly materially undermined the result. A pass may contribute to a body of support, but it does not automatically establish a broad advertising claim.

Failed means the method was executed validly, but the product did not meet the predefined acceptance rule. This is product-negative evidence and should remain visible in searches, briefs, and reviews.

Invalid means a problem with the method, sample, equipment, setup, execution, or recording prevents the run from answering its question. It supports no conclusion about product performance, although it can provide a valuable process lesson.

Inconclusive means the observations are usable but insufficient to decide whether the acceptance rule was met. Results may conflict, the sample may be inadequate, or uncertainty may cross the decision boundary. Record exactly what evidence is missing.

Stopped means the run ended early because of a safety, equipment, sample, time, or protocol issue. Then classify the resulting evidence as invalid or inconclusive where possible, while retaining the stop time and reason.

02

What information belongs in every test record?

Use one record per test question and execution. A spreadsheet can work initially, provided that footage and documents have stable links and nobody replaces earlier outcomes.

Start with identity fields: test ID, product, model, variant, size, batch or lot, brand, project, creator, operator, date, protocol version, related test IDs, status, reason code, and the claim or question examined.

Create a locked planned section containing the exact question, acceptance rule, sample size and selection method, comparison or control, equipment and settings, procedure, run order, duration, environmental conditions, expected output, stop rules, and known limitations. NIST describes experimental design as advance planning intended to support valid, objective conclusions, beginning with a clear objective and the factors being studied [1].

Keep a separate observed section for the actual sample and conditions, deviations, measurements, raw observations, anomalies, equipment or operator problems, stop reason, result against the original rule, final status, and decision owner. Never quietly rewrite the threshold, endpoint, or conditions after watching the footage.

  1. Assign a unique test ID before the run.
  2. Write and lock the planned method.
  3. Record actual execution without correcting the history.
  4. Apply a status against the original acceptance rule.
  5. Add claim boundaries and a retest decision.
  6. Link the record to related protocols, runs, footage, and learning cards.

03

How do you separate product failure from test failure?

Make two decisions in order: first determine whether the execution could answer the question; only then determine whether the product met the rule.

A valid failed test says something unfavorable about the product under the recorded conditions. An invalid test says the execution cannot answer the question. An inconclusive test says the evidence does not justify success or failure. These distinctions prevent a team from dismissing every unfavorable result as a production problem or treating broken setups as proof that a product failed.

Require a specific invalidation reason instead of allowing a vague “bad test” label. Useful reason codes include wrong product, pre-damaged sample, unsuitable or uncalibrated equipment, missing control, conditions outside the protocol, procedure deviation, measurement error, missing critical footage, safety stop, or an acceptance rule created after the run.

Record potentially influential conditions such as prior product use, charge level, warm-up time, water temperature, surface, ambient conditions, operator, test order, camera interruption, and comparison-product condition. NIST identifies operator, time, temperature, and similar nuisance factors as possible sources of variation and discusses blocking or randomization as ways to manage their effects [2].

04

How should footage and anomalies be preserved?

Keep source media attached to the complete test record, with enough indexing for another creator to retrieve setup details, measurements, anomalies, and outcomes without relying on the edited clip.

Store original footage, stills, proxy files, transcripts or searchable descriptions, and edited exports separately. Add timecodes for the setup, sample identification, measurement, deviation, anomaly, stop point, and result. If a critical moment was not captured, record that as missing evidence rather than assuming the event occurred as remembered.

Do not let a polished clip become the sole record. Editing can remove pauses, failed attempts, changed equipment, or environmental context that affects interpretation. Preserve a stable asset ID or checksum where your storage system supports one, and never delete the failed run merely because it cannot be published.

Treat anomalies as searchable fields. “Battery depleted during run,” “comparison sample previously used,” and “camera stopped before endpoint” are more useful than “went wrong.” Specific descriptions allow future briefs to exclude unsuitable methods or plan controlled corrections.

05

How do you stop favorable footage from becoming a misleading claim?

Add a claim-boundary note to every run and review related evidence together. One favorable clip should not be detached from failed, contradictory, invalid, or inconclusive runs.

For each test, write four short entries: observed, supported, not supported, and conflicting evidence. “Observed” states what visibly or measurably happened. “Supported” gives the narrowest statement consistent with the method. “Not supported” identifies broader durability, safety, typicality, or comparative conclusions the test cannot establish. “Conflicting evidence” links every relevant run that weakens or qualifies the result.

FTC guidance says advertisers need a reasonable basis for objective express and implied claims before dissemination, and that the required support depends on the claim [3]. Its health-products guidance also warns against selecting only favorable research while discounting results that do not support the claimed effect [4]. Although your product and claim may require different evidence, the practical library rule is straightforward: reviewers should see the complete relevant record, not only the most convenient clip.

Visual presentation matters alongside literal wording because advertising can communicate implied claims through its overall message [3]. Route technical, legal, health, or safety claims to an appropriately qualified reviewer; a creator library organizes evidence but does not verify compliance.

06

What retest rules prevent the same invalid test from recurring?

Every invalid or inconclusive record should end with a named retest decision, an owner, and the exact change required before another attempt.

Use four decisions. Do not repeat when the question is irrelevant, unsafe, or cannot be answered credibly with available resources. Repeat unchanged when a random interruption occurred but the method remains suitable. Repeat with controlled correction when a specific execution problem can be fixed without changing the question or acceptance rule. Redesign when the method itself cannot answer the question.

A retest receives a new ID. Link it to the prior run, identify the corrected variable, and preserve the old record. If the question, threshold, sample, or procedure changes materially, create a new protocol version rather than presenting the new execution as a simple rerun. Replication and randomization can help separate random error from method or model shortcomings [5].

Be especially careful with extreme-condition demonstrations. NIST notes that excessive stress in accelerated life testing can produce failure mechanisms that would not occur under actual-use conditions [6]. Record why the stress represents normal use—or state plainly that it does not. For a more detailed invalidation workflow, use the Product Test Failure Troubleshooting Protocol linked below.

07

How should the library be structured for retrieval?

Build three linked layers: reusable protocols, individual test runs, and concise learning cards. A learning card summarizes the lesson but must always link back to the complete and potentially contradictory evidence.

The protocol library stores the method, acceptance rule, equipment, required conditions, and version history. The test-run library contains every execution across all statuses. Learning cards turn those records into brief-ready instructions such as “Do not repeat this setup,” “Control starting charge,” or “Current footage cannot support a durability claim.”

Useful filters include product, category, claim type, protocol version, status, invalidation reason, batch, environment, failure mode, anomaly, footage availability, retest eligibility, review state, and “do not repeat.” Search results should expose conflicting tests rather than ranking passed outcomes first.

ATIYO is not a laboratory, testing provider, or claims-verification service. It can preserve the roadmap, brief, brand context, assets, iterations, decisions, and reusable learnings surrounding product-test creative. Media performance remains in the ad platform; ATIYO preserves creative context and learnings. Campaign experiments are a separate activity from product-evidence records, even when both inform future creative.

Frequently asked questions

Questions about this workflow

Should failed and invalid tests be stored together?

They can share one test-run library, but they need separate statuses and reason codes. A failed valid test is evidence about the product under recorded conditions; an invalid test is evidence that the execution could not answer the question.

Can an inconclusive result be used in a product claim?

Do not present an inconclusive run as proof of success or failure. Record its observations, explain what remains unknown, connect conflicting evidence, and obtain appropriate review before using any related claim.

Should edited footage be deleted when a test fails?

No. Preserve the original source and identify any edited version as an output. Failed footage may document a product outcome, while invalid footage can reveal setup problems that future creators should avoid.

When is a retest a new test rather than a repeat?

Treat it as a new protocol version when the question, acceptance rule, sample design, conditions, or procedure changes materially. It should still link to the earlier run so the history remains visible.

Can ATIYO determine whether a product claim is substantiated?

No. ATIYO is not a testing or claims-verification service. It organizes creative context, footage, decisions, and learnings. Qualified legal, technical, or subject-matter reviewers must assess whether evidence supports a particular claim.

Primary and official sources

Sources used in this guide

External product facts were checked against the organizations’ own documentation. Features can change; confirm current details before making a purchase or campaign decision.

  1. NIST: What is experimental design? Supports defining the objective, factors, and experimental plan before interpreting outcomes.
  2. NIST: Randomized block designs Discusses nuisance factors, blocking, and randomization.
  3. FTC: Advertising FAQ's—A Guide for Small Business Explains reasonable-basis expectations and consideration of express and implied advertising claims.
  4. FTC: Health Products Compliance Guidance Used for its discussion of evaluating the totality of relevant evidence rather than selecting favorable findings.
  5. NIST: Glossary of DOE Terminology Supports the discussion of replication and randomization.
  6. NIST: Accelerated life tests Warns that excessive stress can create failure mechanisms not present under normal-use conditions.

Move the plan out of scattered sheets

Run the roadmap, briefs, assets, and learnings in ATIYO.

ATIYO keeps the brand context and production decisions connected. It does not buy media, connect to ad accounts, or invent performance results.