Supplier sample testing protocol
How do you compare two product samples fairly on camera?
A fair on-camera comparison is designed before filming begins. Define one decision question, use the same measurable criteria and setup for both samples, alternate or randomize test order, record every planned run, and distinguish genuine filming errors from product failures. Preserve the raw footage and limit the conclusion to the units, conditions, and criteria actually tested.
Direct answer
What is the short answer?
Write the protocol before seeing either result. Label the products neutrally, match or document all relevant conditions, and decide in advance how many runs you will complete. If testing sequentially, alternate or randomize which product goes first. Record every run and permit retakes only for predefined invalidating events such as camera, measurement, or procedure failures—not because a product performed badly. Then separate recorded observations from your interpretation and publish a conclusion that names the samples, conditions, number of trials, and limitations. NIST defines repeatability in terms of successive measurements under the same measurement conditions, so changing the operator, environment, method, or equipment weakens a claim that the test was repeated consistently (NIST).
01
What should you decide before touching the samples?
Start with one narrow question that can produce a recorded answer, not a general contest over which supplier is “better.”
Useful questions include: “Which tested dispenser produces the more consistent measured dose?”, “Which tested unit heats 250 milliliters of water faster?”, or “Which tested fabric shows less visible damage after five standardized abrasion cycles?” A question such as “Which product has better quality?” is too broad to test fairly in one video.
Write down the primary outcome, measurement method and unit, secondary observations, stopping rule, product-failure rule, invalid-run rule, and number of planned runs. Do this before opening the camera app or conducting an informal trial that could influence the formal procedure. Experimental guidance recommends specifying objectives, endpoints, measurements, randomization, and analysis in advance rather than adapting them after results appear (NCBI Bookshelf).
For supplier selection, distinguish disqualifying criteria from preferences. A leak may be an automatic failure; a slightly louder motor may be a secondary observation. This prevents an attractive result on a minor feature from outweighing a failure that matters to customers.
- State the decision the comparison will inform.
- Choose one primary metric and its unit or scoring rubric.
- Set identical start and stop triggers for both samples.
- Define product failure and invalid-run conditions separately.
- Fix the number of runs before recording any results.
- Lock the protocol with a date or version number.
02
Which test conditions need to match?
Match every condition that could plausibly influence the primary result; if you cannot match it, measure or record it.
Confirm that both samples represent the variants you might actually sell: model, size, material, advertised specification, and included components. Give them neutral labels such as Sample A and Sample B so supplier reputation does not influence handling or commentary.
Use the same preconditioning, including charging, washing, resting, filling, cooling, drying, or acclimation. Measure consumables rather than estimating them visually. Keep the test surface, orientation, power source, measuring tool, operator, timer trigger, and reset procedure consistent.
Lock the camera position, focal length, focus, exposure, white balance, lighting placement, and lighting output. These controls do not improve the underlying product test, but they prevent visual treatment from making one sample appear brighter, larger, cleaner, or more detailed.
Record environmental variables that could matter, such as temperature, humidity, airflow, ambient noise, or time between runs. NIST describes operator and time-related conditions as possible nuisance factors: variables that are not the subject of the test but can affect its outcome (NIST Engineering Statistics Handbook). Say the samples used “the same documented setup” rather than claiming perfectly identical conditions unless you measured enough to support that wording.
03
Should the products be tested simultaneously or sequentially?
Test simultaneously when the products cannot influence each other; otherwise use sequential runs with controlled order.
A simultaneous test is useful when both products can start from equivalent conditions and operate independently. Keep a timer, scale, ruler, or measuring vessel visible where practical. Confirm that one product cannot heat, shade, vibrate, splash, obstruct, or otherwise affect the other.
Sequential testing is preferable when both products require the same outlet, instrument, fixture, operator, or camera position. It introduces order effects: the operator may improve with practice, equipment may warm up, batteries may decline, residue may remain, or the environment may drift.
Counter those effects by alternating order—A then B in round one, B then A in round two—or by using a coin flip or random-number generator to select the first product for each round. Record the order before starting and do not revise it after a poor result. Randomization helps prevent uncontrolled factors from consistently favoring one sample, while blocking is used to keep important controllable factors together (NIST Engineering Statistics Handbook).
- Choose simultaneous or sequential testing based on possible interference.
- For sequential tests, write the order before each round.
- Use the same reset process between runs.
- Allow the same cooling, charging, or drying interval.
- Log any order or reset deviation immediately.
04
How do you reduce operator and judging bias?
Use neutral sample labels, scripted actions, measurable criteria, and observable scoring anchors.
Script operator actions when handling can change the outcome: number of pumps, insertion angle, pressure, stirring pattern, assembly sequence, or cleaning method. If exact force or speed matters, use a suitable measuring tool or fixture rather than asking someone to reproduce it by feel.
For subjective criteria such as ease of use, finish, softness, or visible wear, define the scoring rubric beforehand. For example: 0 means the product cannot complete the task; 1 means completion with a major interruption; 2 means completion with a minor interruption; and 3 means completion without interruption.
When practical, ask a second person to score the footage without being told which supplier provided each sample. Do not substitute reactions such as “this obviously feels premium” for an observable result. Subjective observations can still inform supplier selection, but they should remain separate from the primary measured outcome.
05
What should be recorded for every run?
Create an uninterrupted audit trail connecting the protocol, sample, footage, measurement, and validity decision.
Begin the session with a slate showing the date, session ID, sample identifiers, comparison question, protocol version, planned rounds, order method, and key setup conditions. For each run, show or announce the run number, sample label, order position, start condition, result, and classification: valid result, invalid run, product failure, or disputed result.
Keep the camera running through setup and awkward moments where practical. If recording stops, start the next clip with a fresh slate. Preserve original files rather than keeping only the edited winners. The Office of Research Integrity notes that assessing image authenticity depends on original data and the context and effects of manipulation (ORI).
A simple run log should contain timestamps or filenames, measured result, scorer, deviations, retake decision, and reason. The log should reconcile with every planned run, including runs that never reach completion.
06
When is a retake fair, and when is it cherry-picking?
A retake is fair only when a predefined event invalidates the test or its recording; poor product performance is normally a result, not a reason to erase the run.
Valid retake reasons can include a stopped camera, incorrect sample label, failed shared fixture, malfunctioning timer or scale, outside interruption, or a material departure from the written procedure. Record the invalid run and its reason before beginning the replacement.
A jam, leak, breakage, difficult assembly, slow result, or unattractive outcome is not automatically a filming error. If the operator followed the protocol, preserve it as a product result. Do not quietly repeat one sample until it succeeds while showing the other sample’s first attempt.
If operator error is genuinely uncertain, mark the original run as disputed, explain why, and perform the additional run without deleting the first. In research-integrity terminology, omitting or changing results so that a record no longer accurately represents what occurred is falsification (ORI). Your supplier test is not necessarily formal research, but the same recordkeeping principle prevents cherry-picking.
07
How should observations and conclusions be separated?
Write what the camera or instrument recorded first; interpret the business significance in a separate statement.
An observation might read: “Sample A completed three valid runs in 42, 44, and 43 seconds; Sample B completed them in 51, 49, and 52 seconds.” The interpretation is: “Sample A was faster in these recorded trials under this setup.” Do not jump from one tested unit to “Supplier A always makes a faster product.”
Use qualified language for failures too. “The tested Sample B unit leaked three milliliters after the third cycle” is narrower and more defensible than “Supplier B’s products leak.” Repeatability under matched conditions does not establish that every unit, batch, operator, environment, or use case will produce the same result (NIST).
For a supplier decision, combine the recorded comparison with information the camera cannot establish: production tolerances, inspection terms, compliance documentation, packaging consistency, replacement policy, and batch-level quality checks. A clean sample win is evidence about the units tested, not proof of future manufacturing consistency.
08
How can you edit the comparison without making it misleading?
Edit for comprehension, but disclose transformations and apply equivalent treatment to both products.
Reasonable edits include trimming dead time, labeling runs, synchronizing angles, enlarging an instrument reading, and adding a clearly labeled replay or slow-motion segment. Mark time compression, omitted waiting periods, magnification, reconstruction, and replay on screen.
Do not give one product brighter lighting, tighter framing, warmer color, more setup assistance, favorable playback speed, or cleaner audio unless the difference is intrinsic to what was recorded and relevant to the test. Keep comparable moments visible for comparable durations.
If the footage becomes marketing, the demonstration communicates an advertising claim. The FTC states that a “right before your eyes” demonstration must truthfully depict what the product can do and discusses disclosure of material alterations (FTC). Review the entire edit—including captions, voiceover, thumbnails, and implied takeaways—rather than checking only the literal words of the conclusion.
09
What should the final evidence package contain?
Deliver the edited comparison with enough source material for another person to understand what was tested, how it was tested, and why the conclusion is limited.
Package the protocol version, supplier and sample identifiers, matched-conditions sheet, run order, complete log, raw footage, measurements, scoring sheets, deviations, retake explanations, edited export, and approved conclusion. Keep confidential supplier identifiers separate from public-facing neutral labels if necessary.
A useful conclusion template is: “We tested one unit from each supplier in [number] planned runs under the documented conditions. Sample A recorded [result], while Sample B recorded [result]. This supports choosing Sample A for [specific criterion], but it does not establish the performance of every unit or production batch.”
If published as advertising, support objective express and implied claims before publication. The FTC’s comparative-advertising policy supports truthful comparisons but emphasizes clarity and disclosure where needed (FTC). Qualifications should be prominent, understandable, and close to the relevant claim; FTC guidance explains that disclosures may need both visual and audio presentation when the claim is conveyed in both forms (FTC).
- Confirm the test question and rules were written in advance.
- Check that sellable variants and preparation methods matched.
- Reconcile every planned run with the footage and log.
- Verify that product failures were not classified as filming errors.
- Retain original clips and document every retake.
- Separate measurements, observations, and interpretations.
- Name the tested units, conditions, and limitations in the conclusion.
- Check that the edit does not imply a broader result than the evidence supports.
10
How can ATIYO support the comparison workflow?
ATIYO can preserve the protocol, brand context, assets, iterations, decisions, and reusable learnings associated with the comparison; it does not conduct or independently validate the test.
Create the evidence request before filming, attach the protocol and scoring criteria to the brief, and keep raw clips, logs, edited versions, deviations, and approved claim language together. This makes it easier to see why a take was retained, why a retake was permitted, and which conclusion the evidence supports.
The Brain and Static Studio share one user-provided connection through OpenRouter, fal.ai, or Kie.ai. Any generated analysis or creative treatment still needs human review against the source evidence. ATIYO does not connect to ad accounts, buy media, calculate ROAS, or know performance unless a user records it.
Media performance remains in the ad platform. ATIYO preserves creative context and learnings; it does not replace platform reporting or prove that one creative caused a performance outcome.
Frequently asked questions
Questions about this workflow
Is one uninterrupted video enough to prove the comparison was fair?
No. An uninterrupted recording can improve transparency, but it does not prove that preparation, order, consumables, operator actions, or selection of the session were fair. Pair the footage with a prewritten protocol, conditions sheet, complete run log, and retained originals.
What if one sample fails during setup?
Apply the predefined rule. If the operator materially departed from the procedure or shared equipment failed, classify the run as invalid and log the retake. If the procedure was followed and the sample jammed, leaked, broke, or could not be assembled as instructed, preserve it as a product failure or disputed result rather than silently replacing it.
How many runs should I film?
Choose the number before observing results and base it on the decision, product variability, cost, and risk. A few repeated demonstrations can reveal inconsistency in the tested units, but they do not establish batch-wide or supplier-wide performance.
Can I publish only the best-looking take?
You may use a representative take for a concise edit only if the complete run record supports the same conclusion and the selection does not hide contradictory outcomes. Preserve all runs, explain the selection method, and disclose relevant variability.
Should supplier names appear in the test?
Neutral A and B labels can reduce handling and judging bias. Keep a private key connecting those labels to supplier and sample identifiers so the evidence remains traceable. If publishing a named comparison, verify that the claims and qualifications accurately reflect the documented evidence.
Primary and official sources
Sources used in this guide
External product facts were checked against the organizations’ own documentation. Features can change; confirm current details before making a purchase or campaign decision.
- NIST TN 1297: Appendix D1. Terminology Definitions of repeatability, reproducibility, and measurement conditions.
- In Vivo Assay Guidelines — NCBI Bookshelf Guidance on advance objectives, endpoints, randomization, and blinding.
- NIST Engineering Statistics Handbook: Randomized Block Designs Principles for controlling nuisance factors and randomizing uncontrolled factors.
- ORI: Samples Context on original image data, authenticity, and manipulation.
- ORI: Definition of Research Misconduct Definition addressing omission or manipulation that misrepresents the record.
- FTC: Less than meets the eye? Guidance concerning product demonstrations and material alterations.
- FTC Statement of Policy Regarding Comparative Advertising Policy concerning truthful, clear comparative advertising.
- FTC Health Products Compliance Guidance Detailed guidance on substantiation and clear, conspicuous disclosures.
Move the plan out of scattered sheets
Run the roadmap, briefs, assets, and learnings in ATIYO.
ATIYO keeps the brand context and production decisions connected. It does not buy media, connect to ad accounts, or invent performance results.