Learn how to design a creative testing survey, from exposure order and scale anchors to cell sizing and fielding QA that keeps results readable.

White outline of Goldie, the SurveyMonkey mascot

A creative testing survey measures how a target audience responds to specific ad concepts before those concepts go live. The survey is the instrument: the screener, the exposure order, the scales, and the open ends that turn a set of headlines, images, or video cuts into comparable numbers.

Creative testing is the wider practice around that instrument. It covers what you measure, how much budget to set aside, how long a round takes, and how results feed the media plan. If you need that broader picture first, start with the guide to creative testing, then come back here to build the questionnaire. This page is about the survey itself: what you ask, in what order, with which answer options, and how you check the data before you trust it.

That distinction matters because most creative tests fail at the questionnaire, not the strategy. Teams pick the right concepts, field to the right audience, and then compare two scores that were never comparable in the first place.

A creative test produces one number per concept per metric. Everything that number means comes from how the question was built. Small design mistakes do not add noise evenly, they bend results in one direction, which is worse: you get a confident answer that points the wrong way.

Here are the flaws that show up most often in review, what each one corrupts, and how to fix it.

Design flawWhat it corruptsThe fix
Unlabeled slider or 0 to 100 scaleCross-concept comparability. Respondents anchor differently, so a 62 in one cell is not a 62 in another.Use a fully labeled five-point or seven-point scale with the same anchors in every cell.
Showing every concept to every respondentIndependent scores. By concept three, ratings reflect comparison and fatigue, not response to the ad.Use a monadic design so each respondent sees one concept, or cap sequential monadic at three concepts.
Purchase intent asked only after exposureAttribution. You cannot tell whether the concept moved intent or the audience already had it.Ask a baseline intent item before exposure, then re-ask the same item after.
Diagnostic battery before recallUnaided recall. Once you name the brand or the message in a question, recall is contaminated.Always ask unaided recall first, then aided takeaway, then diagnostics.
Too many concepts in one surveyData quality across the whole survey. Length drives dropoff and straightlining.Keep the concept count low and the question count per concept lower.
Uneven cell sizesStatistical readability. A cell with half the sample has roughly a wider margin of error.Set a fixed target per cell and quota to it.
Leading question wordingEvery score downstream. "How much did you love this ad?" is not a measurement.Use neutral stems and a balanced scale with an equal number of positive and negative points.
No attention or speeder checksEffect size. Inattentive completes flatten differences between concepts.Add a trap item and a minimum time threshold before the exposure block.

Every creative testing survey is assembled from four decisions. Get these right and the rest of the questionnaire mostly writes itself.

Exposure design decides who sees what. Monadic gives each respondent a single concept, which keeps every score independent and makes it the default for creative testing across the industry, as YouGov notes in its guidance on the method. Sequential monadic shows each respondent several concepts in randomized order, which costs less sample but introduces order effects.

The practical limits are well documented. Drive Research recommends testing no more than three concepts in a sequential monadic design, and points out that pushing 20 concepts into one survey produces roughly 100 questions and about a 60-minute interview. Attest advises no more than five or six ideas per survey, ideally two or three, and no more than five questions per asset.

Comparative cells, where you deliberately place two concepts side by side and ask the respondent to choose, answer a different question. They tell you which concept wins a head-to-head, not how either performs on its own. Run them as a final tiebreaker block after the monadic scores, never as a substitute.

Order is not cosmetic. Each block in a creative testing survey has to come before the block it would otherwise contaminate. Here is a question sequence you can copy directly into your survey and adapt:

  1. Screener. Category usage, purchase recency, and demographics needed for quotas. Terminate out of scope respondents here so you do not pay for them.
  2. Attention and speeder check. One trap item placed before exposure, plus a minimum time threshold on the screener page.
  3. Pre-exposure baseline. "How likely are you to consider [category or brand] the next time you buy [product]?" on your standard five-point or seven-point scale. Ask about brand familiarity here too if you plan to read results by prior awareness.
  4. Exposure. One concept, shown full screen, with a forced timer so respondents cannot click past it.
  5. Unaided recall. "Thinking about the ad you just saw, what do you remember about it?" as an open text field. Nothing before this question may name the brand or the message.
  6. Aided message takeaway. "Which of the following was the main message of the ad you just saw?" with a randomized list of candidate messages plus "none of these."
  7. Diagnostic battery. Four to six labeled scale items on the metrics you care about, presented in a randomized grid.
  8. Post-exposure purchase intent. The same item you asked in step three, word for word, on the same scale.
  9. Open end. "What, if anything, would you change about this ad?" This is where the reason behind a low score usually appears.
  10. Remaining demographics. Anything not needed for quotas goes at the end.

The pre-post pairing in steps three and eight is the piece most teams skip.

Asking the same intent item before and after exposure gives you a within-respondent shift rather than a single post-exposure number, and it separates concepts that raised intent from concepts that were simply shown to a high intent audience.

Keep the wording, scale, and anchor labels identical in both instances. If you change one word, you have measured two different things.

This is where most creative testing surveys quietly break, and it is the part almost nobody publishes. Write the anchors out.

Five-point purchase intent. Use this when you want a familiar, low effort item and you plan to report Top 2 Box:

  1. Definitely would not buy
  2. Probably would not buy
  3. Might or might not buy
  4. Probably would buy
  5. Definitely would buy

Seven-point agreement. This is a standard Likert scale format, and it gives you more room to detect differences between similar concepts:

  1. Strongly disagree
  2. Disagree
  3. Somewhat disagree
  4. Neither agree nor disagree
  5. Somewhat agree
  6. Agree
  7. Strongly agree

Three rules govern the choice.

  • First, label every point. Numbered-only scales and 0 to 100 sliders let respondents set their own reference points, which destroys comparability across cells.
  • Second, keep the scale balanced, with the same number of positive and negative options and one neutral midpoint.
  • Third, use the same scale for the same construct in every cell and every wave, because a five-point score and a seven-point score cannot be compared even after rescaling.

On scoring: Top 2 Box, the share of respondents choosing the top two options, is the standard readout for creative testing and it is what the ad testing guide uses.

It is easy to explain and it maps to a decision. Means are more sensitive to small differences but they hide distribution, so a concept that polarizes can post the same mean as a concept nobody cares about.

Report Top 2 Box as the headline and keep the mean and the full distribution alongside it. Pick one metric as your decision rule before you field, not after you see the data.

Plan sample as cells, not as completes. The working target for creative testing is about 200 completes per cell, which gives you enough base to read differences between concepts without overspending. If you want to check the margin of error that base buys you, run the number through a sample size calculator before you field.

The math flows from your exposure design. A monadic test of two concepts at 200 per cell needs 400 completes. A sequential monadic test of four concepts at 200 per cell needs 800, because sequential designs multiply cells by rotation position rather than collapsing them.

Build the budget as a table before you write a single question:

InputExample AExample B
Concepts24
DesignMonadicSequential monadic
Completes per cell200200
Total completes needed400800
Incidence rate in target audience50%25%
Screener starts required8003,200

Incidence rate is the input teams forget. If only a quarter of the general population qualifies for your screener, you need four times as many starts as completes, and that is what actually sets the cost and the field time.

Estimate incidence from a prior round or from category penetration data before you commit to a concept count, and check the targeting options on the market research panel to see how tightly you can define the audience without gutting your incidence.

Length is the other constraint on cell count. The monadic design guide frames survey length as concepts multiplied by metrics, and holds the total under 30 questions. Run that multiplication first. If it breaks the ceiling, cut concepts rather than cutting the diagnostics that tell you why a concept lost.

  1. Define the decision the test has to settle. Write one sentence naming which concept goes into which placement and what score would change your mind. If you cannot write it, you are not ready to field.
  2. Lock your concepts and your stimulus files. Finish the creative first. Match format, aspect ratio, and production polish across concepts so you are testing the idea, not the render quality.
  3. Choose the exposure design. Default to monadic. Move to sequential monadic only when sample is tight, and cap it at three concepts.
  4. Calculate cells and sample. Multiply concepts by completes per cell, divide completes by your incidence rate, and confirm the total fits your budget before you build.
  5. Draft the questionnaire in exposure order. Follow the 10-block sequence above. Randomize grid items and answer options, keep every scale fully labeled, and repeat the intent item verbatim before and after exposure.
  6. Set quotas and randomization. Balance cells on the demographics that matter to your category, and confirm the platform assigns concepts evenly rather than filling one cell first.
  7. Run a soft launch. Field about 10% of your total sample, then stop and check dropoff by page, median completion time, straightlining rates, and whether the stimulus rendered on mobile. Fix problems here, because after full launch the fix costs you the whole sample.
  8. Field to full sample and monitor daily. Watch cell balance and incidence against your estimate, and flag any cell that is filling far slower than the others.
  9. Clean before you analyze. Remove speeders below your time threshold, respondents who failed the trap item, and straightliners across the diagnostic grid. Document what you removed and why.
  10. Read Top 2 Box, then the shift, then the open ends. Rank concepts on your pre-declared metric, check the pre-post intent movement, then read the open ends for the reason behind the ranking.

You do not have to write a creative testing questionnaire from a blank page. A few starting points cover most rounds:

  • The ad copy testing survey template. The ad copy testing survey template has been used more than 3,000 times. It gives you a working question set for headline and copy variants, which is the fastest way to see how exposure and diagnostic questions fit together.
  • Ad testing question banks. The guide to ad testing surveys documents the standard question formula, a five-point Likert item you can copy verbatim, the five core ad metrics, and the 30-question ceiling for a single survey. Use it as your question library, then use this page to sequence what you pull from it.
  • Monadic design references. Cell math and randomization setup live in the guide to monadic versus sequential monadic design. Read it before you commit to a cell structure, because that decision sets your sample budget.
  • Concept testing methodology. If you are testing product or positioning ideas rather than finished ads, the four methodologies are laid out in the guide to concept testing.
  • Automated ad testing. SurveyMonkey LaunchPad handles the fielding side for image concepts: up to 10 PNG or JPG concepts per round, a monadic methodology by default, and results in as little as one hour up to 48 hours.
  • The wider research stack. Creative testing usually sits alongside brand tracking, pricing work, and concept screening, all of which are covered under market research solutions.
  • How do you write a creative testing survey?
  • What questions should you ask in a creative testing survey?
  • How many respondents do you need for a creative test?
  • What is the difference between monadic and sequential monadic testing?

The questionnaire is the part of creative testing you fully control.

Fix the exposure order, label every scale point, repeat the intent item before and after exposure, and size your cells against a real incidence rate, and the numbers you get back will hold up in the room where the media decision gets made.

Everything else in a creative test is downstream of those four choices.

Explore the product to see how automated ad testing fields a monadic creative test in as little as an hour.

Woman wearing a hijab, looking at research insights on laptop

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

A man and woman looking at an article on their laptop, and writing information on sticky notes

A logo testing survey only works with the right methodology. Compare monadic vs. comparative design, sample size, and scoring.

Smiling man with glasses using a laptop

Explore 30+ logo testing questions organized by funnel stage, design element, and methodology, then build your survey from a template.

Woman reviewing information on her laptop

See real concept testing examples across logos, packaging, names, and ads, plus how to turn each into a survey today.