Ad testing survey questions: the complete question bank
A complete bank of ad testing survey questions, organized by category, with response types, so you can build a rigorous test fast.
If you searched for ad testing survey questions, you probably already have a creative sitting in a folder and a deadline sitting on your calendar. What you need isn't a definition of ad testing. You need the actual questions, organized so you can pull the right ones for your format and skip the ones that won't tell you anything useful.
This guide gives you more than 30 distinct ad testing survey questions, grouped into nine functional categories, each tagged with a recommended response type so you know exactly how to build it.
After the question bank, you'll find a table mapping each category to the business decision it actually informs, a section on survey architecture (sequencing, monadic versus comparative design, sample size, and format-specific adjustments), and answers to the questions research teams ask most often about pre-tests, post-tests, and tracking studies.
Each question below includes a suggested response type. Ratings scales are typically five- or seven-point Likert or semantic differential scales unless noted.
Pull the questions relevant to your ad format and research goal rather than using every single one; more on right-sizing your survey is in the taxonomy section below.
If you want a faster path, you can import these questions into the ad and copy testing survey template and start editing right away. If you want the full context on methodology and timing, read the guide to ad testing first.
These questions confirm that viewers understood what you intended them to understand, before you spend a media budget finding out the hard way.
These questions capture the feeling the ad leaves behind, which is often a stronger predictor of recall and sharing than the message itself.
Brand recall questions test whether people connect the ad to the correct brand, which matters more in cluttered channels like social feeds and audio.
Relevance questions predict whether the right people will pay attention, which is a separate question from whether the ad is well-made.
These isolate which specific creative element is carrying or dragging down performance, so revisions target the right piece instead of the whole ad.
These questions tie creative performance to the outcome that ultimately matters to the business: whether the ad moves someone toward a purchase.
Use these only in a comparative or sequential design, where respondents see more than one concept. They answer a different question than the categories above: not "is this good" but "which is best."
Format changes what's worth asking. Layer these on top of the core categories above rather than replacing them.
That's 41 distinct questions across nine categories and response types ranging from open text to ranking, comparative choice, and rating scales. You won't use all 41 in any single study, and you shouldn't try to.
A launch decision for a single finished commercial calls for a different subset than an early screen of six static concepts, and a mid-flight tracking wave calls for a different subset again. The taxonomy section below explains how to select a workable subset for your specific situation, including a sensible ceiling on total question count.
Every category above maps to a specific business risk. Skipping a category doesn't just leave a gap in your report; it leaves a specific, nameable decision unsupported.
| Question category | Business outcome or KPI it informs | Risk of skipping it |
| Comprehension and message clarity | Whether the campaign message will land as intended at scale | You launch an ad that entertains but doesn't communicate the offer, and conversion suffers with no clear explanation why |
| Emotional response and ad appeal | Ad recall, watch-through rate, and social sharing potential | You ship a technically correct ad that nobody remembers or feels anything about, wasting media spend on low-recall impressions |
| Brand recall and brand linkage | Share of voice and media efficiency, especially in cluttered or short-format channels | The ad performs well in testing but gets misattributed to a competitor or category in market, so the spend builds someone else's brand |
| Relevance and audience fit | Targeting accuracy and media plan efficiency | You optimize creative for a general audience and miss the segment most likely to convert, inflating cost per acquisition |
| Creative-element diagnostics | Which specific asset (headline, visual, CTA) needs revision before the next production cycle | You greenlight a full reshoot or redesign when only the call to action needed a fix, burning budget and time |
| Purchase intent and persuasion | Projected lift in consideration or conversion, used to forecast campaign ROI | You launch a well-liked ad that never translates into purchase intent, and the campaign underperforms against revenue goals |
| Comparative and concept-selection | Which concept to greenlight for production or media spend | Without a forced-choice or ranking mechanism, you're left comparing average scores across separate monadic cells, which understates true preference strength |
| Channel-specific variants | Whether creative needs adaptation before running across video, social, audio, and static placements | A concept validated only as video gets repurposed as a six-second social cut or audio spot with no evidence it still works in that format |
| Pre-test, post-test, and tracking sets (see taxonomy below) | Whether the ad is working before spend, during the flight, and after it, and whether wear-out is setting in | You catch a problem only after the media budget is spent, or fail to notice an ad wearing out mid-flight |
The question bank above only works if it's assembled into a survey that respects a few structural rules. Get these wrong and even a perfectly worded question set will produce noisy or misleading data.
Order matters more in ad testing than in most other survey types, because early questions can prime answers to later ones. Follow this general sequence:
A monadic design shows each respondent exactly one ad concept and asks about it in isolation. A comparative (or sequential monadic) design shows the same respondent multiple concepts, either side by side or in sequence, and asks them to choose or rank.
| Monadic | Comparative | |
| Setup | One concept per respondent, judged in isolation | Multiple concepts, side by side or sequential |
| Use when | Benchmarking vs. norms; deep per-concept diagnostics; <10 developed concepts | Screening many early-stage ideas; fast directional read; need a head-to-head winner |
| Watch for | No relative ranking | Order bias; randomize concept order |
Choose monadic when you want to benchmark a concept against industry norms, need in-depth diagnostic feedback per concept, or are testing fewer than 10 well-developed concepts.
Choose comparative when you're screening a larger set of early-stage ideas, need a fast directional read, or specifically need to know which concept wins head to head rather than how each performs on its own.
Comparative designs introduce order bias (whichever concept appears first or last can get an artificial lift), so randomize concept order across respondents whenever your survey platform allows it.
Two screening decisions shape everything downstream.
For example: For video, require full playthrough or use a forced-view player rather than trusting a self-reported "I watched it." For static ads, consider a brief timed exposure (five to seven seconds) before the ad disappears, which better simulates a real scroll-past than letting people study the image indefinitely.
For a single-concept monadic test, aim for a minimum of 100 completed responses per concept to get directionally reliable results, with 200 or more per concept if you plan to compare subgroups or need statistical significance for a go/no-go decision.
For comparative or ranking designs among two to four concepts, 150 to 250 total respondents is usually enough, since every respondent evaluates every concept. Smaller niche audiences may require accepting directional rather than statistically significant results; say so explicitly in your report rather than implying false precision.
| Design | Target |
| Single-concept monadic | 100+ completes per concept |
| Monadic with subgroup analysis | 200+ per concept |
| Comparative, 2–4 concepts | 150–250 total |
Sample size and confidence level trade off against each other, so decide up front how much certainty the decision actually requires.
A concept screen meant to cut a list of eight ideas down to three can tolerate a wider margin of error than a final go/no-go on a media buy in the millions.
When the stakes are high, it's worth running the numbers through a sample size calculator before fielding, rather than guessing at a round number and hoping it holds up under scrutiny.
More questions do not mean more insight.
A strong ad test typically runs 12 to 18 questions total: a short screener, one brand-linkage question, three to five diagnostic and emotional-response questions, one or two purchase-intent questions, and one open-text question.
Once a survey exceeds roughly 20 questions, response quality degrades and later answers become less reliable as attention fades.
Pick the questions from the categories above that map to your specific decision, rather than including every category to be thorough.
Once you're past about 20 questions, expect fatigue to start affecting answer quality, especially toward the end of the survey. If your list from the categories above is pushing past that number, cut comparative or channel-specific questions that don't apply to your format before cutting core diagnostic questions like brand linkage or purchase intent.
A pre-test survey runs before an ad goes live, using concepts, storyboards, or finished creative shown in a controlled setting; its questions (comprehension, appeal, purchase intent) exist to inform a go, revise, or kill decision before media dollars are spent.
A post-test survey runs after a campaign has aired, typically among people who were actually exposed to the ad in market; its questions shift toward aided and unaided recall, message takeaway, and any measurable shift in brand consideration compared to a control group who wasn't exposed.
A tracking study runs continuously or at set intervals throughout a flight, repeating a consistent core set of questions (usually recall, favorability, and purchase intent) so you can plot a trend line and catch wear-out or fatigue while there's still time and budget to adjust the media plan.
A leading question smuggles in an assumption or a desired answer. "How much did you love this exciting new ad?" assumes the respondent loved it and pre-labels the ad as exciting. Rewrite it as a neutral scale question: "How would you rate this ad?" with a balanced scale from very negative to very positive. Watch for three common patterns: scales that skew positive (starting at "good" instead of a true negative anchor), questions that name a desired attribute before asking about it, and forced-choice questions that omit a legitimate "neither" or "not sure" option. If you're not sure whether a question is leading, read it aloud and ask whether a respondent could tell which answer you're hoping for; if the answer is yes, rewrite it.
For a pre-test that only evaluates concepts in isolation, a control group is optional; you're comparing concepts against each other or against category benchmarks, not against a "no ad" baseline.
For a post-test or tracking study, a control group matters much more. Without one, you can't tell whether a lift in brand favorability or purchase intent came from the ad itself or from something else entirely, like a seasonal demand shift or a competitor pulling back their own spend.
A simple control group asks the same core questions (brand recall, favorability, purchase intent) to a comparable group of people who weren't exposed to the campaign, so you have a clean baseline to measure the exposed group against.
Use rating scales when you need a number to track over time, compare across concepts, or report in a scorecard. Use open text when you need to know why a rating landed where it did, or when you can't predict the range of likely answers ahead of time (comprehension checks and "what would you change" questions are almost always better as open text).
A well-built ad test usually pairs a small number of open-text questions with a larger set of scaled questions, rather than leaning entirely on one type.
You don't need to build this from a blank page. Start with the questions in the categories above that map to your specific decision (a launch go/no-go, a concept screen, a mid-flight wear-out check) and drop them into a working survey today.
Import these questions into the ad and copy testing survey template to get a ready-to-edit starting point, or explore LaunchPad's image ad testing if you want built-in scorecards, benchmarks, and panel access instead of building the analysis yourself. For the full methodology behind pre-test timing, sample selection, and reporting, read the guide to ad testing. And if ad testing is one piece of a broader research program, see how it fits alongside other market research use cases.
Whichever questions you pull first, resist the urge to ask everything. A tight set of well-sequenced questions, matched to the decision you actually need to make, will tell you more than a long survey ever will.

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

Learn how an accessibility audit turned into a company-wide brand refresh

A diary study is a qualitative research method where people log experiences over time. Learn when to use one and see real examples.

Learn how to run a win-loss analysis with a repeatable framework, real interview questions and a free template. No CI vendor required.