MaxDiff (maximum difference scaling) tells you what your audience values most, not what they say they do. Learn how best-worst scaling works and when to use it over a rating scale.
Summary:
When you ask survey respondents to rate a long list of product features on a 5-point scale, you rarely get the clarity you need. Most respondents simply rate everything as 'important,' leaving you with a sea of uniform data that makes it impossible to know what to prioritize.
MaxDiff analysis, or best-worst scaling, solves this by forcing respondents to make trade-offs, revealing clear, actionable preferences that standard rating scales miss.
In this guide, we’ll walk through how MaxDiff works, why it outperforms traditional methods for prioritization, and how you can use it to make confident product and marketing decisions.
MaxDiff (maximum difference scaling) is a survey technique that identifies the most and least preferred items in a set by forcing respondents to choose rather than rate.
The full name is maximum difference scaling. It's also called best-worst scaling. The methodology was developed by Jordan Louviere in the 1980s and has since become one of the standard tools in market research for any problem that comes down to priority ranking.
The mechanics are straightforward: respondents see small subsets of items (typically 3–5 at a time) drawn from a larger list. For each subset, they select the item they prefer most and the item they prefer least. Because they must commit to a specific best and a specific worst, they can't take the path that most survey respondents take on a standard importance scale: giving everything a 4 or 5 out of 5.
That forced choice is the point. The result is a set of preference scores that rank items relative to each other. You learn not just that customers prefer Feature A, but by how much they prefer it over Feature B, C, and D. That relative data is substantially more useful for prioritization decisions than a list of mean ratings where the top five items are all clustered between 4.1 and 4.4.
Common use cases for MaxDiff include:
Rating scales suffer from a systematic flaw: most respondents rate everything as important, leaving researchers with data that can't distinguish priorities.
It's a well-documented problem. When you ask people to rate 15 product features on a 5-point importance scale, you typically end up with a distribution clustered around 3.5–4.5. The data isn't wrong. It's just not useful.
Every feature looks roughly as important as every other, and the decision you were trying to inform (which three features to build next quarter) is no easier to make than before you ran the survey.
MaxDiff addresses this in four specific ways.
Acquiescence bias is the tendency to rate positively or agree regardless of actual preference. It's amplified on long lists because respondents shift into a pattern rather than evaluating each item individually.
MaxDiff's forced-choice format breaks the pattern. If a respondent has to pick a worst item for every subset, they can't avoid making real distinctions.
Cultures differ significantly in how they use rating scales. Some cultures systematically avoid extreme responses (rarely marking 1 or 5); others show stronger central tendency, clustering near the midpoint.
These tendencies make cross-market comparisons on rating scale data unreliable, because you can't tell whether a difference in average scores reflects a real preference difference or a cultural response style difference.
Because MaxDiff scores are derived from relative choices rather than absolute ratings, they hold up better in cross-cultural research.
Standard ranking questions become cognitively overwhelming above 7–8 items. Asking someone to rank 20 features from most to least important produces low-quality data because the cognitive load degrades judgment as respondents work down the list.
MaxDiff breaks the task into small, manageable subsets. Respondents make a series of simple 3-choice decisions rather than one overwhelming 20-way ranking. You can test 8–30 items reliably.
Utility scores derived from best-worst choices are more precise than mean ratings for the same sample size. You get more information per respondent because each best-worst choice provides data on how several items compare to each other, not just one item in isolation.
This question comes up often in research planning. The short answer:
| Research question | Best method |
| Which of these 20 features should we build first? | MaxDiff |
| How much would customers pay for Feature A if we also improved speed? | Conjoint |
| How satisfied are customers with their last experience? | Rating scale (CSAT) |
| Which of these 6 ad headlines resonates most? | MaxDiff |
| What price point maximizes revenue for a new product tier? | Conjoint or van Westendorp |
A block is a subset of items shown to a respondent in a single MaxDiff question. If you're testing 20 features and each block contains 4 items, a respondent might see 10–15 blocks across the survey. The block design controls which items appear together and how often each item appears across the full study.
Good block design ensures that every item appears roughly the same number of times across all respondents, and that item co-appearances (which items are shown together) are evenly distributed. This balance is what allows utility scores to be compared across items. If Feature A always appeared alongside weak alternatives, its score would be inflated by comparison rather than by genuine preference.
SurveyMonkey generates balanced block designs automatically. If you're building a MaxDiff study in a general survey tool without this feature, you'll need experimental design software or a statistician to construct the rotation manually.
For each block, the respondent makes two selections: the item they prefer most (best) and the item they prefer least (worst). This produces two data points per question rather than one, which is part of why MaxDiff is statistically efficient. A respondent completing 12 blocks of 4 items each makes 24 discrete choices, not 12.
The structure also pushes respondents toward genuine engagement. When someone must pick both a best and a worst from a block that includes items they generally like, they're making finer-grained distinctions than a rating scale would require.
Utility scores are calculated from the frequency with which each item is selected as best minus the frequency with which it's selected as worst across all respondents. An item selected as best by 70% of respondents and as worst by 5% has a very different utility score than an item selected as best by 35% and as worst by 30%.
The scores are relative, not absolute. They tell you that Feature A is preferred over Feature B by a specific margin, but they don't tell you whether either feature is "very important" in an absolute sense. For most prioritization decisions, that's exactly what you need to know.
The standard sample size guidance for MaxDiff studies is 200–300 complete responses per segment you want to analyze. More items require larger samples to maintain score stability. If you want reliable sub-group results (comparing feature priorities between enterprise buyers and SMB buyers, for example), plan for 200–300 responses per sub-group, not 200–300 total.
A useful rule of thumb: each item should be evaluated by at least 50–75 respondents across the full study. For a study with 20 items and 4-item blocks, 200 respondents typically provides adequate coverage. For 30 items or complex segmentation, 400 or more is more appropriate.
Calculate the sample size for your specific MaxDiff study.
MaxDiff and conjoint analysis are related but address different research questions. MaxDiff ranks single items (features, messages, benefits) by preference. Conjoint analysis evaluates multi-attribute product profiles, modeling how combinations of attributes (price, brand, and features simultaneously) influence preference and purchase intent.
Use MaxDiff to answer "which of these features is most important to our customers?" Use conjoint to answer "how much revenue would we give up by removing Feature A, and would a lower price offset that loss?" If your research question involves trade-offs between multiple attributes of the same product, conjoint is the right tool. If your research question is fundamentally a ranking problem across a list of items, MaxDiff is typically simpler, faster, and easier to explain to stakeholders.
A MaxDiff study requires five decisions: your item list, block design, sample size, respondent targeting, and scoring approach.
Here's a step-by-step process.
MaxDiff works for 8–30 items. Below 8 items, a simple ranking question is sufficient and easier for respondents. Above 30 items, respondent fatigue degrades data quality and sample size requirements increase sharply.
Before you finalize the list, trim to the most relevant items. An item you're 90% sure will rank last is still worth including if the result would change a decision. An item you included "just to see" probably doesn't belong. Every item added requires more respondent time and more sample to maintain score stability.
Every item should be written in the same grammatical structure and at roughly the same level of specificity. If some items are specific ("Two-day delivery guarantee") and others are vague ("Better shipping"), your scores will reflect the writing quality as much as the actual preference. Use consistent structure throughout: noun phrase, benefit statement, or short imperative. Avoid adjectives that prime positive responses ("industry-leading," "advanced") in some items but not others.
With SurveyMonkey, this step is handled automatically. The platform balances item appearances and co-appearances so your utility scores are comparable across the full item list. If you're using a different tool without automated block design, you'll need to generate a balanced incomplete block design, which requires statistical software or a pre-built design tool.
Using SurveyMonkey Audience, target by the attributes that define your target segment: job title, industry, company size, purchase behavior, or demographics. A study comparing two segments (enterprise and SMB buyers, for example) needs 200–300 responses per segment, not 200–300 total. Set quotas to ensure even distribution, and monitor during fielding to catch unexpected drop-off.
A well-designed MaxDiff study takes 5–10 minutes to complete. Monitor the median completion time and the completion rate. If respondents are dropping off before finishing, the study is either too long or some item descriptions are unclear enough to cause confusion.
SurveyMonkey provides automated scoring. Review the full ranking, but focus primarily on the top 5 and bottom 5 items. Those are typically where the actionable decisions are: the top items are your priorities; the bottom items are candidates for cutting or deprioritizing. The middle of the distribution is often where items cluster without meaningful differentiation.
Average utility scores across very different groups often produce misleading results. A feature that scores in the middle overall might be the top priority for one persona and near the bottom for another. Run the scoring separately for each key segment before drawing conclusions. If segment-level results diverge significantly, you likely have a personalization or product line decision in front of you, not just a feature prioritization question.
Eight to 30 is the practical range. Below 8 items, a simple ranking question works just as well and is easier for respondents to complete. Above 30 items, respondent fatigue degrades data quality and the sample size required for stable scores becomes impractical for most research budgets.
If you have more than 30 candidate items, run a qualitative screen first to narrow the list before fielding a MaxDiff study.
A ranking question asks respondents to order all items simultaneously from most to least preferred. This works reasonably well for 5–7 items but becomes cognitively overwhelming above that threshold, producing lower-quality data as respondents fatigue toward the end of the list.
MaxDiff breaks the task into small subsets (3–5 items at a time), which produces statistically more reliable scores for longer lists because each individual decision is simple and more deliberate.
The standard guidance is 200–300 complete responses per segment you plan to analyze. If you need reliable sub-group results (by age, industry, role, or any other variable), plan for 200–300 per sub-group.
Under-sampling a segment produces unstable utility scores where the rankings shift significantly when you add more data, which undermines your ability to make confident decisions from the results.
Not with SurveyMonkey. The platform handles block design and utility score calculation automatically, removing the two steps that historically required statistical expertise.
For complex studies involving multiple segments, follow-up conjoint analysis, or custom sub-group modeling, expert research services are available through SurveyMonkey to support more advanced analytical needs.
SurveyMonkey offers MaxDiff studies through its market research platform, combining automated study design with access to a 335M+ person panel.
Running MaxDiff through SurveyMonkey removes the statistical overhead that used to make best-worst scaling inaccessible to teams without a dedicated research function. Specifically:
How customers use it
ClickUp scored a winning Super Bowl commercial by testing multiple versions with over 5,000 respondents in less than a week, identifying the most memorable creative direction.
Kajabi implemented SurveyMonkey market research solutions to conduct comprehensive Usage and Attitudes studies and message testing, resulting in increased customer conversion and lower attrition rates.

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

Learn the four types of customer data businesses collect, how to gather each one responsibly, and how to stay GDPR and CCPA compliant.

VoC marketing turns customer feedback into sharper messaging and campaigns. Learn how to use Voice of the Customer data and try it for free.

Learn what creative testing is, the metrics it measures, and how to test ads, video, and messaging before you launch a campaign.