Psychographic audience analysis: how to read the data you already collected

Turn survey responses into decisions by running a psychographic audience analysis that scores attitude batteries, sizes segments, and validates each cluster.

White outline of Goldie, the SurveyMonkey mascot

Summary:

  • Use top-two-box scores, means, and composite indexes to evaluate data; remove straightliners and reverse-worded errors before clustering.
  • Prove your clusters are real by confirming sufficient base sizes, tight within-segment agreement, and statistical separation between groups.
  • Use actual respondent verbatims to name your segments and report sizes with margins of error for actionable stakeholder decisions.

The field period closed, the responses landed, and you're now looking at a grid of agreement statements with no idea which ones mean anything.

A psychographic audience analysis is everything that happens after collection: scoring the attitude batteries, checking whether your clusters are real, sizing each one, and naming them in words your team will repeat without wincing.

This guide runs that sequence in order, from raw distributions to a segment map you can defend in a stakeholder review. Every threshold here is either published guidance or clearly marked as reasoning.

Read every attitude battery three ways before you commit to one number:

  1. top-two-box for the share who actively hold the attitude
  2. the mean for the direction of the full distribution
  3. a composite index for the construct underneath several items

When those readings disagree, the distribution is telling you something the average hides.

Before any of that, confirm what you fielded. Pull the questionnaire alongside your export and mark which items belong to which construct, which are reverse-worded, and which arrived as filler. Anyone still designing the instrument or collecting psychographic data has a different problem than the one on your desk.

Top-two-box answers a business question: how many people agree strongly enough that you'd change a message for them. Use it anywhere a stakeholder needs a share, not a score.

The mean reads the whole distribution, which makes it better for ranking items inside a battery.

Composite indexes, built by averaging three to five items that measure one construct, are what you feed a clustering routine, since a single statement carries wording noise that an index of four mostly cancels.

Composite index on a 5-point agreement scaleHow to read itWhat to do with it
1.00 to 1.99Actively rejects the attitudeReport it as a defining negative
2.00 to 2.74Leans againstUse as directional contrast
2.75 to 3.24No real position, or midpoint clusteringCheck the distribution before reporting
3.25 to 3.99Leans towardUsable, weak as a definition alone
4.00 to 5.00Owns the attitudeStrong candidate for a defining item

These cut points sit around the scale midpoint by arithmetic, not by published benchmark. Shift them if your scale runs seven points or your category agrees with everything.

On a 5-point scale, a recoded value is six minus the original. On a 7-point scale it's eight minus the original.

Do this before you average anything, or your index will average an attitude against itself and land politely in the middle.

Then correlate each recoded item against the rest of its construct. A negative correlation means you flagged the wrong item, or respondents read that statement as something else entirely.

Calculate each respondent's standard deviation across the battery.

Anyone at or near zero picked one column and rode it down. Reverse-worded items expose the second pattern: someone who agrees both that making coffee is the best part of their morning and that they'd rather skip it gave you acquiescence, not an attitude.

Flag both groups, then run your means with and without them. If removing 4% of the sample moves an item mean by half a point, that item was fragile and your write-up should say so.

Careful Likert scale design reduces the problem without removing it.

Two items can both average 3.0 and mean opposite things. One has 62% of respondents parked on neutral, which makes it dead weight. The other splits 41% disagree and 44% agree with a hollow middle, which makes it the most valuable statement in your battery. Polarizing items separate people, and separation is the whole job. Plot the five bars for every statement, then sort by how much mass sits at the ends.

Clustering routines never fail. They return the number of clusters you asked for, every time, whether or not those groups exist in the population. Your job in attitudinal segmentation analysis is to prove the solution earned its keep before anyone builds a campaign on it. If you're still weighing whether attitudes are the right basis, the broader view of segmentation types covers the alternatives.

What you're checkingWorking thresholdWhere the threshold comes from
Respondents per segment150 to 300Industry guidance (Articos)
Total sample for a segmentation study300 to 500 as a floorIndustry guidance (Segmentation Study Guide)
Segments, consumer studyThree to fiveIndustry guidance (Articos)
Segments, B2B studyTwo to fourIndustry guidance (Martal)
Smallest base you'll reportNo published figure. Set it where margin of error stops supporting the decision.Methodological reasoning
Within-segment agreementNo published figure. Defining items should show a tight spread.Methodological reasoning
Between-segment separationNo published figure. Gaps on defining items should clear significance testing.Methodological reasoning

Decide your minimum reportable base while you're still blind to the findings, because a threshold set after you've seen an interesting small group is not a threshold. Industry guidance commonly lands at 150 to 300 respondents per segment, which leaves room to crosstab that segment against a second variable. Run your intended bases through a sample size calculator and write the number in the analysis plan.

Share of sample and reportability are different tests. A group at 6% of a 1,200-person sample gives you 72 respondents: thin statistically, and awkward commercially, since you'd be asking a team to build a separate treatment for one person in 17. Small segments earn their place by carrying disproportionate value, not by being interesting.

A real segment is tight inside and distant outside. Inside, defining items should show a narrow spread, so most members answer them the same way. Those same items should also put distance between segments, with gaps large enough to clear significance testing at your actual bases. A group that agrees with itself but sits three points from its neighbor is a slice, not a segment.

Run several solutions instead of accepting the first output. Published guidance points to three to five segments for consumer work and two to four for B2B, mostly because past that point the groups stop being distinguishable to the people who use them. Compare solutions on three things: whether every group clears your base floor, whether defining items still separate, and whether you can describe each segment in one sentence without repeating yourself.

Reject and re-run when one cluster swallows more than half the sample, when any cluster falls under your base floor, when two clusters differ on only one item, or when the segments split cleanly along a demographic you also fed into the model. If your four attitudinal segments turn out to be age brackets wearing costumes, the attitudes did no work.

Item discrimination does most of the work. A statement pulling 41% agreement in one segment and 89% in another is doing exactly what you designed it for. A statement landing between 35% and 46% everywhere is a dead statement — it adds distance to no one and quietly pulls your clusters toward each other. Find these before clustering by checking the spread of top-two-box across preliminary groups, then drop them from the model even when the wording was somebody's favorite.

Battery length cuts both ways.

Too few itemsToo many items
Composite indexes rest on one or two statements eachRespondents fatigue and start straightlining
A single ambiguous phrase decides someone's segmentVariance reflects boredom, not attitude

If the last block of statements shows flatter answers than the first, length and order are shaping your data.

Scale points change what clustering can see.

  • Five points: easier for respondents, compresses variance
  • Seven points: gives distance-based routines more room, at the cost of a wider, mushier middle

Neither is wrong, but the choice is fixed in your dataset now, and it explains part of how tightly your segments hold. The mix of psychographic variables matters the same way — a construct measured with two items behaves differently than one measured with six.

Sample composition is the quietest driver. Quota skew, an over-represented source, or one channel that pulled a different kind of respondent will produce clusters that describe your fielding rather than your market.

  • Compare achieved composition against intended quotas before you interpret anything
  • For a validation wave, controlled audience panel sampling keeps composition from becoming the finding
  • Mode matters too: phone respondents agree more readily than self-administered ones, so mixed-mode studies need a mode check on every item

The last driver is over-inclusion. Throw attitudes, behaviors, category spend, and demographics into one distance calculation and the widest-ranging variables dominate, leaving you with a spend-based segmentation you never wanted. Cluster on attitudes. Profile with everything else afterward.

Everything in this section is an illustrative example built to show the arithmetic. These are not SurveyMonkey platform figures, benchmarks, or observed customer data. The numbers are constructed so you can follow each calculation and repeat it on your own export.

The setup: a hypothetical study of at-home coffee buyers in the US, 1,200 completed responses, a 14-item agreement battery on a 5-point scale, plus one open-ended question about what they look for in coffee. After recoding four reverse-worded items and removing straightliners, a four-cluster solution on six composite indexes produced groups of 384 respondents (32%), 336 (28%), 288 (24%), and 192 (16%).

Here's the attitude battery crosstab, showing top-two-box agreement by segment.

Attitude statement (% who agree or strongly agree)Ritual Seekers (n=384)Convenience First (n=336)Price Guardians (n=288)Origin Obsessives (n=192)Total (n=1,200)
Making coffee is part of a morning I look forward to91%22%41%78%58%
I'd rather spend less time on coffee and get going18%89%62%21%49%
I compare prices before I buy coffee44%51%93%38%57%
I want to know which farm or region it came from49%11%14%88%36%
I'll pay more for coffee that tastes better81%29%24%90%54%
Most coffee brands are basically the same38%44%46%35%41%

Now test the differences instead of admiring them. Take price comparison between Ritual Seekers at 44% and Convenience First at 51%. The standard error of that difference is the square root of (0.44 × 0.56 / 384) plus (0.51 × 0.49 / 336), or 3.7 points. At 95% confidence you need roughly 7.3 points to call it, and you have seven. That gap misses, so it belongs in your notes and not your deck. Price Guardians at 93% against Convenience First at 51% on the same item is a 42-point gap, which clears easily.

The last row teaches the harder lesson. Price Guardians at 46% versus Ritual Seekers at 38% is an eight-point gap with a standard error of 3.8, which technically clears the 95% threshold. It's still a dead statement: the range across four segments is 11 points and no segment takes a real position. Significance tells you a difference probably isn't noise. It says nothing about whether the difference is big enough to build on.

Pull the open-ended responses for each cluster separately and read the highest-frequency phrases. In this illustrative dataset, one group kept describing "my ten minutes of quiet before anyone else is up," which produced Ritual Seekers. Another wrote variations of "fast, hot, out the door." A third refused to pay four dollars for a bag of beans. The fourth listed roast levels unprompted. Coding verbatims by segment with text analysis surfaces those phrases fast, and names taken from respondent wording survive stakeholder review better than names invented in a workshop. Each named segment becomes the attitudinal core of a buyer persona other teams can use.

Sizing is share of sample applied to a population you can count, carried with its error range. Ritual Seekers at 32% of a 1,200-person sample have a margin of error of about 2.6 points, so the true share sits between 29.4% and 34.6%. Applied to an illustrative base of 240,000 customers, that's 70,600 to 83,000 people rather than a tidy 76,800. Report the range. Inside Origin Obsessives, at n=192, any percentage carries roughly ±7 points, which is why a five-cluster run that split that group into 124 and 68 got rejected: percentages in a 68-person cell swing about ±12 points.

  • If one segment holds more than half the sample, the solution under-fits. Re-run with more clusters and see whether the big group splits.
  • If a segment falls below your base floor, don't report it as a segment. Fold it into its nearest neighbor or hold it for a boosted wave.
  • If two segments differ on only one statement, they're one segment. Merge and re-run.
  • If segments line up with age, income, or region, the attitudes did nothing. Pull demographics out of the model.
  • If an item ranges less than about 10 points across segments, it's a dead statement. Drop it from the model, keep it for profiling.
  • If a respondent's standard deviation across the battery is near zero, they straightlined. Flag, remove, and re-run the means.
  • If someone agrees with both an item and its reverse, that's acquiescence. Treat the whole battery as unusable.
  • If a mean sits near the midpoint, look at the bars first. A flat 3.0 and a split 3.0 lead to opposite decisions.
  • If a gap doesn't clear significance at your bases, it goes in the appendix. Not the headline.
  • If you can't describe a segment in one sentence, the solution isn't finished. Go back to the verbatims.
  • What sample size do you need for psychographic segmentation?
  • How do you validate psychographic segments statistically?
  • Can psychographic segmentation predict real behavior?
  • How many psychographic segments should you have?
  • How do I test whether a segment difference is significant?
  • What do I do when a segment base is too small to report?
  • Should I weight my data before clustering?
  • How do I present segments to stakeholders who didn't commission the study?

The sequence holds whatever your battery looks like: recode reverse-worded items, screen out straightliners and acquiescence, read each item three ways, build composite indexes, cluster on attitudes only, test the solution against base size and separation, size each segment with its margin of error attached, and name the groups from what respondents wrote. Skip a step and you'll present a demographic split wearing an attitudinal name.

Ready for the next battery? Analyze your survey data and start turning attitude grids into segments your team can act on.