Jobs to be done: a practical guide to the JTBD framework

Learn the jobs-to-be-done framework and how to validate it with a survey, then try a free product testing template today.

White outline of Goldie, the SurveyMonkey mascot

Summary:

  • JTBD frames products as "hired" tools to achieve functional, emotional, and social progress, creating a stable foundation that outlasts temporary personas.
  • Successful implementation moves beyond qualitative interviews to quantitative validation, replacing anecdotal evidence with data-driven market opportunities.
  • The framework guides the product lifecycle by ranking opportunities, screening new concepts, and prioritizing features that solve the specific "job" customers need done.

Jobs to be done starts from an odd but useful premise: people don't buy products, they hire them to get a job done.

A parent doesn't want a drill, they want a hole in the wall for a shelf their kid keeps asking about.

Once you see products as hired help instead of finished goods, the whole way you research, prioritize, and build changes.

This guide covers what the jobs-to-be-done framework actually claims, how to run a JTBD process from interview to statistically valid survey, and how to carry the output into concept testing and feature prioritization instead of letting it die in a slide deck.

Jobs to be done isn't a SurveyMonkey invention, and treating it like generic "customer insight" undersells it.

The framework traces back to Tony Ulwick's Outcome-Driven Innovation work in the 1990s, which reframed product strategy around the measurable outcomes customers are trying to achieve rather than their stated feature requests.

Clayton Christensen popularized the theory more broadly, most notably through the milkshake study he wrote about with Bob Moesta: a fast-food chain wanted to sell more milkshakes, so it asked frequent buyers what they'd change about the drink.

Thicker. Sweeter. More flavors. None of it moved sales.

Only when researchers watched who was actually buying milkshakes, at what time of day, and what else they weren't buying instead, did the real job surface: commuters wanted something filling and one-handed to make a boring drive more interesting.

That's a "job," and it has almost nothing to do with milkshake flavor.

The core claim of JTBD is that a job is stable even when the products that do it change. People have hired candles, then gas lamps, then light bulbs, then LEDs, all to do the same job: push back darkness in a room.

If you build your roadmap around "how do we improve the candle," you eventually lose to whoever notices the job is lighting, not waxing.

This is also why job statements outlive personas. A "harried commuter parent" persona ages and drifts with the person; the job "get my kids out the door on time without a fight" stays put.

  • Functional. The practical task itself, like getting from point A to point B or reconciling a spreadsheet.
  • Emotional. How the person wants to feel, or wants to avoid feeling, while getting the job done, like reducing anxiety before a big presentation.
  • Social. How the person wants to be perceived by others while doing it, like looking competent in front of a client.

A single purchase usually answers all three at once. Buying a minivan solves a functional job (haul four kids and their gear), an emotional job (feel less frazzled on the drive), and a social job (not feel like you've given up on having a personality). Miss any one of the three and you get a product people rate fine in a survey and quietly stop using.

A JTBD process has a predictable failure mode: teams do six or eight qualitative interviews, write some evocative job statements, and then bet a quarter of roadmap on what a handful of people said in a room. The interviews are necessary. They are not sufficient. Here's the fuller sequence.

Start broader than your product. If you sell project management software, the job isn't "manage projects," it's closer to "know, at a glance, whether my team is going to miss a deadline before it's too late to fix it."

Scoping too narrowly, around your existing feature set, just reproduces what you already built.

Talk to people who recently started, switched away from, or seriously considered a solution.

Ask them to walk through the timeline: what was happening right before they went looking, what they tried first, what almost worked but didn't, and what finally got them to act.

You're listening for the moment of struggle that triggered the search, not for feature preferences.

Convert what you heard into a job statement (see the FAQ below for the exact format) and a list of desired outcome statements: specific, measurable statements of progress, such as "minimize the time it takes to confirm a shipment has cleared customs."

Ulwick's outcome-statement format works well here because it's specific enough to put in front of a survey respondent without translation.

This is where a lot of jobs-to-be-done work quietly stalls.

Eight interviews tell you a job exists; they don't tell you how many of your prospects share it, how badly, or how satisfied they are with current alternatives.

Take the outcome statements from your interviews and field them to a properly sized sample of your target market, asking two questions per outcome: how important is this to you, and how satisfied are you with how well you currently achieve it.

Score the gap, importance minus satisfaction, weighted, and you get a ranked, defensible opportunity list instead of a hunch.

The part people skip is sizing that survey so the ranking actually means something.

If your target market has 40,000 people in it and you only need directional confidence, a few hundred responses might do.

If you're deciding where to put a seven-figure engineering investment, you want a much tighter margin of error. 

The SurveyMonkey sample size calculator and margin of error calculator both handle that math without requiring a statistics course, so you can quote a real confidence interval instead of "we talked to some customers."

Once the survey is in, rank outcome statements by opportunity score.

Ulwick's formula weights underserved outcomes over merely important ones: opportunity equals importance plus the gap between importance and satisfaction, so an outcome rated highly important but poorly satisfied rises to the top, while one that's important and already well satisfied by existing options sinks toward the bottom.

Treat the top decile as your roadmap shortlist. This is the step that turns JTBD from a workshop exercise into an input a finance team will actually respect.

The first is treating a feature request as a job.

"I want a dark mode" is a solution someone proposed, not the underlying job; ask why, twice, and you usually land closer to something like "I want to use this at night without waking up my partner."

The second is running the qualitative interviews well and then skipping the quantitative step entirely, because the interviews already felt conclusive.

Eight good conversations can surface the right jobs and still get the ranking wrong, since the loudest or most articulate interviewee isn't necessarily representative of your broader market.

The jobs-to-be-done framework tells you what to build. It doesn't tell you which idea to greenlight first, which concept actually lands, or which features earn their place on the roadmap. Those are three separate research problems, and they all fall under the broader market research use case that SurveyMonkey LaunchPad is built to handle, with a purpose-built solution for each stage.

Once your quantitative survey has ranked outcome statements by opportunity score, you'll usually have more high-opportunity jobs than you can address at once.

Idea screening lets you put several early concepts, each aimed at a different underserved job, in front of your target audience and see which one people actually prefer before you write a line of code.

A concept can look appealing in isolation and still fail the job it's supposed to do.

Concept testing lets you validate whether a specific product concept addresses the functional outcome you identified and whether it lands on the emotional or social dimension too, benchmarked against category norms instead of your own gut feeling.

Feature backlogs tend to grow by committee, not by evidence.

MaxDiff and feature prioritization forces respondents to choose between competing options rather than rate everything a seven out of 10, producing a forced rank of which features move the needle on the jobs you validated earlier.

It's the quantitative cousin of the opportunity score, applied at the feature level instead of the outcome level.

All three solutions draw on the same underlying audience panel, so you're testing against the same target market you scoped the jobs against in the first place, not a convenience sample of whoever answered your email.

Each stage also comes back with results benchmarked against category norms, so "this concept scored a seven out of 10 on appeal" turns into "this concept scored above the 90th percentile for its category," which is the kind of number that actually moves a go or no-go decision.

Here's how the full sequence plays out, using an illustrative company we'll call Farrow Kitchen, a mid-size meal-kit brand. This example is constructed to show the method, not a documented SurveyMonkey customer case.

Farrow's churn was climbing among subscribers in their first three months, and the standard exit survey kept returning "too expensive" as the top reason, which didn't match the fact that most churned subscribers had used every box before canceling.

The team ran 12 switch interviews with recent cancellers and heard a different story: the job wasn't "eat healthy dinners," it was "get dinner on the table on a weeknight without a decision-making tax after an exhausting day."

Price was the reason people gave; decision fatigue was the reason people left.

The team wrote eight outcome statements from those interviews, among them "minimize the number of choices I have to make between opening the box and eating," and fielded them to 900 current and lapsed subscribers using a properly sized survey.

The gap analysis showed that outcome scored in the top two for importance and dead last for satisfaction: a clear, high-opportunity job nobody had prioritized because the churn survey kept pointing at price.

With that ranked list in hand, Farrow ran idea screening on three concepts aimed at the decision-fatigue job (a default weekly plan, a chef's-pick auto-selection, and a simplified three-recipe rotation), then concept-tested the winner against category benchmarks before building it.

MaxDiff on the surrounding feature set confirmed that "fewer choices, faster" outranked several features engineering had already assumed were higher priority.

The 90-day churn rate improved after the change shipped, an outcome the original exit survey had never pointed toward.

The team also went back and checked the eight outcome statements against the concept-testing results, and the two mapped cleanly: the winning concept scored highest specifically on the decision-fatigue outcome, not on general appeal, which is the kind of alignment that's easy to miss if concept testing and the JTBD survey are run as unrelated projects instead of one continuous pipeline.

Teams often reach for jobs to be done, personas, and user stories to solve the same meeting, and then wonder why the artifacts contradict each other. They answer different questions.

Jobs to be doneCustomer personasUser stories
What it capturesThe progress someone is trying to make, independent of who they areA composite sketch of a customer segment: demographics, attitudes, behaviorsA specific interaction with a product feature, from a user's point of view
Typical format"When [situation], I want to [outcome], so I can [benefit]""Maria, 34, marketing manager, values efficiency and hates clutter""As a [role], I want [feature], so that [outcome]"
DurabilityStable for years; the job rarely changes even as products doDrifts as the market and the individuals it's modeled on changeExpires once the feature ships or gets reworked
Best used forStrategy, roadmap prioritization, positioningMarketing tone, messaging, sales enablementSprint planning, backlog tickets
Main risk if misusedTreated as a one-time workshop output instead of a validated, ranked listMistaken for real segmentation data instead of a communication aidMistaken for strategy when it's really a task description

None of the three replaces the others. A healthy setup uses JTBD to decide what's worth building, personas to help teams talk about who it's for, and user stories to break the validated concept into sprints.

  • What is an example of jobs to be done?
  • What are the three types of jobs to be done?
  • How do you write a jobs-to-be-done statement?
  • What is the difference between jobs to be done and customer personas?

The jobs-to-be-done framework earns its keep the moment you stop treating it as a workshop exercise and start treating it as a research pipeline: interviews to surface candidate jobs, a sized survey to rank them by opportunity, and a testing stage to confirm the concept actually lands. 

Skip the middle step and you're back to betting the roadmap on whoever interviewed the loudest customer.

Start with the outcome statements you already have, or the ones you're about to collect, and put them in front of the audience that matters.

Apply with a template: try the product testing survey template to validate your next concept against the job it's meant to do.

Woman wearing a hijab, looking at research insights on laptop

SurveyMonkey can help you do your job better. Discover how to make a bigger impact with winning strategies, products, experiences, and more.

A man and woman looking at an article on their laptop, and writing information on sticky notes

Ethnographic research observes behavior, not just self-reports. Learn the definition, process, key terms, and how to pair it with survey data.

Smiling man with glasses using a laptop

Learn how to design a creative testing survey, from exposure order and scale anchors to cell sizing and fielding QA that keeps results readable.

Woman reviewing information on her laptop

A logo testing survey only works with the right methodology. Compare monadic vs. comparative design, sample size, and scoring.