Users say they'd pay for Feature X in a survey; your analytics show nobody touches the similar features you already shipped. The contradiction isn't a data error — it's a measurement gap between stated preference and revealed behavior, and it resolves once you stop treating either source as ground truth and start scoring both against the underlying job to be done.

Quick answer: The say-do gap exists because surveys measure what people believe about their future behavior, while analytics measure what they actually did. Neither is "wrong" — they're answering different questions. Reconcile them by scoring the underlying job's importance and satisfaction, not by picking a winner.

Why Surveys and Analytics Tell Different Stories

Surveys and analytics disagree because they measure two different things: stated intent versus revealed behavior. A survey captures what a respondent believes, remembers, or wants to signal about themselves; usage data captures what actually happened when the same person faced real friction, real pricing, and a real competing task. The gap is the distance between those two truths.

This isn't a niche annoyance — it's one of the best-documented phenomena in research methodology. Economists call it the attitude-behavior gap; UX researchers call it the say-do gap. Both describe the same mechanism: self-reported intent is a weak predictor of actual behavior, especially for anything involving future hypotheticals ("would you use...") rather than recalled specifics ("what did you do last time...").

A few concrete forces widen the gap:

  • Social desirability bias. Respondents answer in ways that make them look forward-thinking, responsible, or sophisticated — "yes, I'd use an advanced reporting feature" sounds better than admitting they never open reports at all.
  • Optimistic forecasting. People consistently overestimate their future engagement with anything that sounds useful in the abstract, because imagining a feature is frictionless in a way that using it never is.
  • Sampling mismatch. The survey respondents and the analytics population are rarely the same cohort — your most vocal users are overrepresented in surveys and may not resemble your median active user at all.
  • Question framing. "Would you find X valuable?" invites agreement almost regardless of X; it's a leading question dressed as research.

None of this means surveys are useless. It means a survey answer is a hypothesis about a job, not a verified feature request — and the distinction matters enormously for what you do next.

The Classic Example: Feature Requests vs. Feature Usage

The pattern shows up constantly in prioritization backlogs. Fifty customers ask for CSV export in a survey or support ticket. You ship it. Usage analytics six weeks later show 4% adoption. The team concludes the survey "lied" — but more often, the survey correctly identified a real anxiety ("I might get locked into this tool") that CSV export only partially solves, so people stopped asking once they felt safe, whether or not they ever clicked the button.

The Cognitive Science Behind the Gap

The say-do gap isn't a research-design flaw you can eliminate with better wording — it's rooted in how memory and self-perception actually work. Psychologist Daniel Kahneman's distinction between the experiencing self and the remembering self explains why: people don't report what they experienced, they report a reconstructed, edited summary of it, shaped by the most intense moment and the ending, not the average.

That means a survey question about "how often do you use advanced filters" isn't retrieving a log file from memory — it's retrieving a story the respondent has told themselves about their own competence and needs. The story is usually more flattering and more forward-looking than the log file would show.

Jakob Nielsen, co-founder of the Nielsen Norman Group, has argued for decades that usability research should weight what users do over what they say they'd do, precisely because self-report is such an unreliable predictor of behavior in front of an actual interface. His guidance isn't "ignore surveys" — it's "never let a stated preference override an observed one without investigation."

A few mechanisms worth naming explicitly, because they show up differently in different research artifacts:

  1. Recall distortion — usage frequency is almost always overestimated in retrospective self-report, especially for anything more than a few days old.
  2. Aspirational identity — people answer as the version of themselves they want to be (organized, data-driven, security-conscious), not the version logged in your database.
  3. Demand characteristics — respondents infer what answer the researcher "wants" and drift toward it, especially in moderated interviews.
  4. Loss aversion around omission — it feels safer to say "yes, I'd want that" than to say "no," because saying no risks losing something later.

None of these are reasons to distrust your respondents. They're reasons to treat what people say as one input describing a job's shape, and what people do as a second, independent input describing that job's actual friction — and to expect the two to diverge by default, not as an exception.

Diagnosing Which Signal to Trust

The right response to a say-do gap is never "trust analytics, ignore the survey" or vice versa — it's to identify why they diverge, because the reason tells you what to build. A gap caused by low awareness needs a different fix than a gap caused by genuine lack of demand, and the two look identical in a dashboard.

Start by classifying the contradiction. In practice, most say-do gaps fall into one of four buckets:

Gap typeSurvey saysAnalytics showsWhat's actually happeningWhat to do
Awareness gap"I want X"Feature X exists, near-zero usageUsers don't know it shipped, or can't find itFix discovery/onboarding before touching the feature
Friction gap"I want X"Users start the flow, abandon earlyThe job is real; the execution has too much frictionRedesign the flow, don't kill the feature
Aspiration gap"I want X"No usage even after fixing awareness and frictionStated desire was identity signaling, not a real jobDeprioritize; the "job" doesn't exist at meaningful scale
Cohort gap"I want X" (from survey respondents)Low usage (from full active base)Survey sample isn't representative of the broader user baseRe-segment before concluding anything

The table matters more than any single metric because it turns "the data is contradictory" into a falsifiable question you can actually go answer: was it discoverable? Was it usable? Was it wanted? Was it the same people?

A Quick Triage Checklist

Before declaring a say-do gap resolved, confirm you've ruled out the boring explanations first:

  • Did the feature actually ship to the segment that requested it, and were they told?
  • Is the usage event correctly instrumented — is "opened the modal" being counted as "used the feature"?
  • Was the survey question about a hypothetical ("would you use") or a recalled behavior ("did you use")?
  • Are the survey respondents demographically or behaviorally similar to your analytics population?
  • Has enough time passed for habit formation, or are you measuring week one against a feature nobody's discovered yet?

If you can't confidently rule these out, you don't have a say-do gap — you have a measurement problem, and no amount of prioritization theory will fix that. This is also where structured research synthesis earns its keep: coding interview and survey data against consistent tags makes it possible to check whether the "want" is really coming from the same job, or from several different jobs that happen to use the same words.

From Contradiction to Structured Opportunity Scoring

A say-do gap becomes useful the moment you stop asking "which source is right" and start asking "how important is the underlying job, and how well is it currently satisfied." That reframing is the entire premise of Tony Ulwick's Outcome-Driven Innovation (ODI), developed at the consultancy Strategyn.

The ODI formula: Opportunity Score = Importance + max(Importance − Satisfaction, 0) A job rated highly important but poorly satisfied scores highest — flagged as underserved and worth investigating, regardless of how loudly (or quietly) it showed up in a feature-request tally.

Under ODI, a survey doesn't ask "would you want X" — it asks respondents to rate the importance and current satisfaction of specific, measurable desired outcomes tied to a job ("minimize the time it takes to reconcile a monthly report"). Analytics then independently corroborates or contradicts satisfaction.

If a job is rated highly important and poorly satisfied, but usage of the closest existing feature is near zero, that's the awareness or friction gap from the table above — not proof the job doesn't matter. Ulwick has reported that ODI-guided product bets succeed at rates running several times higher than the roughly 17% success rate typically cited for new product launches industry-wide. That's directionally consistent with scoring outcomes rather than counting feature votes, even if the exact multiplier varies by study and industry.

Why Importance-vs-Satisfaction Beats a Feature Vote Count

A raw tally of "50 people asked for CSV export" collapses two different signals into one number. Opportunity scoring keeps them separate on purpose:

SignalWhat it capturesSay-do gap risk
ImportanceHow much the underlying outcome matters to the respondent's jobLow — people are relatively accurate about what matters to them
SatisfactionHow well current tools/features address that outcomeMedium — colored by recency and mood at time of survey
Stated feature requestA specific solution the respondent imagines would helpHigh — conflates the job with one guessed solution
Feature usage rateWhat people actually did with a shipped solutionLow on intent, but blind to unshipped or undiscovered solutions

Read across the rows: the first two capture the job, which is durable and comparatively reliable; the last two capture a guessed or shipped solution, which is where most of the say-do noise concentrates. Score the job, then treat both the stated feature request and the usage rate as evidence about the solution, not the job itself.

Layering In Forces of Progress

Clayton Christensen and Bob Moesta's Forces of Progress model — part of the broader Jobs to Be Done canon — adds a second lens that explains why stated demand doesn't convert to usage: every switch (or non-switch) is a tug-of-war between four forces.

ForceWhat it isWhich signal usually reveals it
PushDissatisfaction with the current way of doing thingsSurvey — people can describe what frustrates them
PullThe appeal of a new solutionSurvey — people can describe what sounds appealing
AnxietyFear the new thing won't work, or will be worseAnalytics — shows up as drop-off, not as a stated fear
HabitInertia of the existing routine or workaroundAnalytics — shows up as repeat use of the old path

A survey is good at capturing push and pull because those live at the level of conscious complaint and desire. It's much worse at surfacing anxiety and habit, because those live below conscious articulation — nobody says "I'm scared to try the export feature," they just quietly never click it. Analytics is the mirror image: excellent at revealing anxiety and habit through drop-off points and repeat workaround use, but blind to push and pull, since it only sees behavior, never motive.

That asymmetry is the real mechanism behind most say-do gaps: surveys overweight push/pull, analytics overweight anxiety/habit. A feature that scores high on the first pair and low on the second will always look wanted in research and unused in the dashboard.

A Practical Workflow for Reconciling the Gap

Turn the diagnosis into a repeatable process rather than a one-off investigation, so the next contradictory dashboard doesn't restart the debate from zero. The workflow below moves from raw signal to a scored, defensible opportunity in five steps.

  1. Capture the job, not the feature request. Run interviews using outcome-oriented prompts rather than "what feature do you want" — a good user interview question bank is built around past-behavior and struggling-moment questions specifically to reduce the say-do gap at the source.
  2. Synthesize consistently. Tag interview and survey verbatims against a shared taxonomy of jobs and desired outcomes so "I want CSV export" and "I need to prove compliance to my boss" get coded as evidence for the same underlying outcome, not two separate requests. This is exactly the discipline covered in guides to user research synthesis and its faster, AI-assisted synthesis variants.
  3. Score importance and satisfaction independently, using the ODI formula, so you get a ranked opportunity list before a single line of usage data enters the conversation.
  4. Cross-check against analytics using the four-gap table, classifying every high-scoring opportunity as awareness, friction, aspiration, or cohort mismatch before deciding what to build.
  5. Map the winners onto the Forces of Progress and the customer journey to find exactly where push, pull, anxiety, or habit is blocking adoption — because "build it" and "make it easier to start" are very different roadmap items even for the same opportunity.

Where This Fits in a JTBD Practice

None of this replaces a working knowledge of Jobs to Be Done theory — it's an application of it. If your team is newer to the framework, a complete guide to Jobs to Be Done covers the underlying theory of jobs, outcomes, and switching. A more practical, execution-focused guide to JTBD walks through running the interviews and workshops that feed the scoring step above.

The say-do gap is, in a real sense, JTBD's founding observation. Christensen built the whole framework around the fact that demographic surveys and stated preferences predicted the famous "milkshake" purchases far worse than understanding the job the milkshake was hired for.

Where Prodinja Fits

This is the exact workflow Prodinja's Customer Jobs workspace is built around. It turns raw interview notes into structured JTBD statements, runs Ulwick-style opportunity scoring on importance versus satisfaction, and maps the result onto Forces of Progress — so a PM can see not just that a job scores high, but which force is holding adoption back.

It's designed to make the reconciliation workflow above something you run consistently on every contradictory signal, rather than a bespoke analysis rebuilt from scratch each time a survey and a dashboard disagree.

Common Mistakes Teams Make With Contradictory Signals

Most say-do gap post-mortems go wrong in one of a few predictable ways, and each one is avoidable once you know to look for it. Recognizing the pattern in your own retro is usually enough to stop repeating it.

  • Treating the loudest survey signal as consensus. Ten enthusiastic verbatims from power users get treated as "customers want this," when they may represent 2% of the active base.
  • Declaring victory on vanity usage metrics. "80% opened the feature at least once" says nothing about whether it solved the job — pair adoption with retention of use, not first-touch.
  • Skipping the awareness check. Teams frequently conclude "nobody wants this" for a feature that was never mentioned in onboarding or release notes.
  • Averaging importance and satisfaction into one number. The ODI formula deliberately weights underserved-but-important outcomes higher than a flat average would — averaging erases exactly the signal you're trying to isolate.
  • Re-running the same survey question after a null result, instead of digging into why stated importance didn't convert, which is usually a Forces of Progress question, not a research-design one.

Key Takeaways

  • The say-do gap is not a data error — surveys measure stated intent, analytics measure revealed behavior, and the two are answering genuinely different questions.
  • Classify the gap before reacting to it: awareness, friction, aspiration, and cohort-mismatch gaps each demand a different fix, and they look identical in a raw dashboard.
  • Score the job, not the feature request — importance and satisfaction ratings, scored independently via Ulwick's Importance + max(Importance − Satisfaction, 0) formula, are far more stable than a tally of feature votes.
  • Forces of Progress explains the asymmetry: surveys are good at surfacing push and pull, analytics is good at surfacing anxiety and habit — a real gap often means you're only seeing half the forces.
  • Kahneman's remembering-self research and Nielsen's usability guidance both point the same direction: weight what people do over what they recall wanting, without discarding the "why" that only a conversation can surface.
  • Build the reconciliation into a repeatable workflow — capture the job, synthesize consistently, score independently, cross-check against usage, then map onto the journey — rather than re-litigating it every time a new contradiction appears.

Frequently Asked Questions

What is the say-do gap in user research?

The say-do gap is the well-documented mismatch between what people report wanting or intending to do in a survey or interview and what they actually do when observed through usage analytics or behavioral data. It exists because self-report is shaped by memory, identity, and social desirability, while analytics only records literal behavior.

Should I trust survey data or analytics data more?

Neither source is inherently more trustworthy — they measure different things, so the right move is to classify why they disagree rather than pick a winner. Surveys are relatively reliable for what matters to someone (importance) and weaker for predicting specific future actions; analytics is the reverse, reliable on behavior but blind to unmet or undiscovered needs.

How does Jobs to Be Done help resolve survey-analytics contradictions?

JTBD reframes the question from "what feature do people want" to "what outcome are they trying to achieve," which is a more stable, less identity-biased unit to measure. Scoring that outcome's importance and satisfaction, then mapping it onto Forces of Progress, usually reveals that a contradiction was really an awareness, friction, or anxiety problem rather than a demand problem.

What is Ulwick's opportunity scoring formula?

Tony Ulwick's Outcome-Driven Innovation formula is Opportunity Score = Importance + max(Importance − Satisfaction, 0), calculated from survey ratings of how important and how satisfied respondents are with a specific, measurable outcome. High-importance, low-satisfaction outcomes score highest, flagging them as underserved opportunities worth investigating further against usage data.

Can analytics alone tell me if a feature request is worth building?

No — analytics can tell you whether a shipped, discoverable feature is being used, but it can't tell you why an unbuilt or undiscovered one might matter, or which force (push, pull, anxiety, habit) is blocking adoption. Pairing usage data with outcome-scored qualitative research is what turns a raw contradiction into a defensible roadmap decision.