
UX QUALITY, UX METHODS, UX METRICS
Interpreting UX Metrics Correctly: Context Over Gut Feeling
7
MIN
Sep 3, 2026
A project manager recently asked me whether two user groups differed significantly. Each group consisted of 15 people. Short answer: a UX metric is only as good as its context, its sample and its methodology. A figure without these three elements is just noise, not insight. Where exactly was the flaw in his question, and what actually makes a metric meaningful?
📌 Key takeaways
A metric is only useful if it is measurable, systematic and linked to a clear objective.
A figure without temporal, comparative or goal-related context tells us nothing.
With small samples such as n=15, ‘significant’ is often the wrong question to ask. The confidence interval is what matters.
Not every metric is a KPI. But every KPI is a metric.
Vanity metrics such as page views look good, but they do not influence decision-making.
SUS, NPS, CSAT and CES have established benchmarks, but also blind spots.
Correlation is not proof of cause and effect, even if two curves run nicely parallel to each other.
What exactly is a UX metric?
A metric is a measurable indicator that can be used to assess progress towards or the status of a goal. In the context of UX, this means: quantifiable data on the quality of the user experience, collected systematically and in a repeatable manner.
Three characteristics distinguish a genuine metric from gut feeling with decimal places. It is measurable – a number rather than an impression. It is systematic – repeatable and comparable over time. And it is meaningful – linked to a genuine objective rather than marketing jargon.
It sounds trivial, but it isn’t. Anyone who formulates ‘improving the user experience’ as a goal does not yet have a metric. Anyone who turns that into ‘increasing the task success rate from 65 per cent to 85 per cent’ does have one. This step of translating an abstract concept into a measurable quantity is called operationalisation. In practice, it is precisely this step that is most often skipped, usually due to time constraints.
Metric or KPI – what’s the difference?
Not every metric is a KPI, but every KPI is a metric. Bounce rate, page views or click-through rate are metrics: operational, broad in scope, useful for context. A KPI, on the other hand, is one of the few metrics that is strategic enough to inform a decision.
A rule of thumb from real-world experience: 3 to 5 KPIs per team are enough; any more is overwhelming. How you select these few KPIs and embed them within the organisation is a separate, broader topic in its own right and will be covered in a separate article. For now, let’s focus on the metric itself.
Why a number without context is worthless
“Our task success rate is 78 per cent.” Sounds like a statement. But it isn’t. Good or bad? Impossible to say without a comparison. Only with context does the figure become information: “Task success rate 78 per cent, last quarter 65 per cent, industry benchmark 72 per cent.” Now everyone in the room knows that it’s an improvement and that it’s above the industry average.
Three types of context turn a figure into a basis for decision-making. Temporal context shows the trend – whether it’s rising or falling compared to the last quarter. Comparative context places the figure alongside a benchmark or the competition. Target context indicates how far one is from the target.
If any one of these three points of reference is missing, the figure remains a fact with no informative value, and no one in the team can make meaningful use of it.
Can we speak of a significant difference with 15 test subjects?
That is precisely the question I was asked recently. Two groups, each with 15 people, had different task success rates. The project manager wanted to know whether the difference was significant.
The honest answer: with such small sample sizes, ‘significant’ is almost never the right question to ask. What’s more important is the confidence interval – that is, the range within which the true value is likely to lie.
A calculation example makes this clearer. With dichotomous data such as task success (successful or not), 9 out of 12 successful attempts result in a success rate of 75 per cent. That sounds precise. However, the 95 per cent confidence interval, calculated using the Wilson score interval (Wilson, 1927), lies between 43 per cent and 95 per cent. A range of over 50 percentage points. With metric data such as a SUS score, the situation looks slightly better: a mean of 75 points across 12 people has a confidence interval of approximately 66 to 84 points. Significantly narrower, but still broad enough to warrant caution.
For my project manager, this meant that whilst the figures were interesting, they did not provide a basis for a firm decision. We increased the sample size for the next round of testing and reframed the question from ‘is this significant?’ to ‘how certain can we be?’. A small difference in wording, but a big difference in the quality of the decision.
How can you spot a vanity metric?
Pageviews look good in any report. But they say nothing about whether someone has solved their problem. That’s the definition of a vanity metric: it looks convincing without actually informing a decision. Total registrations without activation, time on site without context (long because they’re enthusiastic, or long because they’re lost?), social media likes with no link to conversion. Figures that shine but still reveal very little.
The test is simple: if this figure changes, what do you do differently? With a task success rate, the answer is clear – you revise the flow. With page views? Probably nothing. That’s the difference between an action metric and a vanity metric.
How useful are SUS, NPS, CSAT and CES really?
These four acronyms appear in almost every report, and for good reason – they have established benchmarks. Nevertheless, they should be treated with caution.
Metric | What it measures | Strength | Limit |
SUS (System Usability Scale) | Overall impression of usability, 10 questions, score 0 to 100 | Comparable across different products | Does not provide any insight into specific individual issues |
NPS (Net Promoter Score) | Willingness to recommend, one question, scale 0 to 10 | Well suited to tracking over time | Its link to business success is disputed |
CSAT (Customer Satisfaction) | Satisfaction immediately following an interaction | Quick feedback on individual touchpoints | Highly context-dependent, difficult to compare |
CES (Customer Effort Score) | How much effort a task required | Highlights friction points in the process | Less meaningful for measuring pure satisfaction |
None of these four metrics replaces the others. NPS shows loyalty over time; CES shows where the process is faltering. If you only track one of them, you’re looking at a snapshot, not the full picture.
Why correlation does not imply causality
Ice-cream sales correlate with the number of drowning victims. If one figure rises, so does the other. Nevertheless, nobody seriously sells less ice cream in order to save lives. The reason: a third factor – warm summer weather – drives both figures up at the same time.
The same pattern is constantly lurking in UX data. Two metrics move in parallel, and the temptation is great to construct a cause-and-effect chain from them. Three steps can protect against this. First, formulate a hypothesis: what cause-and-effect relationship is being suggested? Then explain the mechanism: why should one thing trigger the other? And finally, validate it, for example through an A/B test or a time-series analysis.
Without this final step, any narrative based on metrics remains merely an assertion. With it, it becomes robust.
How do you ensure that your data contributes to the quality of user research?
Three criteria determine whether a survey is trustworthy at all.
Validity: Are you really measuring what you set out to measure, or just something that happens to correlate with it?
Reliability: If the study were repeated under the same conditions, would it produce the same result?
Representativeness: Does your sample reflect the actual user base, or just those who happened to have time for the test?
If you consistently check these three criteria before launching your next survey, you’ll measure less often, but more accurately. That’s the difference between a metric that deserves trust and one that merely looks convincing.
Frequently asked questions about UX metrics
Is a difference between two groups with n=15 statistically significant? Rarely clear-cut. With such small samples, the confidence interval is usually wide – often exceeding 20 to 30 percentage points for dichotomous data. Rather than asking whether it is ‘significant’, it is worth looking at the range within which the true value is likely to lie.
What is the difference between a UX metric and a KPI?
Every KPI is a metric, but not every metric is a KPI. Metrics are broad and operational, whilst KPIs are the few strategic indicators that underpin a decision. Further details on this will follow in a separate article.
Which UX metric should I introduce first?
That depends on the objective. The Task Success Rate is suitable for usability issues, whilst NPS or SUS are suitable for measuring satisfaction over time. More important than the choice of metric is the question of whether a change in the figure would actually lead to any change in your actions.
How do I recognise a vanity metric?
If, when the figure changes, you cannot imagine what specifically you would do differently, it is probably a vanity metric.
Is a correlation enough to justify a UX improvement?
Not on its own. A correlation can be an initial indication, but it must first be validated – for example, via an A/B test – before it can serve as a justification for an investment.
Conclusion
UX metrics rarely fail because the wrong figure has been chosen. They fail because the context is missing, the sample size is too small, or a correlation is mistaken for causation. All three mistakes can be avoided, without needing a degree in statistics.
Anyone introducing a new metric should first ask themselves: what will I do differently if this figure changes? If there’s no answer, it’s the wrong metric.
On 13 November, in the workshop ‘UX Metrics: Defining, Measuring and Managing’, I’ll show you how your team can progress from the first metric to robust reporting.
Not enough yet? Then read on in our newsletter. It comes out four times a year. It sticks in your mind for longer.
About the author
Tara Bosenick is a UX consultant and co-owner of Uintent. Since 1999, she has been helping companies make their products more user-friendly, using sound research methods and a clear eye for what really matters. As a speaker at conferences such as Mensch & Computer and the World Usability Congress, she shares her knowledge of UX and AI. Her workshops on UX-AI prompting and AI integration embody what makes for good UX: clear benefits, direct applicability and enjoyment of the process.
Related Articles you might enjoy
AUTHOR
Tara Bosenick
Tara has been active as a UX specialist since 1999 and has helped to establish and shape the industry in Germany on the agency side. She specialises in the development of new UX methods, the quantification of UX and the introduction of UX in companies.
At the same time, she has always been interested in developing a corporate culture in her companies that is as ‘cool’ as possible, in which fun, performance, team spirit and customer success are interlinked. She has therefore been supporting managers and companies on the path to more New Work / agility and a better employee experience for several years.
She is one of the leading voices in the UX, CX and Employee Experience industry.




















