top of page
uintent company logo
Contact

UX QUALITY, UX METHODS, UX METRICS

Interpreting UX Metrics Correctly: Context Over Gut Feeling

7

MIN

Sep 3, 2026

A project manager recently asked me whether two user groups differed significantly. Each group consisted of 15 people. Short answer: a UX metric is only as good as its context, its sample and its methodology. A figure without these three elements is just noise, not insight. Where exactly was the flaw in his question, and what actually makes a metric meaningful?


📌 Key takeaways

  • A metric is only useful if it is measurable, systematic and linked to a clear objective.

  • A figure without temporal, comparative or goal-related context tells us nothing.

  • With small samples such as n=15, ‘significant’ is often the wrong question to ask. The confidence interval is what matters.

  • Not every metric is a KPI. But every KPI is a metric.

  • Vanity metrics such as page views look good, but they do not influence decision-making.

  • SUS, NPS, CSAT and CES have established benchmarks, but also blind spots.

  • Correlation is not proof of cause and effect, even if two curves run nicely parallel to each other.


What exactly is a UX metric?

A metric is a measurable indicator that can be used to assess progress towards or the status of a goal. In the context of UX, this means: quantifiable data on the quality of the user experience, collected systematically and in a repeatable manner.

Three characteristics distinguish a genuine metric from gut feeling with decimal places. It is measurable – a number rather than an impression. It is systematic – repeatable and comparable over time. And it is meaningful – linked to a genuine objective rather than marketing jargon.

It sounds trivial, but it isn’t. Anyone who formulates ‘improving the user experience’ as a goal does not yet have a metric. Anyone who turns that into ‘increasing the task success rate from 65 per cent to 85 per cent’ does have one. This step of translating an abstract concept into a measurable quantity is called operationalisation. In practice, it is precisely this step that is most often skipped, usually due to time constraints.


Metric or KPI – what’s the difference?

Not every metric is a KPI, but every KPI is a metric. Bounce rate, page views or click-through rate are metrics: operational, broad in scope, useful for context. A KPI, on the other hand, is one of the few metrics that is strategic enough to inform a decision.

A rule of thumb from real-world experience: 3 to 5 KPIs per team are enough; any more is overwhelming. How you select these few KPIs and embed them within the organisation is a separate, broader topic in its own right and will be covered in a separate article. For now, let’s focus on the metric itself.


Why a number without context is worthless

“Our task success rate is 78 per cent.” Sounds like a statement. But it isn’t. Good or bad? Impossible to say without a comparison. Only with context does the figure become information: “Task success rate 78 per cent, last quarter 65 per cent, industry benchmark 72 per cent.” Now everyone in the room knows that it’s an improvement and that it’s above the industry average.

Three types of context turn a figure into a basis for decision-making. Temporal context shows the trend – whether it’s rising or falling compared to the last quarter. Comparative context places the figure alongside a benchmark or the competition. Target context indicates how far one is from the target.

If any one of these three points of reference is missing, the figure remains a fact with no informative value, and no one in the team can make meaningful use of it.


Can we speak of a significant difference with 15 test subjects?

That is precisely the question I was asked recently. Two groups, each with 15 people, had different task success rates. The project manager wanted to know whether the difference was significant.

The honest answer: with such small sample sizes, ‘significant’ is almost never the right question to ask. What’s more important is the confidence interval – that is, the range within which the true value is likely to lie.

A calculation example makes this clearer. With dichotomous data such as task success (successful or not), 9 out of 12 successful attempts result in a success rate of 75 per cent. That sounds precise. However, the 95 per cent confidence interval, calculated using the Wilson score interval (Wilson, 1927), lies between 43 per cent and 95 per cent. A range of over 50 percentage points. With metric data such as a SUS score, the situation looks slightly better: a mean of 75 points across 12 people has a confidence interval of approximately 66 to 84 points. Significantly narrower, but still broad enough to warrant caution.

For my project manager, this meant that whilst the figures were interesting, they did not provide a basis for a firm decision. We increased the sample size for the next round of testing and reframed the question from ‘is this significant?’ to ‘how certain can we be?’. A small difference in wording, but a big difference in the quality of the decision.


How can you spot a vanity metric?

Pageviews look good in any report. But they say nothing about whether someone has solved their problem. That’s the definition of a vanity metric: it looks convincing without actually informing a decision. Total registrations without activation, time on site without context (long because they’re enthusiastic, or long because they’re lost?), social media likes with no link to conversion. Figures that shine but still reveal very little.

The test is simple: if this figure changes, what do you do differently? With a task success rate, the answer is clear – you revise the flow. With page views? Probably nothing. That’s the difference between an action metric and a vanity metric.


How useful are SUS, NPS, CSAT and CES really?

These four acronyms appear in almost every report, and for good reason – they have established benchmarks. Nevertheless, they should be treated with caution.


Metric

What it measures

Strength

Limit

SUS (System Usability Scale)

Overall impression of usability, 10 questions, score 0 to 100

Comparable across different products

Does not provide any insight into specific individual issues

NPS (Net Promoter Score)

Willingness to recommend, one question, scale 0 to 10

Well suited to tracking over time

Its link to business success is disputed

CSAT (Customer Satisfaction)

Satisfaction immediately following an interaction

Quick feedback on individual touchpoints

Highly context-dependent, difficult to compare

CES (Customer Effort Score)

How much effort a task required

Highlights friction points in the process

Less meaningful for measuring pure satisfaction

 

None of these four metrics replaces the others. NPS shows loyalty over time; CES shows where the process is faltering. If you only track one of them, you’re looking at a snapshot, not the full picture.


Why correlation does not imply causality

Ice-cream sales correlate with the number of drowning victims. If one figure rises, so does the other. Nevertheless, nobody seriously sells less ice cream in order to save lives. The reason: a third factor – warm summer weather – drives both figures up at the same time.

The same pattern is constantly lurking in UX data. Two metrics move in parallel, and the temptation is great to construct a cause-and-effect chain from them. Three steps can protect against this. First, formulate a hypothesis: what cause-and-effect relationship is being suggested? Then explain the mechanism: why should one thing trigger the other? And finally, validate it, for example through an A/B test or a time-series analysis.

Without this final step, any narrative based on metrics remains merely an assertion. With it, it becomes robust.


How do you ensure that your data contributes to the quality of user research?

Three criteria determine whether a survey is trustworthy at all.


  1. Validity: Are you really measuring what you set out to measure, or just something that happens to correlate with it?

  2. Reliability: If the study were repeated under the same conditions, would it produce the same result?

  3. Representativeness: Does your sample reflect the actual user base, or just those who happened to have time for the test?


If you consistently check these three criteria before launching your next survey, you’ll measure less often, but more accurately. That’s the difference between a metric that deserves trust and one that merely looks convincing.


Frequently asked questions about UX metrics

Is a difference between two groups with n=15 statistically significant? Rarely clear-cut. With such small samples, the confidence interval is usually wide – often exceeding 20 to 30 percentage points for dichotomous data. Rather than asking whether it is ‘significant’, it is worth looking at the range within which the true value is likely to lie.


What is the difference between a UX metric and a KPI?

Every KPI is a metric, but not every metric is a KPI. Metrics are broad and operational, whilst KPIs are the few strategic indicators that underpin a decision. Further details on this will follow in a separate article.


Which UX metric should I introduce first?

That depends on the objective. The Task Success Rate is suitable for usability issues, whilst NPS or SUS are suitable for measuring satisfaction over time. More important than the choice of metric is the question of whether a change in the figure would actually lead to any change in your actions.


How do I recognise a vanity metric?

If, when the figure changes, you cannot imagine what specifically you would do differently, it is probably a vanity metric.


Is a correlation enough to justify a UX improvement?

Not on its own. A correlation can be an initial indication, but it must first be validated – for example, via an A/B test – before it can serve as a justification for an investment.


Conclusion

UX metrics rarely fail because the wrong figure has been chosen. They fail because the context is missing, the sample size is too small, or a correlation is mistaken for causation. All three mistakes can be avoided, without needing a degree in statistics.

Anyone introducing a new metric should first ask themselves: what will I do differently if this figure changes? If there’s no answer, it’s the wrong metric.

On 13 November, in the workshop ‘UX Metrics: Defining, Measuring and Managing’, I’ll show you how your team can progress from the first metric to robust reporting.

Not enough yet? Then read on in our newsletter. It comes out four times a year. It sticks in your mind for longer.


About the author

Tara Bosenick is a UX consultant and co-owner of Uintent. Since 1999, she has been helping companies make their products more user-friendly, using sound research methods and a clear eye for what really matters. As a speaker at conferences such as Mensch & Computer and the World Usability Congress, she shares her knowledge of UX and AI. Her workshops on UX-AI prompting and AI integration embody what makes for good UX: clear benefits, direct applicability and enjoyment of the process.

Subscribe to our newsletter

Photo of a desk with charts and notes, overlaid with hand-drawn arrows and circles indicating mean, median, and significance

Quantitative UX Methods: How to Read and Interpret Numbers Correctly

UX METHODS, UX METRICS, BEST PRACTICES

Top-down view of a desk with a centrally placed pink popsicle, rising line chart, and hand-drawn orange curves, arrows, and question marks.

Interpreting UX Metrics Correctly: Context Over Gut Feeling

UX QUALITY, UX METHODS, UX METRICS

Glowing abstract profile cards connected by cyan data lines, with repeated golden silhouettes in a dark digital space.

AI Personas: What They Can Do, What They Shouldn’t Do

AI & UX Research

Barcamp session board covered in Post-it notes arranged in a grid, with hand-drawn arrows and annotations highlighting connections between sessions

Barcamp Guide: How to Pitch Sessions and Organize Your Own Barcamp

TRENDS, BEST PRACTICES

A photographic desk scene with a research report, charts, sticky notes, coffee, and a pen. Loose dark-navy hand-drawn circles, arrows, question marks, an X, and a checkmark create the feeling of a critical research review.

UX Research Quality: Why Good Intentions Aren’t Enough

AI & UX Research

Symbolic digital illustration: A glowing prompt cursor suspended at the center of a dark space, connected to a sparse network of luminous nodes. Some points shine brightly, others fade – a visual metaphor for deliberate, intentional AI use.

Sustainable Prompting: Inspiration for UX Teams

AI & UX Research

Futuristic illustration of three floating AI tools: a glowing spark, a transparent workspace cube with layered documents, and a crystalline gear, connected by golden lines against a deep navy background.

Prompt, Project, or Skill? Which AI Tool Truly Accelerates Your UX Research

AI & UX Research

Glowing futuristic shield made of UI elements repels digital threats in dark space.

UX Research As Risk Management: Why We Finally Need To Change Our Language

BEST PRACTICES, UX QUALITY

Person at desk between chaotic and structured data streams, central light focus

UX & AI: The Best Newsletters and Podcasts – My Personal Selection

AI & UX Research

Futuristic digital illustration: A glowing golden certification seal floating against a deep navy background, surrounded by AR interface fragments and a faint headset silhouette – symbolizing trust and validation in medical technology.

Trust, but Verified: Why Medical Certification Matters for AR, VR, and Mr in Medtech

MEDICAL, UX METHODS

Floating semi-transparent AR interface with minimal medical data and anatomical visuals, glowing in cyan and gold against a dark futuristic background.

Making the Magic Usable: Why Usability Engineering Matters for AR, VR, and MR in Medtech

MEDICAL

A futuristic, symbolic illustration shows a person standing on a glowing bridge between two worlds: on the left, a warmly lit hospital room with a bed and medical equipment; on the right, an immersive digital space featuring a holographic human body with organs glowing in cyan and orange tones. Both sides are connected by flowing streams of light, set against a deep navy blue background with soft violet transitions.

Reality, Reimagined: How AR, VR, and Mr Are Finding Their Way Into Medtech

TRENDS, MEDICAL

A glowing golden trophy floats above a gap, while small figures below work on user research and wireframes, untouched by its light.

Understanding UX AI Benchmarks: What HLE and METR Really Tell Us About AI Tools

AI & UX Research

Futuristic digital illustration on a deep navy background: a human hand holding a warm glowing pencil and a cyan-lit robotic hand both reach toward a radiant central data cluster. Surrounded by stacked documents and a network of connected nodes, the scene symbolizes collaboration between human interpretation and digital information processing.

NotebookLM in UX Research: An Honest Assessment of a Specialized AI Tool

AI & UX Research, BEST PRACTICES

Futuristic glowing cylinder divided into segments by golden barriers.

Introducing Gated Salami Prompting: Why You Should Slice Complex LLM Tasks Into Smaller Pieces

AI & UX Research, BEST PRACTICES

Futuristic square illustration on deep navy background: a glowing golden speech bubble dissolves into particles that partially reassemble incorrectly, surrounded by energy arcs, luminous nodes, and a stylized digital head—symbolizing LLM hallucinations.

Fictitious Quotes, Lost Nuances: The Hallucination Problem in Qualitative Analysis With Llms

AI & UX Research, BEST PRACTICES, UX METHODS

Surreal futuristic illustration of a glowing digital head with data streams, charts, and evaluation symbols representing AI evaluation methodology.

How do we know that our prompt is doing a good job? Why UX research needs an evaluation methodology for AI-based analysis

AI & UX Research, TRENDS, BEST PRACTICES

A surreal, futuristic illustration featuring a translucent human profile with a glowing brain connected by flowing data streams to a hovering, golden crystal.

Prompt Psychology Exposed: Why “Tipping” ChatGPT Sometimes Works

AI & UX Research, BEST PRACTICES

Surreal, futuristic illustration of a person seen from behind standing in a glowing digital cityscape.

System Prompts in UX Research: What You Need to Know About Invisible AI Control

AI & UX Research, UX QUALITY

Abstract futuristic illustration of a person, various videos, and notes.

Summarizing YouTube Videos With AI: Three Tools Put to the Test in UX Research

AI & UX Research, BEST PRACTICES

Related Articles you might enjoy

AUTHOR

Tara Bosenick

Tara has been active as a UX specialist since 1999 and has helped to establish and shape the industry in Germany on the agency side. She specialises in the development of new UX methods, the quantification of UX and the introduction of UX in companies.


At the same time, she has always been interested in developing a corporate culture in her companies that is as ‘cool’ as possible, in which fun, performance, team spirit and customer success are interlinked. She has therefore been supporting managers and companies on the path to more New Work / agility and a better employee experience for several years.


She is one of the leading voices in the UX, CX and Employee Experience industry.

bottom of page