Should AI Products Show Confidence Scores? Designing for Evidence and Uncertainty — featured visual

Should AI Products Show Confidence Scores? Designing for Evidence and Uncertainty

Author: JVDS Design Studio Reading time: about 8 min

"82% confidence" may look scientific, but without calibration and a clear explanation of what the number means, it can create false certainty. Better design usually starts with evidence, scope and verifiability.

Many AI products add a percentage next to the answer to let users "know how certain the model is": Confidence 87%, Confidence 0.82. The interface looks more professional, but such numbers are very easy to be misunderstood. The Explainability + Trust guide of Google PAIR clearly reminds us that numerical confidence levels may be misunderstood by users and need to be verified in real scenarios. Explanations should also prioritize helping users complete tasks rather than aiming to expose the entire interior of the model.

01 Start with the user's decision, then decide whether to show a score

If a confidence number cannot help users take different actions, it is just decoration. For most generative AI products, sources, scope of application, conflict evidence, and "I don't know" are often more useful than an isolated percentage.

02 Why 82% can be dangerous

Users might interpret it as "this answer has an 82% probability of being true", but the actual score could come from classifier probability, retrieval similarity, model self-evaluation, or team-defined rules, with completely different meanings. If the system is not calibrated, precise numbers to single digits actually create non-existent precision.

For generative answers, show evidence quality first — visual illustration

03 For generative answers, show evidence quality first

For example, the answer can indicate: based on three currently valid systems; Last updated on August 2026; Two of them have the same source, while one has a conflict. Users can directly open the supporting paragraph. Such information is closer to the real judgment process and is also easier to verify.

04 Turn uncertainty into a next action

The best uncertainty prompt is not a simple "The result is for reference only", but rather telling users why they are unsure and what to do. For example: "Two different versions of the policy were found. It is impossible to determine the currently applicable version. Please select your region." Or, "Due to the lack of data for the second quarter of 2026, we can only analyze up to June at present."

05 When confidence scores are useful

In tasks such as classification, recognition, and prediction, if the score is verified and the user knows the meaning of the threshold, it can assist in decision-making. For instance, auditors might only manually review results that fall below a certain threshold. However, the interface should explain what the score represents and provide the calibrated action rules, rather than allowing users to guess for themselves.

Layer explanations instead of exposing all model detail at once — visual illustration

06 Layer explanations instead of exposing all model detail at once

Ordinary users usually first need a brief reason: what data was used and what the key basis was. Professional users then expand features, retrieve fragments, model versions or more technical information. Google PAIR advocates interpretation for understanding and controls complexity through progressive disclosure.

07 High confidence must not replace human judgment

In high-risk areas, the interface should retain key evidence, responsible persons and confirmation actions. Even if the model score is very high, a green logo cannot imply that "it can be executed automatically with confidence". Automation strategies should integrate risk, reversibility and authority, rather than merely relying on model confidence.

08 A four-layer framework for communicating trustworthiness

  • The first layer: Clearly define the scope of application of the answer and the data time.
  • The second layer: Display direct sources and evidence fragments.
  • The third layer: Hint at conflicts, omissions and reasons for uncertainty.
  • The fourth layer: The confidence value will only be displayed when it has been calibrated and is for clear decision-making purposes.

09 The purpose of explanation is not to make AI look scientific

Good Explainability UX helps users understand "why I can trust to what extent and what to verify next". If a 92% of users will only confirm more boldly after watching it, then this figure may increase the risk. Only when users can make correct judgments more quickly after reviewing the evidence can explanations truly create value.

Four UI patterns more useful than a confidence percentage — visual illustration

10 Four UI patterns more useful than a confidence percentage

  • Evidence coverage: Show that the answer is supported by several valid sources and whether there are any omissions.
  • Conflict warning: Clearly mark that there are different conclusions among the sources, rather than averaging them into one answer.
  • Scope label: Indicates that the answer is only applicable to a certain region, time, customer type or dataset.
  • Verification suggestion: Inform the user which original text to view next, what data to supplement, or who to confirm.

The common point of this information is that it is actionable. Users often don't know what to do when they see "87%". If you see the message "The latest policy does not cover overseas branches. Please confirm the local HR rules", you can take the next step immediately.

11 If you must show a score, get three things right

First, clearly define: Is this classification probability, retrieval relevance, or business risk score? Second, conduct calibration tests to ensure that 80% of the data is indeed explainable. Thirdly, map the scores to operational rules. For instance, if the score is below the threshold, it must be manually rechecked instead of merely changing the color of the logo. User research also needs to verify whether different roles overly rely on high scores.

Frequently Asked Questions

Is it true that the more accurate the AI confidence level is, the better?

No. Uncalibrated precise numbers may create a false sense of certainty. The definition of the score and its decision-making purpose should be explained first.

Is generative AI suitable for displaying probabilities?

Most open-ended generative tasks find it difficult to compress the "correct answer rate" into a single probability. Sources, evidence coverage, conflicts and missing information are usually more useful.

How to design the "AI uncertainty" state?

Explain the reasons for the uncertainty, the current known range, the missing conditions, and provide clarification, supplementary data, source viewing or the next step of transferring to a human.

Is it true that the more detailed the explanation, the more reliable it is?

Not necessarily. Too many technical details will increase the cognitive burden. Hierarchical/progressive disclosure should be adopted to allow different roles to unfold as needed.

Should the confidence level of high-risk AI be hidden?

It doesn't have to be hidden, but the score cannot be used as a substitute for automatic approval. It should be designed in combination with evidence, risk, manual review and reversibility.

Related ServiceLearn More
UI/UX Design ServicesView Service Details
Project ConsultationContact JVDS Design Studio
Design and Website Development ArticlesRead More Related Articles
Link copied

From Idea to Launch, We Build It Together

Building useful, scalable digital products around user experience

Tell Us About Your Project