Visual guide to usability testing tasks, moderation, evidence, severity, and retesting

How to Run an Effective Usability Test

Author: JVDS Design Studio Reading time: about 8 min

Asking users to look at a page and say “It seems good” is not usability testing. Testing observes whether people can find the entry, understand information, complete actions, and recover from errors in tasks close to real use.

Asking users to look at a page and say “It seems good” is not usability testing. Testing observes whether people can find the entry, understand information, complete actions, and recover from errors in tasks close to real use.

01 Decide Which Behaviors to Observe

A broad goal—“See whether the new version is easy to use”—makes moderators ask everything and finish with opinions such as “The button could be larger” or “The colors look good.”

Turn the goal into observable behavior:

  • Can users find the creation entry without help?
  • Can users choose the right option from the available information?
  • Can users understand an error and correct it?
  • How many returns, failed attempts, or requests for help occur?
  • Do different roles interpret the same term differently?

Write a risk list before testing. Which design failures could block submission, cause incorrect actions, create business loss, or generate development rework? Tasks should prioritize those risks instead of browsing every page evenly.

02 Preparation: Five Essentials for Every Test

Objectives and Success Signals

Define “complete,” “partially complete,” and “failed” for each task. Signals may include outcome, independent completion, critical errors, time, and confidence, but do not quantify everything merely to build a dashboard.

Participant Criteria

Recruit actual users or people highly likely to use the product. Role, expertise, frequency, and context often matter more than a raw count. For a focused iterative test, begin with a small round and add participants if new issues continue to appear.

A Testable Prototype or Product

High fidelity is not always required. Wireframes or clickable flows are enough for information architecture; visual hierarchy, trust, or critical feedback may require a near-final interface. Tell observers which features are simulated, but do not reveal paths in the task itself.

A Moderation Script

A script keeps sessions consistent; it is not a text to read mechanically. Include introduction, informed consent, warm-up, tasks, follow-up questions, closing, and thanks.

An Observation Sheet

Use columns such as “Time—Behavior—Direct Quote—Observer Interpretation—Question to Confirm.” Keeping facts separate from interpretation prevents later disputes.

Visual explanation that a good task is a scenario rather than an instruction

03 Write Scenarios, Not Interface Instructions

The following pairs look similar, but their testing value differs completely.

Weak TaskBetter Task
Click “New” in the upper right and create a projectYou have a new client and need to create a project and invite a teammate to follow it; prepare everything in the system
Select “Pending Review” in the filterYour manager asks you to find applications still unreviewed this week and process one
Change the address and click SaveYou notice the shipping address is wrong; correct it before dispatch

A good task is clear, relevant, and challenging without naming interface labels or the path. Use business language participants understand rather than the product team’s module names.

04 Explain Less and Trace More

The moderator does not teach the user; they make the thought process visible. Explain that “We are testing the product, not you,” and invite participants to describe what they seek and why they hesitate.

When a user stops, use neutral prompts:

  • “What are you looking for now?”
  • “What does this suggest to you?”
  • “What would you do next if I were not here?”
  • “What do you expect to happen after clicking?”

Do not ask “Is this button hard to see?” or “Would it be better on the left?” Such questions put the answer in the user’s mouth.

A Reusable Moderator Introduction

Visual explanation of recording more than task completion

05 Observe More Than “Completed” or “Failed”

Record at least six types of evidence for a task:

EvidenceExample
PathEntry point, detours, and backtracking
UnderstandingInterpretation of labels, states, fees, permissions
BehaviorClicks, search, comparison, repeated review, skipping
ErrorWrong selection, omission, duplicate submission, invalid attempt
RecoveryWhether the user notices and can correct the error
Emotional signalsHesitation, confirmation, concern, surprise—without diagnosing psychology

Observers should not debate design solutions during the session. Record issues and ideas separately so a strong opinion does not bias later observation.

06 Grade Issues by Impact, Frequency, Recovery, and Risk

Counting occurrences alone undervalues rare but high-risk problems. JVDS generally separates severity into four judgments:

1. Task impact: Does it block the core task or add only one step?

2. Scope: Does it affect one user, one role, or most users?

3. Recovery difficulty: Can users recognize and fix it themselves?

4. Business risk: Could it cause incorrect approval, payment, privacy exposure, or an irreversible action?

LevelTypical BehaviorRecommended Response
P0 BlockerCore task cannot be completed or severe security or business risk existsStop launch, fix, and retest
P1 CriticalMost target users fail or recovery is very difficultMust fix in the current release
P2 SignificantTask completes, but understanding is costly, efficiency is poor, or help is requiredSchedule soon and validate the solution
P3 MinorOutcome is unaffected; issue concerns local copy or consistencyCombine with other improvements

This is a project-management method, not a universal industry standard. Teams can adjust levels, but judgment rules must remain consistent and evidence must be retained.

Visual explanation of synthesizing findings on the day of testing

07 Synthesize Quickly on the Same Day

Do not wait a week after every session before reviewing notes. Spend 15 minutes after each session answering:

  • What new issue appeared today?
  • Which issues repeated?
  • What should the next session validate—not immediately change?

After all sessions, cluster related evidence into “Issue—Evidence—Impact—Cause Hypothesis—Recommendation—Validation.” Label whether a cause is inferred or verified so the design team’s explanation is not mistaken for a user fact.

08 Example: A Data Import Flow

A team tests an Excel import feature. The first task says only “Upload the file and complete import.” Users click the button successfully, and the team concludes the flow is fine.

In the redesigned task, participants must download a template, enter two invalid records, upload it, understand validation results, and repair failed rows. Three real issues emerge: users cannot distinguish “file uploaded” from “data imported”; errors do not identify the cell; and every correction requires reuploading the entire file.

The lesson is that tasks must cover the complete work loop, not a page. Otherwise, testing proves only that a button can be clicked.

09 Retesting Matters More Than “Changed”

A solution passing design review does not prove the problem is gone. Retest P0 and P1 issues when possible, especially after changing workflows, terminology, or information architecture. Retesting can be narrower and recruit only relevant roles, but tasks should remain comparable and new issues should be recorded.

Frequently Asked Questions

How many people does one usability-test round need?

No fixed number fits every project. A focused flow with one consistent user type can begin with a small round; major differences in role, device, or context require groups. Continue according to whether critical issues are still emerging and essential user groups are covered.

Can a moderator answer user questions?

It depends on the objective. To test independent support, delay help and record assistance points. To test work after training, provide the agreed support available in the real context. Keep the approach consistent across sessions.

Should testing be remote or in person?

Remote testing suits distributed users, digital products, and rapid iteration. In-person testing reveals devices, paper materials, collaboration, and the work environment. Choose according to real use, not tool convenience.

Does testing require a long report?

No. Iterative work needs a clear issue list, evidence clips, and decision record. Long reports support cross-team learning but cannot replace a review where product, design, and engineering see the evidence together.

10 Put Research and Design Into Product Decisions

ServiceView
UI/UX design servicesView service details
Project inquiryContact JVDS
Design and website articlesRead more related articles
Link copied

From Idea to Launch, We Build It Together

Building useful, scalable digital products around user experience

Tell Us About Your Project