Asking users to look at a page and say “It seems good” is not usability testing. Testing observes whether people can find the entry, understand information, complete actions, and recover from errors in tasks close to real use.
Asking users to look at a page and say “It seems good” is not usability testing. Testing observes whether people can find the entry, understand information, complete actions, and recover from errors in tasks close to real use.
01 Decide Which Behaviors to Observe
A broad goal—“See whether the new version is easy to use”—makes moderators ask everything and finish with opinions such as “The button could be larger” or “The colors look good.”
Turn the goal into observable behavior:
- Can users find the creation entry without help?
- Can users choose the right option from the available information?
- Can users understand an error and correct it?
- How many returns, failed attempts, or requests for help occur?
- Do different roles interpret the same term differently?
Write a risk list before testing. Which design failures could block submission, cause incorrect actions, create business loss, or generate development rework? Tasks should prioritize those risks instead of browsing every page evenly.
02 Preparation: Five Essentials for Every Test
Objectives and Success Signals
Define “complete,” “partially complete,” and “failed” for each task. Signals may include outcome, independent completion, critical errors, time, and confidence, but do not quantify everything merely to build a dashboard.
Participant Criteria
Recruit actual users or people highly likely to use the product. Role, expertise, frequency, and context often matter more than a raw count. For a focused iterative test, begin with a small round and add participants if new issues continue to appear.
A Testable Prototype or Product
High fidelity is not always required. Wireframes or clickable flows are enough for information architecture; visual hierarchy, trust, or critical feedback may require a near-final interface. Tell observers which features are simulated, but do not reveal paths in the task itself.
A Moderation Script
A script keeps sessions consistent; it is not a text to read mechanically. Include introduction, informed consent, warm-up, tasks, follow-up questions, closing, and thanks.
An Observation Sheet
Use columns such as “Time—Behavior—Direct Quote—Observer Interpretation—Question to Confirm.” Keeping facts separate from interpretation prevents later disputes.

03 Write Scenarios, Not Interface Instructions
The following pairs look similar, but their testing value differs completely.
| Weak Task | Better Task |
|---|---|
| Click “New” in the upper right and create a project | You have a new client and need to create a project and invite a teammate to follow it; prepare everything in the system |
| Select “Pending Review” in the filter | Your manager asks you to find applications still unreviewed this week and process one |
| Change the address and click Save | You notice the shipping address is wrong; correct it before dispatch |
A good task is clear, relevant, and challenging without naming interface labels or the path. Use business language participants understand rather than the product team’s module names.
04 Explain Less and Trace More
The moderator does not teach the user; they make the thought process visible. Explain that “We are testing the product, not you,” and invite participants to describe what they seek and why they hesitate.
When a user stops, use neutral prompts:
- “What are you looking for now?”
- “What does this suggest to you?”
- “What would you do next if I were not here?”
- “What do you expect to happen after clicking?”
Do not ask “Is this button hard to see?” or “Would it be better on the left?” Such questions put the answer in the user’s mouth.
A Reusable Moderator Introduction

05 Observe More Than “Completed” or “Failed”
Record at least six types of evidence for a task:
| Evidence | Example |
|---|---|
| Path | Entry point, detours, and backtracking |
| Understanding | Interpretation of labels, states, fees, permissions |
| Behavior | Clicks, search, comparison, repeated review, skipping |
| Error | Wrong selection, omission, duplicate submission, invalid attempt |
| Recovery | Whether the user notices and can correct the error |
| Emotional signals | Hesitation, confirmation, concern, surprise—without diagnosing psychology |
Observers should not debate design solutions during the session. Record issues and ideas separately so a strong opinion does not bias later observation.
06 Grade Issues by Impact, Frequency, Recovery, and Risk
Counting occurrences alone undervalues rare but high-risk problems. JVDS generally separates severity into four judgments:
1. Task impact: Does it block the core task or add only one step?
2. Scope: Does it affect one user, one role, or most users?
3. Recovery difficulty: Can users recognize and fix it themselves?
4. Business risk: Could it cause incorrect approval, payment, privacy exposure, or an irreversible action?
| Level | Typical Behavior | Recommended Response |
|---|---|---|
| P0 Blocker | Core task cannot be completed or severe security or business risk exists | Stop launch, fix, and retest |
| P1 Critical | Most target users fail or recovery is very difficult | Must fix in the current release |
| P2 Significant | Task completes, but understanding is costly, efficiency is poor, or help is required | Schedule soon and validate the solution |
| P3 Minor | Outcome is unaffected; issue concerns local copy or consistency | Combine with other improvements |
This is a project-management method, not a universal industry standard. Teams can adjust levels, but judgment rules must remain consistent and evidence must be retained.

07 Synthesize Quickly on the Same Day
Do not wait a week after every session before reviewing notes. Spend 15 minutes after each session answering:
- What new issue appeared today?
- Which issues repeated?
- What should the next session validate—not immediately change?
After all sessions, cluster related evidence into “Issue—Evidence—Impact—Cause Hypothesis—Recommendation—Validation.” Label whether a cause is inferred or verified so the design team’s explanation is not mistaken for a user fact.
08 Example: A Data Import Flow
A team tests an Excel import feature. The first task says only “Upload the file and complete import.” Users click the button successfully, and the team concludes the flow is fine.
In the redesigned task, participants must download a template, enter two invalid records, upload it, understand validation results, and repair failed rows. Three real issues emerge: users cannot distinguish “file uploaded” from “data imported”; errors do not identify the cell; and every correction requires reuploading the entire file.
The lesson is that tasks must cover the complete work loop, not a page. Otherwise, testing proves only that a button can be clicked.
09 Retesting Matters More Than “Changed”
A solution passing design review does not prove the problem is gone. Retest P0 and P1 issues when possible, especially after changing workflows, terminology, or information architecture. Retesting can be narrower and recruit only relevant roles, but tasks should remain comparable and new issues should be recorded.
Frequently Asked Questions
How many people does one usability-test round need?
No fixed number fits every project. A focused flow with one consistent user type can begin with a small round; major differences in role, device, or context require groups. Continue according to whether critical issues are still emerging and essential user groups are covered.
Can a moderator answer user questions?
It depends on the objective. To test independent support, delay help and record assistance points. To test work after training, provide the agreed support available in the real context. Keep the approach consistent across sessions.
Should testing be remote or in person?
Remote testing suits distributed users, digital products, and rapid iteration. In-person testing reveals devices, paper materials, collaboration, and the work environment. Choose according to real use, not tool convenience.
Does testing require a long report?
No. Iterative work needs a clear issue list, evidence clips, and decision record. Long reports support cross-team learning but cannot replace a review where product, design, and engineering see the evidence together.
10 Put Research and Design Into Product Decisions
| Service | View |
|---|---|
| UI/UX design services | View service details |
| Project inquiry | Contact JVDS |
| Design and website articles | Read more related articles |