“Test with five people” can be a reasonable starting point—or a source of dangerously false confidence. Sample size is not determined by the method’s name. It depends on what you need to learn, how much users differ, how precise the conclusion must be, and the cost of a wrong decision.
“Test with five people” can be a reasonable starting point—or a source of dangerously false confidence. Sample size is not determined by the method’s name. It depends on what you need to learn, how much users differ, how precise the conclusion must be, and the cost of a wrong decision.
01 Distinguish Two Goals: Finding Problems and Estimating Proportions
Qualitative research usually asks what problems exist, why they occur, and how users understand them. Quantitative research asks how many people are affected, how large the difference is, and whether it is stable. The former requires coverage of important differences and progress toward information saturation; the latter needs enough observations to control error and support comparisons.
| Goal | Typical Question | Sample Basis | Wrong Approach |
|---|---|---|---|
| Find problems | Why are users stuck? Which needs remain unmet? | Key-segment coverage, issue discovery, and saturation | Using headcount to prove a market proportion |
| Compare solutions | Is option A better than option B? | Expected difference, variance, confidence, and statistical power | Claiming a significant lift from a very small sample |
| Estimate a metric | What are the success, satisfaction, or usage rates? | Allowed error, confidence level, and population variation | Reporting only the mean without an interval |
| Iterate continuously | Does this round still contain high-risk issues? | Round objective, issue severity, and repeated testing | Recruiting a large group once and waiting until development is complete |
02 Practical Starting Points by Method
The numbers below are planning starting points, not universal industry standards. More heterogeneous audiences, more complex tasks, and riskier decisions require larger samples or additional rounds.
| Method | Common Starting Point | When to Add More | Main Output |
|---|---|---|---|
| Exploratory interviews | Begin with 5–8 people per key segment | New themes continue; roles, markets, or experience differ materially | Needs, motivations, language, and context |
| Formative usability testing | 5–8 people per relatively homogeneous segment per round | Flows are high risk, issues are rare, or assistive-technology users must be covered | Issue discovery and rapid iteration |
| Quantitative usability benchmark | Often 20–40 or more before evaluating precision | Versions or groups must be compared, or narrower error is required | Success, time, scale scores, and intervals |
| Survey | Calculate from population, proportion, error, and planned groups | Regions, customer tiers, or roles must be compared | Attitudes, proportions, and relationships |
| Card sorting or tree testing | Exploration may start with 30–50; stable structures often need more | Content is complex, segments differ, or statistical clustering is required | Categorization mental models and navigation findability |

03 When Can Interviews Stop? Look at Saturation, Not Only Headcount
Saturation does not mean two users in a row said the same thing. It means new interviews rarely change key themes, user differences, or the decision. After each session, record whether a new theme appeared, whether conditions around a known theme expanded, whether an assumption was overturned, and whether the product decision changed.
| Interview | New Themes | Additional Conditions | Decision Changed? | Recruit More? |
|---|---|---|---|---|
| 1–3 | Many | Many | Yes | Continue |
| 4–6 | Fewer | Important differences remain | Partially | Fill specific segment gaps |
| 7–8 | Very few | Mostly repeated evidence | No | Pause and validate |
If a project includes administrators, members, approvers, and external customers, five people across all four roles cannot justify saying “we reached saturation.” Evaluate saturation within the key segments relevant to the research question.
04 Why “Five Users Find 85% of Problems” Is Not a Universal Rule
That familiar claim depends on particular assumptions about issue probability and discovery goals. Real product problems do not all occur at the same rate. Permission exceptions, refund failures, assistive-technology compatibility, and infrequent high-risk flows may require specific participants and more scenarios to uncover.
Formative testing is better suited to “small samples, multiple rounds”: test a defined flow, fix high-risk problems, then run the next round. Compared with testing one obvious defect on 20 users at once, three rounds of five to eight people usually support iteration more effectively.

05 Quantitative Research Calculates Precision, Not a Convenient Round Number
When estimating task success, satisfaction, or a survey proportion, design the sample around acceptable error. Smaller samples produce wider intervals. Cutting the error in half generally requires roughly four times as many observations.
Eight successes among ten participants do not prove that the population success rate is exactly 80%. The confidence interval is wide: useful for identifying an obvious issue, but not for making a precise public claim. To compare two versions, define the smallest difference worth detecting and use it in a power calculation.
06 Segmentation Quickly Multiplies the Required Sample
A project may say it is recruiting 20 people while planning to compare new and returning users, managers and members, domestic and overseas markets, and mobile and desktop. After segmentation, only two or three remain in each group—insufficient for either qualitative coverage or quantitative comparison.
Remove unnecessary comparisons before adding participants. Research design is not about fitting every difference into one test; it is about selecting the differences the current decision truly requires.

07 Let Risk Determine Investment
| Decision Risk | Example | Recommended Strategy |
|---|---|---|
| Low risk and reversible | Button copy or information order | Test quickly with a small sample, then monitor behavior after launch |
| Medium risk affecting core conversion | Registration, trial, or purchase flow | Multiple usability rounds plus key-metric validation |
| High risk involving money or compliance | Payments, lending, healthcare, or permissions | Cover exception scenarios and special populations; add expert and quantitative validation |
| Long-term strategic decision | New market or product positioning | Multiple methods, segments, and research stages |
08 Recruitment Quality Matters More Than a Few Extra People
- Use real or highly similar target users instead of relying entirely on coworkers and friends.
- Base screening on behavior and experience, not only age and gender.
- Record role, usage frequency, device, permissions, and relevant context so differences can be interpreted.
- Reserve 10%–20% backups for no-shows, ineligible participants, and technical issues.
- For specialized populations, assistive-technology users, or professional roles, prioritize fit over sample size.
09 Sample Planning for Three Common Projects
Scenario 1: Redesigning a Corporate Website Inquiry Form
Use analytics and support records to locate the issue, then recruit five to eight real target customers for task testing. Cover markets separately when behavior differs substantially. After launch, monitor form starts, completions, qualified leads, and errors.
Scenario 2: Evaluating a New Approval Flow in Enterprise Software
Cover requesters, approvers, and administrators rather than blending roles. Start with roughly five participants per key role in one round, focusing on states, permissions, returns, and exceptions. Expand the sample if efficiency must be compared quantitatively.
Scenario 3: Estimating Customer Satisfaction with a Survey
Define the population, expected response rate, acceptable error, and customer tiers to compare before calculating sample size. If only a small number respond, report the sample and likely bias honestly rather than presenting it as the view of all customers.
Frequently Asked Questions
Are six user interviews enough?
If the question is focused, users are relatively homogeneous, and new themes have clearly declined, six may be enough to form the next hypothesis. It is far from enough for multiple roles, markets, or complex workflows.
Does usability testing always require five people?
No. Five is a common formative starting point for quickly discovering frequent issues. Add participants or rounds for rare, high-risk, or segmented problems and when quantitative metrics are required.
Does a survey need at least 100 responses to be valid?
There is no fixed threshold. One hundred may answer a rough question but fail to support multiple group comparisons. Calculate from population, expected proportion, error, and planned analysis.
Does a small-sample study still have value?
Yes, when the conclusion matches the evidence. Small samples can reveal problems, generate hypotheses, and explain context, but should not estimate precise proportions or represent all users.
| Service | View |
|---|---|
| UI/UX design | View service details |
| Project inquiry | Contact JVDS Design Studio |
| Design and web articles | Read more articles |