After launch, things go quiet. The site is live, the copy is written, the business turns back to day-to-day work. Whether the content actually lands, nobody finds out: people who get stuck on a page rarely complain — they leave. That gap can be closed with surprisingly little effort. Five people, twenty to thirty minutes each, a handful of clearly worded tasks and an observer who keeps quiet: that is all the cheapest quality gain after launch requires. This article covers who to invite, how to phrase tasks instead of opinion questions, what to log, how observations turn into concrete text edits — and which questions five people explicitly cannot answer.
Why five people are enough
The number five is not a gut-feel rule of thumb. It comes from an analysis of many projects by the Nielsen Norman Group: a single test participant uncovers on average 31 percent (Nielsen Norman Group) of the usability problems present. Because findings overlap — the second person trips over some of the same spots as the first — the yield does not rise evenly, it flattens quickly. After five people it sits at roughly 85 percent (Nielsen Norman Group). To find every problem in one version, around 15 people (Nielsen Norman Group) would be needed, and the last ten mostly repeat what the first five already showed.
For a business without a test lab, that is the decisive message: the insight gained per hour invested is highest at the very start. Half a day of effort delivers by far the largest part of what an elaborate study would have produced. The remainder is fine-tuning and can be collected in the next round.
If the budget stretches to fifteen test participants, it does not belong in one large study but in three studies of five people each — with a revision between the rounds.
One condition applies, though: the audience has to be reasonably homogeneous. Five people are enough when all of them approach the site with a comparable intention. If a website serves two clearly different groups — private clients and a company's purchasing department, say — the Nielsen Norman Group recommends three to five people per group (Nielsen Norman Group) rather than five in total. For most trade, service, practice and hospitality businesses the circle is narrow enough that one round of five people holds up.
What the test shows — and what it does not measure
The practical test does not replace solid copywriting — it shows where the copy stalls. Anyone still working on the basics should first go through the rules for clear website copy and test afterwards. Anyone still planning a site should settle purpose and audience in the website brief first; the test then checks whether that intention actually arrived.
Who to invite — and who not to
The selection decides the return. A representative survey of 1,004 people (Bitkom Research) aged 16 and over, conducted in 2025, shows how people in Germany look for a trade business: 83 percent (Bitkom Research) ask friends, family or colleagues, 41 percent (Bitkom Research) gather information locally, 26 percent (Bitkom Research) look at company websites and 15 percent (Bitkom Research) read online reviews. At the same time, 40 percent (Bitkom Research) treat a missing website as a reason to rule a business out, and 55 percent (Bitkom Research) find businesses with digital services more attractive.
The consequence for recruiting is clear: the website is rarely the first point of contact, but frequently the vetting one. It is opened by people who have already heard a name and now want to see whether the business fits. So the test should not run with the enthusiastic, but with the vetting — and with those who have the least practice.
An existing customer
Knows the service but not the site. This person notices at once when the website promises something different from what the business actually does — and says so, because the relationship can carry it.
A business from the neighbourhood
Not from your own trade. Someone who runs a website themselves reads in a structured way and spots gaps in process and contact route without swimming along in the jargon.
Someone without domain knowledge
A family member or a friend from a completely different occupation. This reliably reveals which terms need explaining and which abbreviations run empty.
An older person
In the 65 to 74 age group, 10 percent (Federal Statistical Office of Germany) were offline most recently, against 3 percent (Federal Statistical Office of Germany) across all 16- to 74-year-olds. The rest are online, but often work with larger type, a slower pace and less routine.
Someone who only uses a phone
No laptop, no second screen, on the move, one-handed. This person finds the places where copy tips over on a small display — tables, long paragraphs, phone numbers you have to copy out.
Who is unsuitable: you
Anyone who wrote, commissioned or signed off the site cannot test it. Knowledge of the structure cannot be switched off. For the same reason the team is only of limited use as a reviewing instance.
Five people means five individual appointments, not a group session. Groups produce agreement: the first opinion sets the frame, the others fall in line, and the business ends up hearing one opinion instead of five observations. Individual sessions barely take longer in total, because a single run rarely needs more than twenty to thirty minutes.
How to ask without making a fuss
- Arrange sessions individually, in a quiet room, without an audience from the team.
- Use the participant's own device where possible — type size, browser and habits are part of the finding.
- State up front that the website is being tested, not the person.
- No briefing on the structure of the site, no advance explanation of the services.
- Name a time frame and keep to it; after thirty minutes attention is spent.
- Ask whether you may take notes — notes are enough, a recording is dispensable.
Tasks instead of opinion questions
The most common mistake sits in the question. „How do you like the site?" produces politeness, not insight. People asked for a verdict deliver a verdict — usually a friendly one, often about colours, rarely about comprehension. Set a task instead and you get to watch behaviour. And behaviour cannot be softened out of politeness.
| Instead of this question | Set this task | What it makes visible |
|---|---|---|
| How do you like the site? | Find out whether we cover your area. | Whether the service area is findable and unambiguously named |
| Is everything understandable? | Tell me in one sentence what this business offers. | Whether headings name the service instead of circling it |
| Is the page well organised? | Request an appointment for next week. | Whether the contact route sits where the decision is made |
| Are the prices acceptable? | Estimate roughly what a job like this costs here. | Whether any price range is recognisable at all |
| Does the business look reputable? | Find out who runs the business and how long it has existed. | Whether trust is evidenced or merely asserted |
| Does this work on a phone? | Call us from this page without copying the number out. | Whether the phone link exists and is recognisable as one |
Good tasks have three properties. They have a verifiable outcome — the person found the service area or did not. They use the customer's language, not the site's: writing „range of services" into the task gives away the search term and makes the task worthless. And they describe a situation, not a function: „Your heating failed at the weekend — find out who you can reach and when" is more precise than „look for the emergency service".
- Open without guidance: „Look at the home page for twenty seconds. What does this business do, and for whom?"
- Service area: „Find out whether we come to you."
- Scope of service: „You need [specific occasion]. Find out whether we do that."
- Price range: „Estimate the order of magnitude, and show me what you base that on."
- Getting in touch: „Request an appointment — using the route you would actually choose."
- Reassurance: „Suppose you want to ask something first. How do you go about it?"
Three phrasings that void the test
Every task needs an ending
Watch and stay quiet
The hardest exercise in the whole method lasts twenty minutes and consists of saying nothing. The urge to help is strong: you watch someone searching at the wrong end and want to cut it short. But every piece of help erases exactly the information the session was arranged for. Explain, and from that moment on you are testing your own explanation instead of the page.
- Read the task out, then fall silent. Even through long pauses.
- Reflect questions back: „What would you do if I were not sitting next to you?"
- Do not point at the screen, do not scroll along, do not sigh.
- No judgement in your face — approval and disappointment are both read.
- Leave a short break between tasks instead of making the clock visible.
- Ask follow-up questions only after the last task, and only about observed moments.
Asking people to think aloud helps: „Just say whatever goes through your mind." It feels artificial for the first two minutes and then becomes incidental. The gain is considerable, because a silent hesitation turns into a sentence — „I am not sure whether this applies to me" is a finding, a three-second pause is only a suspicion.
The four sentences that suffice during the test
After the last task there is room for questions, but the rule still holds: ask about observations, not opinions. „You stopped at this point — what caused that?" is a question. „What could have been done better?" is an invitation to invent design proposals, which no test participant is responsible for.
What belongs in the log
Log behaviour with a timestamp and a location, not your interpretation. Five signals carry most of the insight: hesitation of more than three seconds, scrolling back, reading the same passage twice, a question directed at the observer, and abandoning the task. Add to that a click on something that is not clickable at all — a reliable sign that expectations about the interface are not being met.
| Observation | What it usually means | Where the cause sits |
|---|---|---|
| Hesitation of more than three seconds | The heading does not match the expectation | Subheading, first sentence of a section |
| Scrolling back | A piece of information came too late | Order of the sections |
| Reading the same passage twice | Sentence structure or jargon slows things down | Word choice, sentence length, nested clauses |
| A question to the observer | The page does not answer the question | Missing content, not missing design |
| A click on a non-clickable element | Interface expectations are disappointed | How links and buttons are marked up |
| Abandoning the task | The route is longer than the patience | Number of steps to the goal |
One sheet of paper per person is enough. What matters is that every line names a place you can later change — „heading on the services page" is a place, „feels cluttered" is not. After five sessions the sheets are laid side by side and all lines concerning the same place are merged.
Participant 3 | Phone | Task: request an appointment
00:20 scrolls past „range of services" -> heading does not name the service
00:45 stops, scrolls back, reads twice -> paragraph too long, term unclear
01:10 asks: „do you come out to us?" -> service area sits too far down
01:35 taps a photo, expects a link -> element looks clickable
02:05 finds the form, stalls at field 2 -> field label not understandable
02:40 gives up: „I would call at this point" -> phone number missing here
Task completed: no
Places affected: 4 text spots, 1 ordering issue, 1 form fieldOne counting rule has proven itself: what two out of five people show goes on the list. What three or more show gets changed before the next round. What only one person shows is noted and deliberately watched in the following round instead of being acted on immediately — a single finding can also come down to a bad day.
From observation to text edit
This is where most tests fail: observation is clean and then nothing changes, because the translation from „hesitated" into „write it differently" is missing. That translation is craft, not creativity. Every observation corresponds to a bounded intervention, and in the vast majority of cases it is an intervention in the copy, not in the design.
| Observation in the test | Concrete change | How success is checked |
|---|---|---|
| Three of five hesitate at „range of services" | Heading names the service: „Heating, bathrooms and plumbing" | The next round reads the heading without pausing |
| Two ask about the price | Add a price range with a from-price and a reference unit | The price question disappears from the task |
| Four look for the service area | Town and radius in the first paragraph and in the footer | The service-area task takes under fifteen seconds |
| Everyone scrolls past the company history | Move the section behind services and process | The route to the contact option gets shorter |
| Two do not understand „refurbishment in occupied buildings" | Replace the jargon with the activity | The follow-up question does not appear in the next round |
| One cannot find the phone number | Phone link at the end of every section | The call succeeds without searching |
Headings name the service
„Our range of services" becomes a list of what is actually done. Headings are the only copy everyone reads — so they carry the information, not the mood.
Replace jargon
Every term two people stumbled over is replaced by the activity or explained in a subordinate clause at first mention. The test supplies the list; no guesswork required.
Add a price range
Where people ask about the price, an order of magnitude is missing. A from-price with a reference unit and one sentence on what moves the price answers the question without committing to a figure.
Reverse the order
Scrolling back is an ordering problem. What people search for belongs ahead of what you want to tell them: services and area first, then process, then company history.
Name town and area early
The most frequent follow-up question in practical tests amounts to „do you come out to us?" (project experience). Town, radius and travel time belong in the first paragraph and in the footer of every page.
Contact route at the point of hesitation
The contact route belongs where the decision is made — at the end of every section that explained a service, not exclusively on a page of its own.
Two changes per observation is one too many. Change the heading, the order and the imagery at once and the next round cannot tell you what worked. One change per finding, then test again — slower in theory and faster in practice. How to keep those changes in place over time is covered by the maintenance routine for small firms.
The test changes copy, not taste
What five people cannot answer
The limits of the method matter as much as its yield, because overstretched results end up carrying expensive decisions they cannot carry. Five people deliver causes, not distributions. Anything starting with „how many" lies outside.
- Matters of taste. Whether someone likes a photo has nothing to do with the task and is answered by five people neither representatively nor reliably.
- Colour and layout debates. Contrast and legibility can be checked, preferences cannot. One belongs in the cross-check, the other in no debrief at all.
- Statistical claims. „60 percent of our customers cannot find the form" cannot be derived from three out of five people. For robust figures the Nielsen Norman Group names at least 20 participants (Nielsen Norman Group).
- Willingness to pay. What someone would be prepared to pay is not revealed by a test situation. The test only shows whether the price information was found and understood.
- Search terms. Which words people actually type is shown by search data, not by five sessions — for that, finding the search terms customers use is the right route.
- Rankings between drafts. Which of two drafts is „better" is not decided by a round of five people; it only shows what ails both of them.
Three sentences that do not follow from the test
Anyone who genuinely needs numbers needs a different method. For stable heatmaps the Nielsen Norman Group names 39 participants (Nielsen Norman Group), for a card sort at least 15 participants per group (Nielsen Norman Group). Those are orders of magnitude a small business can hardly manage — and, as a rule, does not need, because the coarse comprehension problems become visible long before that.
Cross-check with read-aloud and keyboard
Five sighted people with a mouse and a phone cover part of reality. The other part only shows up when a page is operated without a mouse or read aloud — and that is exactly where a striking number of sites fail. The third accessibility test report on German online shops examined 65 shops (Aktion Mensch) in 2025; only 20 of them, or 30 percent (Aktion Mensch), were fully operable by keyboard. The year before it was 15 out of 71 (Aktion Mensch), in 2023 17 out of 78 (Aktion Mensch). The assessment drew on 14 criteria (Aktion Mensch) from the Web Content Accessibility Guidelines across eight subject areas.
Accessibility on the internet is essential for ten percent of the population, necessary for around 30 percent and helpful for 100 percent.
That these are not isolated cases is shown by WebAIM's automated analysis of one million home pages: 95.9 percent (WebAIM) showed detectable failures against the guidelines, on average 56.1 errors (WebAIM) per page, with 83.9 percent (WebAIM) carrying text at insufficient contrast. Two findings a practical test cannot surface at all, because participants compensate unconsciously — they hold the phone closer instead of complaining.
The cross-check costs a quarter of an hour and prevents the edits from the test undermining accessibility. It follows the Web Content Accessibility Guidelines 2.2, published as a W3C Recommendation in the version of 12 December 2024 (W3C), which add nine new success criteria (W3C) compared with version 2.1.
- Tab through the edited page: is the order sensible, and is the focus visible at all times? Focus must not be entirely hidden by overlaid content (success criterion 2.4.11, W3C).
- Switch on the read-aloud function built into the operating system and call up the list of headings: does the sequence of headings describe the structure of the page?
- Have the edited headings read aloud: do they describe their section even without the surrounding text?
- Check new links: does the link text alone say where it leads? A „click here" reference is empty without its surroundings.
- Measure buttons and phone links: at least 24 by 24 CSS pixels (W3C) target size under success criterion 2.5.8.
- Check the contrast of edited text: 4.5:1 for text (W3C), 3:1 for interface and graphical elements (W3C).
- Zoom the page to 200 percent: does the edited paragraph stay readable without horizontal scrolling?
Why the cross-check belongs to the test
How comprehension and accessibility can be handled in the same piece of copywriting is covered in the article on accessible, readable copy; which obligations have applied to many businesses since June 2025 is set out in the overview of the European Accessibility Act in Germany. For businesses found mostly on phones, the notes on mobile usability are worth a look as well.
How often to repeat
A single test is a snapshot. The real return comes from repetition, in small rounds: three runs of five people with a revision in between deliver more than one run of fifteen (Nielsen Norman Group). The reason is simple — in round two nobody trips over what was fixed in round one, which makes the problems visible that the first obstacle had been hiding.
| Occasion | Number of participants | Time required | What is checked |
|---|---|---|---|
| Four to six weeks after launch | 5 | half a day | Home page, main service page, contact route |
| After every substantial text change | 3 to 5 | two hours | only the changed passages |
| For a new service or a new audience | 3 to 5 per group | half a day | the new area in context |
| Once a year as a fixed routine | 5 | half a day | the full route from search to enquiry |
- Implement the changes between rounds — otherwise the second round repeats the first.
- Invite fresh people for every round; anyone who knows the site finds the old spots too quickly.
- Keep the tasks identical so the rounds stay comparable, and only add tasks for new content.
- Keep the logs; after three rounds it becomes clear which places are stubborn.
- Do not schedule rounds in advance — the best occasion is a change that has just been finished.
The best moment for the first round
Why fast implementation decides the value
The value of a test is created not by observing but by changing. And that is where it is decided whether the method survives everyday life: a heading corrected on test day is verified the next day. A heading that enters a queue as a change request is verified in four weeks — by which point the log, the context and the motivation have all faded. In practice, test rounds rarely fail because of participants and frequently because of the implementation loop (project experience).
| Step after the test | Implementation in your own hands | Implementation through an external loop |
|---|---|---|
| Change a heading | same day, straight in the editor | describe the change request, wait for a slot |
| Swap the order of two sections | a few minutes | coordination, because the effect needs explaining |
| Add a price range | same day, legal pages untouched | follow-up questions on price statements, another round |
| Replace a term across five pages | one sitting, one approval | bundled request, because single edits are too small |
| Second test round | possible within days | only after the first changes are signed off |
| Effort per round of changes | no separate commissioning required | agreed separately depending on the contract |
This is precisely where the ways of running a website differ more sharply than they do on the starting price. The comparison of implementation routes sets out how quickly changes actually become visible after launch; anyone who wants to see the path from first draft to publication will find it under how building a Flow Site works. How the three usual routes — self-build, agency and AI-supported build — differ in effort and response time is set out in builder, agency or AI.
And because most findings cluster around the contact route, two levers are worth reviewing before the first round: the contact form as a source of enquiries and the question whether and how prices are shown on the website. Both subjects regularly turn up as the first hurdle in practical tests. If you are still unsure which pages should exist at all, the baseline is set out in which pages a business website needs.
Five people, half a day, three changes
Sources and studies