Skip to content
PageSpeed 100 as the delivery default
Inhalte

Sharpen Your Content: A Practical Test With Five People

How five participants reveal where your website copy stalls: set tasks instead of opinion questions, observe in silence, and publish the edit the same day.

15 min read UsabilityTextarbeitFeedbackBarrierefreiheit

After launch, things go quiet. The site is live, the copy is written, the business turns back to day-to-day work. Whether the content actually lands, nobody finds out: people who get stuck on a page rarely complain — they leave. That gap can be closed with surprisingly little effort. Five people, twenty to thirty minutes each, a handful of clearly worded tasks and an observer who keeps quiet: that is all the cheapest quality gain after launch requires. This article covers who to invite, how to phrase tasks instead of opinion questions, what to log, how observations turn into concrete text edits — and which questions five people explicitly cannot answer.

A practical test with five people: from observation to editSet tasks, stay quiet, log what happens, then rewriteFive test participantsExisting customerLaptopNeighbouring firmLaptopPerson without jargonPhoneAn older personTabletPhone-only userPhoneOne separate session per personCumulative findings85 %of findings1 person31 %3 people67 %5 people85 %Discovery rate per person: 31 %What gets rewrittenHesitation at “range of services”Heading names the actual serviceScrolling back to find a pricePrice range below the serviceQuestion: “do you cover us?”Town and area in the first lineEdit published the same dayOne session per person: 20 to 30 minutesSet tasks, do not explain1Set the task2Observe only3Log hesitation4Edit the text5Live same dayThree rounds of five people beat one round of fifteen (Nielsen Norman Group)

Why five people are enough

The number five is not a gut-feel rule of thumb. It comes from an analysis of many projects by the Nielsen Norman Group: a single test participant uncovers on average 31 percent (Nielsen Norman Group) of the usability problems present. Because findings overlap — the second person trips over some of the same spots as the first — the yield does not rise evenly, it flattens quickly. After five people it sits at roughly 85 percent (Nielsen Norman Group). To find every problem in one version, around 15 people (Nielsen Norman Group) would be needed, and the last ten mostly repeat what the first five already showed.

For a business without a test lab, that is the decisive message: the insight gained per hour invested is highest at the very start. Half a day of effort delivers by far the largest part of what an elaborate study would have produced. The remainder is fine-tuning and can be collected in the next round.

If the budget stretches to fifteen test participants, it does not belong in one large study but in three studies of five people each — with a revision between the rounds.

Paraphrased from: Nielsen Norman Group (Jakob Nielsen), Why You Only Need to Test with 5 Users

One condition applies, though: the audience has to be reasonably homogeneous. Five people are enough when all of them approach the site with a comparable intention. If a website serves two clearly different groups — private clients and a company's purchasing department, say — the Nielsen Norman Group recommends three to five people per group (Nielsen Norman Group) rather than five in total. For most trade, service, practice and hospitality businesses the circle is narrow enough that one round of five people holds up.

What the test shows — and what it does not measure

A test with five people is a qualitative method. It shows where people get stuck and why, and it produces evidence you can look at. It does not produce percentages about „our customers", no verdict on taste and no ranking of drafts. Anyone deriving figures from it is overstretching the method. The value lies in the observed causes, not in how often they occur.

The practical test does not replace solid copywriting — it shows where the copy stalls. Anyone still working on the basics should first go through the rules for clear website copy and test afterwards. Anyone still planning a site should settle purpose and audience in the website brief first; the test then checks whether that intention actually arrived.

Who to invite — and who not to

The selection decides the return. A representative survey of 1,004 people (Bitkom Research) aged 16 and over, conducted in 2025, shows how people in Germany look for a trade business: 83 percent (Bitkom Research) ask friends, family or colleagues, 41 percent (Bitkom Research) gather information locally, 26 percent (Bitkom Research) look at company websites and 15 percent (Bitkom Research) read online reviews. At the same time, 40 percent (Bitkom Research) treat a missing website as a reason to rule a business out, and 55 percent (Bitkom Research) find businesses with digital services more attractive.

The consequence for recruiting is clear: the website is rarely the first point of contact, but frequently the vetting one. It is opened by people who have already heard a name and now want to see whether the business fits. So the test should not run with the enthusiastic, but with the vetting — and with those who have the least practice.

An existing customer

Knows the service but not the site. This person notices at once when the website promises something different from what the business actually does — and says so, because the relationship can carry it.

A business from the neighbourhood

Not from your own trade. Someone who runs a website themselves reads in a structured way and spots gaps in process and contact route without swimming along in the jargon.

Someone without domain knowledge

A family member or a friend from a completely different occupation. This reliably reveals which terms need explaining and which abbreviations run empty.

An older person

In the 65 to 74 age group, 10 percent (Federal Statistical Office of Germany) were offline most recently, against 3 percent (Federal Statistical Office of Germany) across all 16- to 74-year-olds. The rest are online, but often work with larger type, a slower pace and less routine.

Someone who only uses a phone

No laptop, no second screen, on the move, one-handed. This person finds the places where copy tips over on a small display — tables, long paragraphs, phone numbers you have to copy out.

Who is unsuitable: you

Anyone who wrote, commissioned or signed off the site cannot test it. Knowledge of the structure cannot be switched off. For the same reason the team is only of limited use as a reviewing instance.

Five people means five individual appointments, not a group session. Groups produce agreement: the first opinion sets the frame, the others fall in line, and the business ends up hearing one opinion instead of five observations. Individual sessions barely take longer in total, because a single run rarely needs more than twenty to thirty minutes.

How to ask without making a fuss

The most effective sentence is the most honest one: „We have reworked the website and are looking for five people to spend twenty minutes solving a few tasks on it. It is not about your opinion, it is about where things get stuck. There is nothing you can do wrong — if something does not work, that is on us." A small thank-you is customary and unproblematic, as long as it is not tied to a particular verdict.
  • Arrange sessions individually, in a quiet room, without an audience from the team.
  • Use the participant's own device where possible — type size, browser and habits are part of the finding.
  • State up front that the website is being tested, not the person.
  • No briefing on the structure of the site, no advance explanation of the services.
  • Name a time frame and keep to it; after thirty minutes attention is spent.
  • Ask whether you may take notes — notes are enough, a recording is dispensable.

Tasks instead of opinion questions

The most common mistake sits in the question. „How do you like the site?" produces politeness, not insight. People asked for a verdict deliver a verdict — usually a friendly one, often about colours, rarely about comprehension. Set a task instead and you get to watch behaviour. And behaviour cannot be softened out of politeness.

Instead of this questionSet this taskWhat it makes visible
How do you like the site?Find out whether we cover your area.Whether the service area is findable and unambiguously named
Is everything understandable?Tell me in one sentence what this business offers.Whether headings name the service instead of circling it
Is the page well organised?Request an appointment for next week.Whether the contact route sits where the decision is made
Are the prices acceptable?Estimate roughly what a job like this costs here.Whether any price range is recognisable at all
Does the business look reputable?Find out who runs the business and how long it has existed.Whether trust is evidenced or merely asserted
Does this work on a phone?Call us from this page without copying the number out.Whether the phone link exists and is recognisable as one

Good tasks have three properties. They have a verifiable outcome — the person found the service area or did not. They use the customer's language, not the site's: writing „range of services" into the task gives away the search term and makes the task worthless. And they describe a situation, not a function: „Your heating failed at the weekend — find out who you can reach and when" is more precise than „look for the emergency service".

  1. Open without guidance: „Look at the home page for twenty seconds. What does this business do, and for whom?"
  2. Service area: „Find out whether we come to you."
  3. Scope of service: „You need [specific occasion]. Find out whether we do that."
  4. Price range: „Estimate the order of magnitude, and show me what you base that on."
  5. Getting in touch: „Request an appointment — using the route you would actually choose."
  6. Reassurance: „Suppose you want to ask something first. How do you go about it?"

Three phrasings that void the test

„As you can see, the menu is up here" anticipates the result. „Find the ‚Request appointment‘ button" gives away the solution and only tests the eye. „Did you find that difficult just now?" invites reassurance. If such a phrasing slips out, note it in the log — that particular point no longer counts for this person.

Every task needs an ending

A task is complete when the person says they are done — or when they give up. Both are results, and giving up is the more valuable one. It marks the point where genuine prospects drop out too, except that they do not bother telling anyone about it.

Watch and stay quiet

The hardest exercise in the whole method lasts twenty minutes and consists of saying nothing. The urge to help is strong: you watch someone searching at the wrong end and want to cut it short. But every piece of help erases exactly the information the session was arranged for. Explain, and from that moment on you are testing your own explanation instead of the page.

  • Read the task out, then fall silent. Even through long pauses.
  • Reflect questions back: „What would you do if I were not sitting next to you?"
  • Do not point at the screen, do not scroll along, do not sigh.
  • No judgement in your face — approval and disappointment are both read.
  • Leave a short break between tasks instead of making the clock visible.
  • Ask follow-up questions only after the last task, and only about observed moments.

Asking people to think aloud helps: „Just say whatever goes through your mind." It feels artificial for the first two minutes and then becomes incidental. The gain is considerable, because a silent hesitation turns into a sentence — „I am not sure whether this applies to me" is a finding, a three-second pause is only a suspicion.

The four sentences that suffice during the test

„Please read the task back to me the way you understood it." — „Feel free to say out loud what you are looking for." — „What would you have expected at this point?" — „Just carry on the way you would at home." Everything else is either help or evaluation, and both distort the finding.

After the last task there is room for questions, but the rule still holds: ask about observations, not opinions. „You stopped at this point — what caused that?" is a question. „What could have been done better?" is an invitation to invent design proposals, which no test participant is responsible for.

What belongs in the log

Log behaviour with a timestamp and a location, not your interpretation. Five signals carry most of the insight: hesitation of more than three seconds, scrolling back, reading the same passage twice, a question directed at the observer, and abandoning the task. Add to that a click on something that is not clickable at all — a reliable sign that expectations about the interface are not being met.

ObservationWhat it usually meansWhere the cause sits
Hesitation of more than three secondsThe heading does not match the expectationSubheading, first sentence of a section
Scrolling backA piece of information came too lateOrder of the sections
Reading the same passage twiceSentence structure or jargon slows things downWord choice, sentence length, nested clauses
A question to the observerThe page does not answer the questionMissing content, not missing design
A click on a non-clickable elementInterface expectations are disappointedHow links and buttons are marked up
Abandoning the taskThe route is longer than the patienceNumber of steps to the goal

One sheet of paper per person is enough. What matters is that every line names a place you can later change — „heading on the services page" is a place, „feels cluttered" is not. After five sessions the sheets are laid side by side and all lines concerning the same place are merged.

log.txt
Participant 3  |  Phone  |  Task: request an appointment

00:20  scrolls past „range of services"        -> heading does not name the service
00:45  stops, scrolls back, reads twice        -> paragraph too long, term unclear
01:10  asks: „do you come out to us?"          -> service area sits too far down
01:35  taps a photo, expects a link            -> element looks clickable
02:05  finds the form, stalls at field 2       -> field label not understandable
02:40  gives up: „I would call at this point"  -> phone number missing here

Task completed: no
Places affected: 4 text spots, 1 ordering issue, 1 form field

One counting rule has proven itself: what two out of five people show goes on the list. What three or more show gets changed before the next round. What only one person shows is noted and deliberately watched in the following round instead of being acted on immediately — a single finding can also come down to a bad day.

From observation to text edit

This is where most tests fail: observation is clean and then nothing changes, because the translation from „hesitated" into „write it differently" is missing. That translation is craft, not creativity. Every observation corresponds to a bounded intervention, and in the vast majority of cases it is an intervention in the copy, not in the design.

Observation in the testConcrete changeHow success is checked
Three of five hesitate at „range of services"Heading names the service: „Heating, bathrooms and plumbing"The next round reads the heading without pausing
Two ask about the priceAdd a price range with a from-price and a reference unitThe price question disappears from the task
Four look for the service areaTown and radius in the first paragraph and in the footerThe service-area task takes under fifteen seconds
Everyone scrolls past the company historyMove the section behind services and processThe route to the contact option gets shorter
Two do not understand „refurbishment in occupied buildings"Replace the jargon with the activityThe follow-up question does not appear in the next round
One cannot find the phone numberPhone link at the end of every sectionThe call succeeds without searching

Headings name the service

„Our range of services" becomes a list of what is actually done. Headings are the only copy everyone reads — so they carry the information, not the mood.

Replace jargon

Every term two people stumbled over is replaced by the activity or explained in a subordinate clause at first mention. The test supplies the list; no guesswork required.

Add a price range

Where people ask about the price, an order of magnitude is missing. A from-price with a reference unit and one sentence on what moves the price answers the question without committing to a figure.

Reverse the order

Scrolling back is an ordering problem. What people search for belongs ahead of what you want to tell them: services and area first, then process, then company history.

Name town and area early

The most frequent follow-up question in practical tests amounts to „do you come out to us?" (project experience). Town, radius and travel time belong in the first paragraph and in the footer of every page.

Contact route at the point of hesitation

The contact route belongs where the decision is made — at the end of every section that explained a service, not exclusively on a page of its own.

Two changes per observation is one too many. Change the heading, the order and the imagery at once and the next round cannot tell you what worked. One change per finding, then test again — slower in theory and faster in practice. How to keep those changes in place over time is covered by the maintenance routine for small firms.

The test changes copy, not taste

Five observations dictate word choice, order and completeness. They do not dictate that the home page needs a different photo, that the type should feel bigger or that a colour ought to be „fresher". The moment the debrief turns to design, the method has been abandoned — and the next finding drowns in the discussion.

What five people cannot answer

The limits of the method matter as much as its yield, because overstretched results end up carrying expensive decisions they cannot carry. Five people deliver causes, not distributions. Anything starting with „how many" lies outside.

  • Matters of taste. Whether someone likes a photo has nothing to do with the task and is answered by five people neither representatively nor reliably.
  • Colour and layout debates. Contrast and legibility can be checked, preferences cannot. One belongs in the cross-check, the other in no debrief at all.
  • Statistical claims. „60 percent of our customers cannot find the form" cannot be derived from three out of five people. For robust figures the Nielsen Norman Group names at least 20 participants (Nielsen Norman Group).
  • Willingness to pay. What someone would be prepared to pay is not revealed by a test situation. The test only shows whether the price information was found and understood.
  • Search terms. Which words people actually type is shown by search data, not by five sessions — for that, finding the search terms customers use is the right route.
  • Rankings between drafts. Which of two drafts is „better" is not decided by a round of five people; it only shows what ails both of them.

Three sentences that do not follow from the test

„People want it this way." — Five people are not people in general, they are five pieces of evidence. „The problem is solved." — It is solved when the next round stops tripping over it. „We need a complete rebuild." — Whether that pays off is a separate calculation; when it adds up is covered in when a new website pays off.

Anyone who genuinely needs numbers needs a different method. For stable heatmaps the Nielsen Norman Group names 39 participants (Nielsen Norman Group), for a card sort at least 15 participants per group (Nielsen Norman Group). Those are orders of magnitude a small business can hardly manage — and, as a rule, does not need, because the coarse comprehension problems become visible long before that.

Cross-check with read-aloud and keyboard

Five sighted people with a mouse and a phone cover part of reality. The other part only shows up when a page is operated without a mouse or read aloud — and that is exactly where a striking number of sites fail. The third accessibility test report on German online shops examined 65 shops (Aktion Mensch) in 2025; only 20 of them, or 30 percent (Aktion Mensch), were fully operable by keyboard. The year before it was 15 out of 71 (Aktion Mensch), in 2023 17 out of 78 (Aktion Mensch). The assessment drew on 14 criteria (Aktion Mensch) from the Web Content Accessibility Guidelines across eight subject areas.

Accessibility on the internet is essential for ten percent of the population, necessary for around 30 percent and helpful for 100 percent.

Aktion Mensch, BITV-Consult, Google and Stiftung Pfennigparade, third test report on the accessibility of online shops (2025)

That these are not isolated cases is shown by WebAIM's automated analysis of one million home pages: 95.9 percent (WebAIM) showed detectable failures against the guidelines, on average 56.1 errors (WebAIM) per page, with 83.9 percent (WebAIM) carrying text at insufficient contrast. Two findings a practical test cannot surface at all, because participants compensate unconsciously — they hold the phone closer instead of complaining.

The cross-check costs a quarter of an hour and prevents the edits from the test undermining accessibility. It follows the Web Content Accessibility Guidelines 2.2, published as a W3C Recommendation in the version of 12 December 2024 (W3C), which add nine new success criteria (W3C) compared with version 2.1.

  1. Tab through the edited page: is the order sensible, and is the focus visible at all times? Focus must not be entirely hidden by overlaid content (success criterion 2.4.11, W3C).
  2. Switch on the read-aloud function built into the operating system and call up the list of headings: does the sequence of headings describe the structure of the page?
  3. Have the edited headings read aloud: do they describe their section even without the surrounding text?
  4. Check new links: does the link text alone say where it leads? A „click here" reference is empty without its surroundings.
  5. Measure buttons and phone links: at least 24 by 24 CSS pixels (W3C) target size under success criterion 2.5.8.
  6. Check the contrast of edited text: 4.5:1 for text (W3C), 3:1 for interface and graphical elements (W3C).
  7. Zoom the page to 200 percent: does the edited paragraph stay readable without horizontal scrolling?

Why the cross-check belongs to the test

Changes coming out of a practical test are almost exclusively text changes — and text is exactly what touches accessibility. A new heading can break the outline level, a new link can lose its context, an inserted price range can end up as an image instead of text. The quarter of an hour after each round of edits stops an improvement in one place from becoming a step backwards in another.

How comprehension and accessibility can be handled in the same piece of copywriting is covered in the article on accessible, readable copy; which obligations have applied to many businesses since June 2025 is set out in the overview of the European Accessibility Act in Germany. For businesses found mostly on phones, the notes on mobile usability are worth a look as well.

How often to repeat

A single test is a snapshot. The real return comes from repetition, in small rounds: three runs of five people with a revision in between deliver more than one run of fifteen (Nielsen Norman Group). The reason is simple — in round two nobody trips over what was fixed in round one, which makes the problems visible that the first obstacle had been hiding.

OccasionNumber of participantsTime requiredWhat is checked
Four to six weeks after launch5half a dayHome page, main service page, contact route
After every substantial text change3 to 5two hoursonly the changed passages
For a new service or a new audience3 to 5 per grouphalf a daythe new area in context
Once a year as a fixed routine5half a daythe full route from search to enquiry
  • Implement the changes between rounds — otherwise the second round repeats the first.
  • Invite fresh people for every round; anyone who knows the site finds the old spots too quickly.
  • Keep the tasks identical so the rounds stay comparable, and only add tasks for new content.
  • Keep the logs; after three rounds it becomes clear which places are stubborn.
  • Do not schedule rounds in advance — the best occasion is a change that has just been finished.

The best moment for the first round

Four to six weeks after launch. Earlier there is no distance from your own copy; later the business has grown used to the site and mistakes weaknesses for quirks. If a site is still being planned, the test works on the draft version too: tasks can be set on an unfinished structure just as well — and the timeline of a website project allows for an acceptance phase anyway.

Why fast implementation decides the value

The value of a test is created not by observing but by changing. And that is where it is decided whether the method survives everyday life: a heading corrected on test day is verified the next day. A heading that enters a queue as a change request is verified in four weeks — by which point the log, the context and the motivation have all faded. In practice, test rounds rarely fail because of participants and frequently because of the implementation loop (project experience).

Step after the testImplementation in your own handsImplementation through an external loop
Change a headingsame day, straight in the editordescribe the change request, wait for a slot
Swap the order of two sectionsa few minutescoordination, because the effect needs explaining
Add a price rangesame day, legal pages untouchedfollow-up questions on price statements, another round
Replace a term across five pagesone sitting, one approvalbundled request, because single edits are too small
Second test roundpossible within daysonly after the first changes are signed off
Effort per round of changesno separate commissioning requiredagreed separately depending on the contract

This is precisely where the ways of running a website differ more sharply than they do on the starting price. The comparison of implementation routes sets out how quickly changes actually become visible after launch; anyone who wants to see the path from first draft to publication will find it under how building a Flow Site works. How the three usual routes — self-build, agency and AI-supported build — differ in effort and response time is set out in builder, agency or AI.

And because most findings cluster around the contact route, two levers are worth reviewing before the first round: the contact form as a source of enquiries and the question whether and how prices are shown on the website. Both subjects regularly turn up as the first hurdle in practical tests. If you are still unsure which pages should exist at all, the baseline is set out in which pages a business website needs.

Five people, half a day, three changes

That is the realistic yield of a first round: three to five places where the page failed its job, and one change for each that can be described in minutes. No tool, no licence, no consulting day. The only condition is that the changes actually happen — ideally on the same day, while the observation is still fresh.

Sources and studies

This article is based on data from: the article „Why You Only Need to Test with 5 Users" by Jakob Nielsen (Nielsen Norman Group) with the analysis of the average discovery rate per participant and the saturation curve; the article „How Many Test Users in a Usability Study?" by the Nielsen Norman Group with recommended participant numbers for qualitative and quantitative studies, card sorting and eyetracking; the Web Content Accessibility Guidelines (WCAG) 2.2 of the World Wide Web Consortium (W3C) in the version of 12 December 2024, in particular success criteria 1.4.3, 1.4.11, 2.4.7, 2.4.11 and 2.5.8; the third test report „So barrierefrei sind Online-Shops in Deutschland" by Aktion Mensch, BITV-Consult, Google and Stiftung Pfennigparade from 2025 (65 e-commerce sites assessed, test period 5 March to 4 May 2025) together with Aktion Mensch publications on digital participation; the representative survey of 1,004 people aged 16 and over conducted by Bitkom Research in 2025 on searching for trade businesses; the WebAIM Million report on the automated analysis of one million home pages; and the surveys of the Federal Statistical Office of Germany (Destatis) on internet use among 16- to 74-year-olds. Own project experience from practical tests of small business websites is included as well.