Skip to content
PageSpeed 100 as the delivery default
KI-Sichtbarkeit

AI Search: How Small Businesses Get Cited

Where answer engines take their information from, which statements a page has to carry, and what a small business can do technically to be cited — without empty promises.

14 min read KI-SucheSichtbarkeitAntwortmaschinenrobots.txt

A growing share of questions is now answered without anyone opening a website. The answer appears directly in the search results or in a chat window, sometimes with a short source reference and sometimes without one. For a small business that shifts the task: it is no longer only about reaching the first results page, but about making sure a machine can find your details, understand them and name you when it answers. This article explains, calmly and in order, where answer engines take their information from, which statements a page has to carry, and which technical conditions sit behind that. It also says plainly what cannot be steered: citations cannot be bought and cannot be promised. Poorly structured pages, however, are passed over systematically — and that part can be changed.

AI search: how a small business gets citedClear statements, consistent business data, readable deliveryAnswers, not click pathsThe company pageServices named concretelyWhat, for whom, in which areaPrices with a unitHourly rate, call-out, flat feeContact, address, hourswritten identically everywhereFAQ from real questionsAnswers in full sentencesDate and accountabilityPage date, named contactAccess and deliveryrobots.txt set on purposeAllow or exclude fetchersllms.txt providedShort map of the contentText sits in the HTMLnothing loads afterwardsClean server responses200 instead of redirectsBusiness data from one sourceinserted into every outputThe answer given to a personA local business offers aweekend emergency service.Call-out 45 euros, hourly rate89 euros. Mon–Fri 7 to 17.Backed by a source referenceexample.comUpdated 07/2026Not something to promiseWhether you get cited is decidedper query by the machine.50 %use AI chats at least sometimesinstead of search engines (Bitkom)66 %of 16- to 29-year-olds useAI chats for answers (Bitkom)2 %of home pages serve a validllms.txt file (Web Almanac)

What is shifting in search right now

The starting point can be quantified. In a representative survey of internet users in Germany, 50 percent (Bitkom) said they use AI chats at least occasionally instead of a classic search engine, while 47 percent (Bitkom) stayed exclusively with classic search. Among 16- to 29-year-olds, the share using AI chats at least occasionally was 66 percent (Bitkom). For a business with five or with fifty employees this does not mean that search engines lose their role. It means that a growing share of questions first passes through an intermediate layer that reads, summarises and decides on its own which source it names. Which building blocks a company website needs for that is summarised in the overview of what XICflow offers.

The reason for the shift is practical. 33 percent (Bitkom) of users say they reach an answer faster through an AI chat than through classic search, and 73 percent (Bitkom) consider the results helpful. Only 36 percent (Bitkom), however, feel that those answers are backed by enough links. That gap is exactly where the work lies for a small business: if you want to be named, you have to make it easy for a machine to treat your page as a citable source. How this interlocks with classic visibility in your own catchment area is described in the article on local visibility for trades and hospitality.

One question has become two

The old question was: how do I reach the first results page? Alongside it stands a second one today: how does my page become the basis for an answer that somebody reads without ever opening it? Both questions rest on the same foundation — complete, unambiguous and current statements that can be read without detours. Answering the second one improves the first at the same time.

Where answer engines take their information from

The path is less spectacular than the term artificial intelligence suggests. At their core, answer engines draw on the same inventory as classic search: pages a crawler was allowed to fetch, that were accepted into an index, and whose text can be shown as a snippet. Only then does the summary come on top. Google's documentation states this explicitly for the AI features in Search.

There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.

Google Search Central, documentation on AI features in Search

That statement is both a relief and a demand. A relief, because there is no hidden switch to flip and no separate discipline that needs its own budget. A demand, because it puts the fundamentals in charge: fetchability, indexability, comprehensibility. The same documentation also notes that indexing and serving are not assured, even when every recommendation has been met (Google Search Central). Anyone who hears a firm promise about citations should therefore be sceptical.

Fetchable

The crawler is allowed to load the page, receives a clean status code and finds the text directly in the delivered HTML — no login, no chain of redirects, no waiting for parts that load afterwards.

Comprehensible

The page answers a recognisable question in complete sentences. Service, location, responsibility, price range and contact route appear as statements, not as a mood board of adjectives.

Attributable

Details can be traced to a responsible party: the business, the address, the date of the information, a named contact. What can be attributed can also be named as a source.

Unambiguous pages instead of marketing prose

An answer engine extracts statements from a page, meaning connections between a subject and a property. The sentence "We are your reliable partner for all things bathroom" contains none. The sentence "We renovate bathrooms within a radius of 40 kilometres, mostly in existing buildings, with a typical build time of two to four weeks" contains four. The difference is not a matter of style but of usability. Marketing prose is empty for machines, and it is nearly as empty for people.

Typical wordingWhat can be extractedRobust version
We offer you first-class serviceNothing verifiable: no object, no place, no timeframeRepairs on gas heating systems, response on the same working day until 4 p.m.
Expertise from a single sourceNo scope of work, unclear what is includedPlanning, execution and handover in house, electrical work through a fixed partner firm
Fair pricesNo order of magnitude, no unit of referenceHourly rate 89 euros net, call-out fee 45 euros net within 30 kilometres
Serving customers nationwideUsually contradicts everyday practiceService area: Hildesheim district, Hanover region, Peine district
Available at any timeNot usable, collides with opening hoursMonday to Friday 7 a.m. to 5 p.m., weekend emergency line on a separate number

Why precision counts twice here is shown by another figure from the same survey: 42 percent (Bitkom) of AI chat users say they have already received false information, and 57 percent (Bitkom) therefore verify the answers before using them. That verification often ends on the website of the business that was named. A page that confirms exactly what the answer said gains trust at that moment. A page that stays silent or says something different loses the enquiry — even though it was cited.

The sentence test for every page

Read a page and note down every sentence that carries a verifiable statement: a service, a place, a number, a deadline, a responsibility, a contact route. If fewer than five sentences remain, the page is practically empty for an answer engine. The fix is not more text but more substance: what exactly, for whom, where, in what timeframe, in what order of magnitude.

Name the services and prices that actually exist

A collective page that handles eight services in two lines each is hard for answer engines to use: title, description and heading can only be assigned once, and the text stays too thin for every individual service. One page per service solves this without inflating the structure. It describes the process, names the service area, answers the three questions that come up on the phone again and again, and ends with a contact route. In practice a small business needs five to ten such pages, not fifty.

  • One page per service, with a descriptive address containing the term customers actually use
  • The scope in complete sentences: what is included, what is not included, what runs through partner firms
  • The service area with concrete towns and districts rather than a vague regional label
  • Typical duration, typical process and the conditions that should be met before the appointment
  • An order of magnitude for the price, clearly labelled as an hourly rate, a flat fee, a starting price or a guide value
  • The details on VAT, call-out fees and additional costs, so that a single number cannot be misread
  • A section with the questions that are asked about exactly this service

Prices are the point where many businesses hesitate — and at the same time the point answer engines like to cite, because a number is an unambiguous statement. The way out lies in labelling, not in leaving them out: a starting price, an hourly rate or a range with a clear reference unit holds up as long as the conditions stand next to it and the figure is maintained. Price statements are subject to the same legal requirements as in a shop, for instance regarding total prices shown to consumers; which mandatory details a website needs anyway is set out in the article on imprint and privacy policy. What a transparent price presentation can look like is shown by the XICflow pricing overview.

Business details that read the same everywhere

Answer engines reconcile details. When the company name, address, phone number and opening hours differ slightly between the home page, the contact page, the imprint and external profiles, the result is not an error but uncertainty — and uncertainty leads to a different, less ambiguous source being named instead. The most effective lever is therefore unspectacular: every business detail should come from a single source and be generated from there everywhere. How that path runs from configuration to delivered page is described in the overview of how XICflow works.

  • Company name including legal form in one defined spelling, identical to the commercial register and the imprint
  • Address with the street type written out, without shifting abbreviations and without a PO box instead of the business address
  • Phone number in a single consistent format, linked so it can be dialled, with old lines actively retired
  • An email address that is genuinely read, complemented by a form as a second route
  • Opening or service hours per location, including different hours on public holidays and during company holidays
  • The service area in words: the towns, districts and counties where the business actually works
  • A date for the details, so it stays visible when someone last checked them

The most common silent error

Business details rarely diverge on purpose; they diverge over time. An extension is added, a name suffix disappears, the legal form changes, the summer opening hours stay in place. When the footer, the contact page, the location pages, the legal texts and the machine-readable markup are maintained separately, they drift apart within a few months (project experience). A shared data source prevents this with a single edit.

Opening hours deserve a separate remark, because they are among the most frequently requested details of all and at the same time go stale the fastest. Public holidays, bridge days and company holidays belong in the system before the date in question, not after it. Anyone running several locations should store them per location rather than collected on a single contact page; which structure has proven itself for that is set out in the article on local visibility for trades and hospitality.

Real questions, real answers

An FAQ section is the closest format to what an answer engine does, because it already has the shape in which answers are given: a question, and beneath it a self-contained reply. What matters is where the questions come from. Invented questions in an advertising tone are worthless; what works are the questions that are genuinely asked on the phone, by email and in the first conversation. They can be collected within two weeks if somebody in the business writes them down.

  1. For two weeks, note every question that comes in from outside — on the phone, in the form, at the counter, on site
  2. Merge duplicates and sort by frequency rather than by preference
  3. Write the question down the way it was asked, not in the internal jargon of the business
  4. Give the answer in three to five sentences, complete and without referring on to another page
  5. Spell out numbers, deadlines and conditions so the answer holds up out of context
  6. Assign the questions to the relevant service page instead of stacking everything on one collective page

A welcome side effect: the same answers relieve the phone and shorten the path to an enquiry. How to capture and process those enquiries cleanly afterwards is described in the article on the contact form as a source of enquiries. And anyone who wants to go beyond plain questions and publish subject matter regularly will find the relevant considerations in the article on blogging for small businesses.

Make currency and accountability visible

Answer engines prefer information that can be placed in context. That includes a visible date: when was this content published, when was it last revised? It also includes a responsible party: who stands behind the statement, which business, which person, at which address? The details that are legally mandatory anyway already provide half of that foundation; which ones those are in detail is set out in the article on imprint and privacy policy. The second half is editorial discipline: a date per page, not only in the blog.

In practice the most common problem shows up elsewhere: pages that have stood unchanged for years often contain details the business itself abandoned long ago — old services, old prices, old contacts. Those details keep being read and keep being cited. A half-yearly pass through the ten most important pages costs a few hours and prevents most of these cases (project experience). Among the pages that go stale particularly fast is the careers page; which details should hold up there is described in the article on the careers page for attracting skilled staff.

What a date has to achieve

An automatically running date that resets to today on every request is worthless and damages credibility. What helps is a publication date that stays put and a modification date that only moves when something substantive was actually changed. Both belong visibly on the page and additionally in the machine-readable markup.

The technical side: robots.txt, llms.txt and clean delivery

The robots.txt file sits in the root directory of a website and tells automated fetchers which areas they are allowed to load. The method has been described as a standard since 2022 (IETF, RFC 9309) and is therefore no longer specialist knowledge. It is well established in practice: around 85 percent (Web Almanac) of the examined requests returned a regular status code, roughly 13 percent (Web Almanac) ran into nothing, and 77 percent (Web Almanac) of the files that do exist work with a general rule for all fetchers.

What is new is that these files increasingly contain dedicated rules for the fetchers of AI providers. The share of files with a rule for individual ones of those programs rose noticeably within a year, for example from 2.9 to 4.5 percent (Web Almanac) for the most frequently named and from 1.9 to 3.6 percent (Web Almanac) for the second most frequent. For a small business this is a commercial decision, not a technical one: anyone who wants to appear in AI answers should deliberately allow the relevant fetchers. Anyone who wants to keep their content out of them can just as deliberately exclude them — and accepts appearing less often in such answers.

robots.txt is not access protection

A disallow rule in robots.txt controls fetching, not inclusion in an index. Google's documentation states explicitly that a blocked address can still appear in search results if other pages link to it — though without a description (Google Search Central). Anyone who genuinely wants content to stay private needs protection at server level. Which areas of a website should not be openly reachable in the first place is covered in the article on keeping a website's attack surface small.

One nuance matters because it is often misunderstood: some providers separate the fetcher used for search from a distinct token covering the use of content in AI models. For its separate token, Google states explicitly that blocking it does not affect inclusion in Search and is not used as a ranking signal (Google Search Central). Anyone who wants to limit model usage while staying visible in search can therefore decide the two separately — provided the rules in the file are written out cleanly and kept apart.

Alongside this, a second and considerably younger proposal has spread: llms.txt, a short text file that summarises the structure and the most important addresses of a website in readable form. It is not yet widespread — a valid file of this kind was found on around 2.1 percent (Web Almanac) of the home pages examined. Google also points out that no new machine-readable files are needed for its own AI features (Google Search Central). The effort is small nonetheless, and the benefit lies mainly in describing your own content inventory in an orderly way once. As a substitute for good pages it is useless; as an addition it does no harm.

Something else matters considerably more: the content has to sit directly in the delivered HTML. Pages that load their text only via JavaScript are sluggish for people on slow connections and a risk for fetchers, because at the moment of the request the text is simply not there yet. Static delivery solves both problems in one step — the text is in the file the server hands over. What that means for loading times and scoring is set out in the article on PageSpeed and static delivery. Clean server responses belong to the same picture: a regular status code instead of a chain of redirects, one reachable address per piece of content, and no error page pretending to be a success.

Machine-readable markup in the page head is the last building block, and deliberately worth only one paragraph here. It describes in a standardised vocabulary what the page contains — in its current release that vocabulary covers 823 types (schema.org). For answer engines it is no magic ingredient, but it is an additional, unambiguous reading of the same details. Which markup fits which page type and how to check the result is covered in depth in the article on structured data in search results.

What you can expect and what you cannot

This is the point for sobriety. Citations in AI answers cannot be bought, cannot be booked and cannot be promised by anyone. Which sources an answer engine names is decided per query, and the same question can lead to two different source lists on two different days. The providers themselves put it plainly: indexing and serving are not assured (Google Search Central). Anyone advertising a firm commitment about placements in answers is promising something they do not control.

ExpectationRealistic assessment
We will be named in AI answers from now onNot steerable. What is steerable is whether the page qualifies as a source at all
A tool makes the website AI-readyThere is no separate discipline, only complete details, clean delivery and maintained data
More text helps moreMore verifiable statements help. More prose without substance helps neither people nor machines
Citations replace website visitsSome visits fall away, others arrive pre-qualified. The page has to confirm the answer
This is a one-off projectDetails go stale. A regular pass through the most important pages keeps the inventory reliable

What can be steered, by contrast, is the opposite direction: a page without clear statements, with contradictory business details, with content hidden behind JavaScript and without a recognisable date will reliably be passed over by answer engines. That is the actual message of this article. It is less about buying an advantage than about removing an avoidable disadvantage.

An order of work you can actually get through

The following steps are deliberately sorted by effect per effort. A small business can work through them across a few weeks without the website sitting in a half-finished state in between. The sequence matters: first get the details right, then the structure, then the technology.

  1. Define the business details and have them inserted into every output from a single source
  2. One page per service, with verifiable statements on scope, area, process and order of magnitude
  3. Add price information, clearly labelled and with the conditions next to it
  4. Collect real customer questions for two weeks and assign them as FAQ sections to the matching pages
  5. Introduce a visible publication and modification date per page, and actually maintain it
  6. Check robots.txt: which fetchers are allowed, which are deliberately excluded, is the file reachable
  7. Add llms.txt as a short overview and check delivery: is the text in the HTML, are the status codes clean
  8. Set a half-yearly appointment to walk through the ten most important pages and replace anything outdated

In XICflow several of these points are part of generation rather than of rework. Pages are delivered statically, the entire text sits directly in the HTML and nothing is loaded afterwards. robots.txt and llms.txt are produced per website from the configuration, as is the machine-readable markup. Business details live in one place and are inserted into the footer, the contact page, the legal texts and the markup, so a changed phone number takes effect everywhere with a single edit. FAQ sections are a building block of their own that generates the matching markup at the same time. What that looks like in a finished result is shown by the example pages in the demos; how the approach differs from hand-built websites is set out in the side-by-side comparison.

The effort sits in the content, not in the technology

The technical conditions for visibility in AI answers can be created in an afternoon and then stay in order. The real work lies in writing down your own promise precisely enough that somebody without prior knowledge can understand it and pass it on. Doing that once, seriously, improves the website, the sales conversation and the prospect of being named in an answer.

How to build up such a set of pages from a briefing without ending up with a text kit that carries no substance is described in the article on creating a website with AI for small businesses.

Sources and studies

This article is based on data from: the representative Bitkom survey on internet search and the use of AI chats in Germany (1,156 respondents aged 16 and over), the Bitkom study on the use of artificial intelligence in the German economy, the Google Search Central documentation on AI features in Search, on robots.txt and on Google's crawlers, the Google Web Almanac based on the HTTP Archive (SEO chapter), the schema.org vocabulary, the IETF specification RFC 9309 on the Robots Exclusion Protocol, and our own project experience.