What makes an AI assistant cite one source over another?

A model picks a passage, not a page. The winner is the one that answers on its own, is reachable without JavaScript, rests on a verifiable fact and carries a visible date. Position in Google results is neither a condition here nor a guarantee.

SeoScore

The question usually gets asked as "how do I get to the first position". But an AI answer has no positions — it has passages, and they are collected from more than one page.

Four things decide whether yours is among them.

1. Reachability

The most banal and most common reason: the crawler never got the text. Either robots.txt kept it out, or the content is drawn by JavaScript, which most AI crawlers do not execute.

There is nothing to optimise here — either the text is in the server response or it is not.

2. Self-containment

A model lifts a paragraph and drops it into an answer without the context of your page. So a paragraph starting "as we mentioned above" or "this depends on what we discussed" is unusable as a citation, even when it is correct.

The test is simple: cover the whole page, leave one paragraph, and ask whether it answers. If it does not, it will not be taken.

3. Corroboration

A claim that exists nowhere else is a risk to the model. A claim backed by Google's documentation, a standard or a study is safe.

Which leads to something most people do not do: external links to primary sources help you, not your competitor. They make your text checkable.

The other half is the same idea in reverse — your own numbers. A measurement you ran yourself and described precisely is the one fact for which a link to you is mandatory.

4. Recency

A visible date and dateModified in the schema are not cosmetics. For questions whose answer changes — prices, versions, rules — a model prefers the newer source even when the older one is more thorough.

What this does not mean

It does not mean position in Google stopped mattering: several assistants retrieve through search, so the index remains the entry condition. But among the first ten results, the ones cited are not the top ones — they are the ones whose passages fit.

Related articles

The same subject from another angle — each answers its own question:

How we measure this is described under what we check.

Sources

What this text rests on — primary sources, not retellings:

Four things that decide it — Reachability — did the crawler get the text at all; Self-containment — does the passage stand alone; Corroboration — is the fact found elsewhere too; Recency — is there a visible date.

Frequently asked

Does domain authority matter?

Indirectly. Authoritative domains get indexed more readily and get cited by others more often, and that is what corroboration is. But an "authority score" is not a Google metric, and nobody hands one to a model.

Does longer text get cited more?

No. What gets cited is a passage, not a document. Three thousand words without headings is one passage, and an answer to a single question cannot be lifted out of it.

Do I need `llms.txt`?

It does no harm and costs little, but no provider has publicly committed to reading it. Treating it as the solution would be a mistake; treating it as a sign of tidiness is fair.

How do I check how a model sees me?

Just ask it. The second way: give the assistant your page and ask it to answer using only that — you will see whether the text actually answers the question you think it answers.

Check yours

A free scan shows which of the things described here your site already does, and which it does not.

Scan a website →