1. Where the 40 per cent comes from — and what it does not say

The number goes back to the founding paper of the field: Aggarwal and colleagues introduced the term Generative Engine Optimization in 2024 and measured how text changes affect visibility in generated answers. Adding verbatim quotations lifted a visibility score there from 19.3 to 27.2 — a relative gain of roughly 41 per cent.

The critical survey of 2026 puts that value in its place: it is a relative maximum on one metric under one particular configuration. Not an average, and not a value that could be reproduced on the open web. Anyone who reads the number as a promise is reading it wrong.

More decisive is a property of the experimental setup that almost always gets lost in the retelling: the measurements were taken on documents that had already been retrieved into the context. So the study says what additionally helps a text that has already been found. It says nothing about how a text gets found in the first place.

All known GEO effects are conditional on retrieval. What is not fetched cannot be cited, whatever you change in the text.

2. The two factors that work robustly

If you search the literature for what several independent papers show consistently, two things remain: fit between question and document and position in the context.

That is backed, among others, by a factorial design with around 252,000 runs that identifies both as the primary determinants of first citation. Neither is a GEO trick; both are classic quantities from information retrieval — the same ones that have counted in search engine optimisation for years.

In practice this means: a page that answers a concrete question concretely has most of the work behind it. Everything else is refinement.

3. What works moderately — and what it depends on

Two groups of measures have a medium level of evidence:

  • Extractable evidence — statistics, definitions, prices, clearly delimited facts. They work, but with varying strength depending on subject area and search intent.
  • Structure and freshness — headings, dates, clean sections. The effects are heterogeneous and depend on where in the processing chain the measurement is taken.

"Moderate" here means: observed several times, but not in every domain and not at every scale. Anyone who builds in such elements is doing nothing wrong. Anyone who expects a fixed percentage from them is overestimating the data.

4. What demonstrably does not work

Here the evidence is clearest, and it contradicts several popular pieces of advice.

Keyword stuffing does not carry over. What worked for a time in classic search engine optimisation has no effect in generative answer engines — the finding runs through several papers.

Generic recipes almost always fail. A benchmark called C-SEO Bench systematically tested methods against subject areas in 2025. Of 54 combinations of method and domain, a grand total of three were significantly positive. That is not noise around a positive mean; it is the opposite of a universal lever.

And optimisation can do harm. Perhaps the most important single finding comes from an arena study of 2026: optimisations that touched only the body text lowered presence in the top 20 by around 9 per cent, top-10 presence after reranking by 16 per cent, and actual citation by 6 per cent. Anyone who rewrites text for the machine can make it less usable for the retrieval stage before the citation stage is ever reached.

5. Why GEO measurements are so shaky

One reason so many contradictory field reports survive in this area is a matter of measurement:

  • Same question, different sources. A 2026 investigation found agreement between cited sources of only 0.34 to 0.42 across 24 hours — for an identical question. The recommended minimum is seven to eight repetitions per question before anything has been measured at all.
  • Small rephrasings flip the result. In a test across 30 question pairs, every single pair changed the cited domains on one model as soon as the question was rephrased.
  • The engines disagree with each other. The overlap of cited domains between different answer engines lies between 0.11 and 0.18. Optimising for one does not optimise for the other.

So anyone who measures once measures nothing. And anyone reading an agency figure without the number of repetitions is reading a snapshot.

6. The denominator problem: per cent of what, exactly?

One finding relativises most visibility reports at a stroke: in 57.8 per cent of the evaluated ChatGPT runs no search was triggered at all. The model answered from itself.

Anyone calculating rates only on the runs with a citation is therefore leaving out more than half the cases — and arriving at a figure that looks friendlier than the situation is. For practice this means: part of the answers in your subject area cannot be reached by any measure, because no source is consulted at all.

7. What this means for the work

Five sentences that hold up given this state of research:

  1. Findable first, optimisable second. Every measured GEO effect presupposes that the document was retrieved. Technical retrievability is not a preliminary step but the condition.
  2. Fit beats tricks. A page that answers a concrete question completely is using the only robust lever.
  3. No recipes across industries. Three out of 54 — plan measures for your subject area, not from a checklist.
  4. Do not tinker with body text without checking the retrieval stage. Body optimisation without regard to being found has measurably done harm.
  5. Measuring means repeating. Below seven to eight queries per question you get an anecdote, not a statement.

8. The honest remainder: what nobody knows

The critical survey ends with an admission that appears in no vendor brochure: there is so far no study showing a stable causal effect on organic findability that holds over time and across several platforms. The literature predominantly measures what happens after retrieval.

That is no reason to do nothing. It is a reason to reverse the order: first create the conditions under which your page can be considered at all — then write the text so that it answers the question. That is exactly what the second part of this series is about.

And it is a reason for caution with promises. A position paper from May 2026 additionally warns of undisclosed commercial influence: GEO makes it possible to embed advertising messages in apparently neutral content instead of labelling them. In China a case was documented in 2026 in which manipulated sources led language models to recommend invented products.

9. FAQ: frequently asked questions about GEO and AI citations

Does Generative Engine Optimization deliver 40 per cent more visibility?

No, not as a general statement. The number comes from a 2024 study and describes a relative gain on one particular metric, under one particular configuration, on documents that had already been retrieved. As an expected value for your own website it is of no use.

Which GEO measures are actually proven?

Two things are robustly proven: the fit between question and document, and the position in the context. Moderately proven are extractable evidence such as statistics and definitions, plus clear structure and freshness — with marked differences between subject areas.

Can GEO harm visibility?

Yes. An arena study of 2026 found that optimisations touching only the body text lowered presence in the top 20 by around 9 per cent and actual citation by 6 per cent. Anyone who rewrites for the answer stage can worsen the retrieval stage.

Why do I keep getting different sources for the same question?

Because generative answer engines do not answer deterministically. Across 24 hours the agreement of cited sources for an identical question was only between 0.34 and 0.42. Even rephrasing the question can swap the source list entirely.

How often do I have to measure before a number is reliable?

The minimum is seven to eight repetitions per question, plus several phrasings of the same question and several points in time. A single query is a snapshot, not a measurement.

Does GEO apply equally to all answer engines?

No. The overlap of cited domains between the engines lies between 0.11 and 0.18. An optimisation for one engine carries over to the others only in small part.

10. Sources and status

Status of the evaluation: September 2026. The basis is two freely accessible scientific papers and the individual studies evaluated in them.

Further reading on this site: Generative Engine Optimization — the fundamentals and llms.txt in 2026: barely read — and what actually works instead.

Stephan Michalik
About the Author
Stephan Michalik
Founder Grünberg.Digital. · CEO Flio Germany GmbH

Maximum performance through the synergy of experience and innovation: As Founder of Grünberg.Digital. and CEO of Flio Germany GmbH – a leading business incubator and enabler – Stephan Michalik designs holistic online marketing strategies. Whether precise SEA, high-revenue email marketing, or high-converting landing pages: He seamlessly combines these core disciplines with cutting-edge AI. The result: highly efficient, AI-powered marketing ecosystems for maximum digital advantage.

LinkedIn