A page can rank well on Google and never appear in an AI-generated answer. A brand can appear in an AI-generated answer without its own website being cited. Its website can be cited while the brand itself is barely mentioned. A company can be mentioned frequently and still be described incorrectly. It can even become a recommended option while the evidence supporting that recommendation comes almost entirely from third-party publications.
Those outcomes are easy to collapse into a single idea called “AI visibility”. They should not be.
Much of the early discussion around Generative Engine Optimization (GEO) has inherited the mental model of traditional search: a user asks something, a system ranks sources, and the goal is to move upward. That model is convenient because the SEO industry already knows how to think about rankings.
It is also increasingly inadequate.
The better way to understand AI visibility is as the result of a multi-stage information pipeline. A source has to survive several distinct competitions before it becomes part of an answer, and success at one stage does not establish success at the next. That distinction changes how we should think about GEO strategy, experiments, content, citations, and even the products being built to measure this new search environment.
The ranking model breaks down surprisingly quickly
Imagine a hypothetical user asks:
What are the best platforms for monitoring a brand’s visibility across ChatGPT, Google AI Overviews, and Perplexity?
A conventional search mindset immediately asks which pages rank for “AI visibility platform”
A generative system can face a much broader task. It may need to understand what the user means by visibility, identify relevant tools, compare supported engines, evaluate use cases, retrieve evidence about pricing or functionality, assemble competing claims, and decide which sources are suitable to support the final answer.
Google now officially documents part of this process for its own generative Search experiences. Its current guidance explains that AI features can use retrieval-augmented generation, where information retrieved through Google’s Search systems is supplied to a model to help ground a response. Google also describes query fan-out, where the system may issue multiple related searches to gather information across subtopics rather than relying only on the user’s literal query. Google’s generative AI optimization guide and its documentation for AI features in Search are unusually explicit on this point.
This immediately creates a more useful mental model:
User intent → search or retrieval activation → query interpretation and possible fan-out → candidate retrieval → ranking or reranking → context selection → grounding → generation → citation → brand representation → user outcome
That sequence should not be mistaken for a claim that every commercial answer engine implements exactly the same architecture. ChatGPT, Google AI Mode, Perplexity, Gemini, Claude, and other systems have proprietary components whose details are not publicly known.
The value of the model is conceptual. It gives us separate failure points.
Retrieval and citation are different competitions
This is where many GEO discussions become imprecise.
Suppose an engine can access your page. That establishes accessibility, not visibility.
Suppose the page has been indexed by a search system. That establishes another form of eligibility, not retrieval for a particular question.
Suppose the source is retrieved as a candidate. That still does not prove it enters the final context supplied to the model.
Suppose its information is used in the answer. That does not guarantee the source will receive visible attribution.
And suppose the source is cited. That still does not tell us whether the associated brand is prominent, accurately represented, or recommended.
Google’s documentation is useful because it clearly preserves some of these boundaries. For AI Overviews and AI Mode, Google says supporting pages need to satisfy normal Search eligibility requirements. It does not introduce a separate technical schema or special “AI markup” requirement for inclusion. Google also says the two products can use different models and techniques and can surface different supporting links. Google Search Central therefore gives practitioners a strong reason to measure AI Overviews and AI Mode separately rather than combining them into a generic “Google AI” metric.
OpenAI documents a different part of the problem. Publishers that want their content discoverable in ChatGPT Search should allow OAI-SearchBot. OpenAI explicitly distinguishes this search crawler from GPTBot, which is associated with controls around potential model training. OpenAI’s publisher documentation is important because it separates two questions that are frequently confused: whether an engine can retrieve current web information, and whether content may be used for model development.
Again, crawl permission establishes opportunity. It does not establish selection.
Perplexity’s developer platform makes the separation even easier to see. Its Search API exposes ranked web retrieval as a product in its own right, while generated answers with citations are handled through its broader agent capabilities. We should not assume the consumer application is internally identical to the public API, but the architecture is still a useful demonstration of a broader information-retrieval principle: finding candidate evidence and composing an answer from evidence are separate operations.
This matters because optimization advice should identify which operation it is supposed to influence.
“Make the page crawlable” addresses one layer.
“Make the information relevant to the query space” addresses another.
“Make the evidence easy to understand and support” may help further downstream.
“Earn third-party coverage” potentially changes the external evidence environment.
None of those should casually be called a “ranking factor” unless we have evidence that establishes exactly that.
The most famous GEO statistic illustrates the problem
The foundational paper, GEO: Generative Engine Optimization, published at KDD 2024 by Pranjal Aggarwal and colleagues, was important because it converted a vague idea into an optimization problem.
The researchers created GEO-Bench and tested different interventions to source content. These included approaches involving quotations, statistics, citations, fluency, authoritative language, and other modifications. The experiments showed that source visibility inside generated answers could change significantly, with improvements reaching roughly 40% in some experimental conditions.
That result became one of the most repeated numbers in GEO.
It also became one of the easiest to misuse.
The paper does not establish a universal rule that adding statistics to an arbitrary webpage increases ChatGPT visibility by 40%. It does not establish a stable production ranking factor shared by commercial engines. And it does not prove that a content rewrite will cause a page that was previously absent from retrieval to become organically discoverable.
Those distinctions matter because the experimental environment determines what causal question a study actually answers.
A useful corrective comes from the 2026 preprint Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026). The survey reviews dozens of GEO studies and argues that much of the strongest evidence concerns what happens after candidate material has already become available to a generative system.
That is still important evidence. If changing how information is written changes whether an already-available source is used or cited, content design clearly matters.
But “this source wins more often after entering the competition” and “this source becomes more likely to enter the competition from the open web” are different claims.
The latter is much harder to prove.
The survey’s broader conclusion is uncomfortable but useful: GEO research still lacks strong longitudinal evidence demonstrating a stable, cross-platform causal intervention that predictably improves organic generative discoverability and downstream user behavior across live commercial engines.
For practitioners, that should increase rigor rather than decrease ambition.
The correct reaction is not that optimization is impossible. It is that we should become much more precise about which stage an experiment is measuring.
Query fan-out changes what “relevance” means
Traditional keyword research encouraged us to map one query or a cluster of closely related queries to a page.
Generative systems can complicate that mapping.
Take a hypothetical prompt:
Which GEO platform is best for a large agency managing 50 client brands?
The literal phrase “best GEO platform” captures only part of the information need.
A system performing query fan-out could potentially explore related concepts such as agency reporting, multi-brand management, supported AI engines, pricing, competitor tracking, multi-domain support, workflow features, or enterprise capabilities.
Google officially documents query fan-out in its own AI Search experiences. What we cannot claim is that we know the exact fan-out queries generated for every request, because those internals are not fully exposed.
The practitioner implication is still significant.
A useful unit of GEO planning may be an intent space rather than a keyword.
For the example above, the question becomes:
Does enough credible evidence exist across the web to establish this company as a strong answer to the different subproblems contained inside the buyer’s request?
That changes content strategy.
A company may have a perfectly optimized “best GEO software” page and still have weak evidence around agency workflows, multi-brand operations, integrations, pricing, third-party validation, or comparative performance.
If the answer engine retrieves evidence along several paths, visibility can depend on the completeness of the information environment around the entity.
This is a reasonable inference from documented query fan-out. It is not evidence of a universal “topic authority score” inside every answer engine.
That distinction between evidence and inference needs to become habitual in GEO.
Different engines should be treated as different environments
Another assumption worth abandoning is that “AI search” behaves like one channel.
Research increasingly suggests otherwise.
A 2025 empirical study, Generative Engine Optimization: How to Dominate AI Search, examined source behavior across generative search systems and reported meaningful differences in source selection, source diversity, freshness, language behavior, and sensitivity to query phrasing.
One of its interesting observations is the apparent importance of earned-media sources in the systems tested.
That finding is useful. Turning it into “AI engines always prefer earned media” would be irresponsible.
An empirical tendency in a particular benchmark is different from an official ranking rule. Models change. Search systems change. retrieval stacks change. Datasets differ. Geography and language differ. Query classes differ.
Even Google explicitly says AI Overviews and AI Mode can use different models and techniques and may surface different links. If two products from the same company should not automatically be merged into one measurement bucket, treating ChatGPT, Perplexity, Google, Gemini, and Claude as a single homogeneous channel makes even less sense.
This is why serious GEO measurement has to become engine-specific and probabilistic.
One answer is an anecdote, not a measurement
Generative output is stochastic.
Run the same prompt multiple times and the wording may change. Sources can change. Competitors may appear or disappear. Even semantically similar formulations of the same intent may retrieve different supporting evidence.
This creates a measurement problem that traditional rank tracking largely avoided. A SERP position can fluctuate, but asking “what rank did I hold?” still produces a relatively discrete observable.
Generated answers are multidimensional.
Consider a brand that appears in 7 out of 10 runs.
In five runs it is cited.
In three it is one of the primary recommendations.
In two it is mentioned only as an alternative.
Its own domain receives two citations, while a review site receives three.
The engine describes the product accurately in six runs and inaccurately in one.
What is its “AI rank”?
There is no single answer.
A more defensible measurement model would preserve separate observables such as:
Mention rate. Citation rate. Recommendation rate. Citation-domain distribution. Prominence. Narrative accuracy. Competitive share. Stability across runs. Stability across prompt paraphrases. Cross-engine variation.
Only after preserving those dimensions should a product consider creating a composite score.
This has a direct product implication for platforms such as Seerly. A useful AI visibility product should be able to explain why a brand’s observed performance is weak rather than simply report that the score is 42.
For example:
Your brand appears in 14% of agency-oriented prompts, while two competitors appear in more than 60%.
Most competitor citations in this intent cluster come from third-party category pages rather than vendor websites.
Your brand is frequently mentioned for monitoring but rarely recommended for multi-client workflows.
Those are diagnoses. A single score is a summary.
The distinction matters because only the diagnosis tells a team what to investigate next.
What this means in practice
The pipeline model changes several familiar optimization decisions.
Technical SEO still matters
For Google, this is officially documented. Google’s generative Search features still depend heavily on normal Search infrastructure, including crawling, indexing, page eligibility, and established Search quality systems.
The temptation to abandon SEO for exotic “LLM optimization” is therefore misplaced.
Good technical foundations remain necessary.
What changes is that technical eligibility becomes the beginning of the analysis rather than the end.
Content should create evidence, not merely target phrases
Google’s current generative AI optimization guidance emphasizes useful, unique, non-commodity information.
There is an intuitive reason for this.
If 500 pages repeat the same generic claim, an answer engine has little reason to depend on any one of them.
Original research, first-hand evidence, proprietary datasets, credible comparisons, clear product documentation, and expert analysis create information that is harder to substitute.
This does not give us a guaranteed citation formula. It creates a stronger information asset.
That is a more defensible goal.
Distribution matters because your website is not the whole evidence environment
If third-party sources repeatedly appear in the answers relevant to your category, your GEO strategy cannot stop at your own domain.
This is where digital PR, expert contributions, independent reviews, research distribution, and credible industry publishing potentially become relevant.
But the strategy should begin with observation:
Which sources actually appear for the intents that matter?
Only then should a team decide which publications, communities, datasets, or external sources deserve investment.
Starting with a universal list of “domains LLMs love” reverses the logic.
Measurement should use prompt families
A single tracked prompt is too fragile.
A stronger approach starts with a commercial or informational intent and creates several semantically equivalent formulations.
For example:
Best platforms for monitoring AI search visibility
Tools for measuring whether a company appears in ChatGPT and Google AI results
Software for tracking brand mentions and citations across answer engines
The exact set should reflect genuine user language, not artificial permutations.
Each variation can then be tested repeatedly across relevant engines.
The result becomes a distribution rather than an anecdote.
Recommendations should be treated as experiments
Suppose research suggests that statistics, citations, or stronger evidence presentation improve source utilization.
A responsible GEO recommendation system should say:
Published research provides evidence that this intervention can affect source utilization under certain conditions. Test it against a baseline for your prompt set.
It should not say:
Add three statistics to improve ChatGPT rankings.
One statement respects the actual evidence. The other manufactures certainty.
The unresolved question at the center of GEO
The field has already accumulated evidence that generative systems respond differently to different sources, content structures, prompts, and information environments.
What remains much harder is proving the complete causal chain:
I changed this page → the engine retrieved it more often → it entered context more often → it influenced more answers → the brand was represented more favorably → users acted differently.
Most GEO claims prove only part of that sequence, if they prove anything at all.
That is not unusual for a young discipline involving proprietary systems.
It does, however, create an important standard for anyone working seriously in the field:
Whenever someone says a tactic “improves AI visibility,” ask which stage improved, on which engine, under what methodology, across how many prompts, over how many runs, and compared with what control?
Those questions are more valuable than another list of supposed ranking factors. The strongest mental model for AI visibility is therefore a pipeline with observable boundaries.
A brand may be accessible but not retrieved. Retrieved but not selected. Selected but not cited. Cited but barely mentioned. Mentioned but misrepresented. Represented accurately but not recommended. Recommended without producing measurable commercial impact.
Once those stages are separated, GEO becomes much easier to reason about.
More importantly, it becomes much harder to fake certainty.