GEO is not the new SEO. It is a licensing decision with a formatting problem attached.

Every few months a discipline arrives with an acronym and a price list attached, and the honest question is whether there is anything underneath it. GEO — generative engine optimization, sometimes AEO — is unusual in that there genuinely is. It comes from a 2023 research paper rather than an agency deck, the paper built a benchmark, and it measured how much the techniques actually move the needle.

The measurement is the part that gets left out of the pitch, because it is not the number a pitch wants.

Where the term comes from, and what it measured

Aggarwal et al. introduced both the term and GEO-bench, a benchmark of diverse user queries paired with source documents, and demonstrated visibility improvements of up to roughly 40% inside generated answers. That is a real, replicated, published result and it is worth taking seriously.

It is also bounded in a way the marketing around it rarely reproduces. Forty per cent more visibility within a generated answer is not a rewrite of who gets recommended in your market. It is an improvement in how likely your existing page is to be selected and quoted, applied to a page that was already a plausible candidate. Nothing in the paper suggests you can format your way from invisible to authoritative.

The second finding matters more and is quoted even less: which technique works varies significantly by domain. The tactic that lifts a technical comparison page is not the tactic that lifts a local service page. Anyone selling you a single universal GEO playbook either has not read the paper or is hoping you have not.

What it actually shares with SEO, which is most of it

The technical fundamentals do not change at all. Fast pages, clean semantic HTML, a proper heading structure, schema markup — these help a classic crawler and a retrieval system for the same reason, which is that both need to work out what a page says and whether to trust it.

Google is unusually direct about this in its own documentation: AI Overviews and AI Mode draw on the same index and the same quality systems as classic Search, and there is no separate AI submission process. There is no second index to get into. If the page is not indexed and eligible to be shown with a snippet, no amount of GEO work reaches it.

So the useful framing is additive, not alternative. If your SEO foundation is weak, GEO does not route around it — it inherits the weakness. Most of the “we are not being cited anywhere” cases I have looked at were not content problems at all. They were a page that renders nothing without JavaScript, or a bot rule in a CDN that somebody switched on eighteen months ago and nobody remembers.

What GEO genuinely adds on top

Structured data becomes more directly consequential than it is in classic SEO. FAQ, HowTo, and clear Organization and Article markup give a retrieval system an unambiguous way to extract a fact rather than infer one. Schema.org is the shared vocabulary underneath all of it, and it is the least glamorous, most reliably useful work in this entire area.

Content that answers one specific, narrow question outperforms broad pages, and the mechanism is worth understanding rather than just obeying. These systems are lifting a passage. A page with a direct answer near the top of each section gives them something clean to lift; a page that spends three paragraphs warming up gives them a choice between quoting the preamble and quoting nobody. Structure is doing real work here, not cosmetic work.

Then there is the half you cannot edit. Being mentioned elsewhere — press, directories, industry roundups, comparison pieces — feeds retrieval in a way that is harder to measure than a backlink and is nonetheless real. Consistency matters here more than volume: answer engines cross-reference sources rather than trusting one, so a business described three different ways in three places gets attributed less confidently than one described identically everywhere. That is an unglamorous afternoon of work on your directory listings, and it beats another rewrite of your homepage.

The ceiling nobody puts in the proposal

Here is the number that should shape the budget conversation. Pew Research Center ran a panel study of 68,879 searches and found that when an AI summary is present, 8% of visits result in a click on a traditional result, against 15% when no summary appears. Clicks on a link inside the summary itself: 1%.

Read that carefully, because it cuts both ways and both directions are real.

The pessimistic reading is that being cited is worth far less traffic than being ranked used to be. A citation is not a visit. If your entire model depends on click volume from informational queries, AI summaries are a structural problem that no amount of GEO fixes, because the technique optimises for being chosen inside a box that people mostly do not click.

The optimistic reading is that the clicks are collapsing whether or not you participate. If summaries appear on your queries, the traffic is going regardless, and the choice is between being the cited source inside the answer and not being mentioned at all. Being named in the answer a buyer reads has value that does not show up in your analytics — which is uncomfortable to budget for and still true.

What I would not do is let anyone sell GEO to you on traffic projections. The mechanism it improves is citation, and citation converts to sessions at roughly 1%.

The prerequisite that rarely gets mentioned

Before any of the formatting work matters, the engines have to be allowed to use you, and that is a separate decision from being crawled.

Google splits it with a robots.txt token called Google-Extended. It is not a crawler — Googlebot still does the fetching. It governs whether what was fetched may be used to train future Gemini models and for grounding. And Google’s own wording on the consequence is unusually unambiguous: Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.”

That single sentence turns this into a licensing decision rather than an SEO one, which is a much better decision to be making. You can stay fully indexed and fully ranked while opting out of the generative uses. It is a real choice, not a trap with a ranking penalty hidden in it.

Which way to go depends on what you sell. If being cited inside AI answers is the point — you want the mention, the traffic was never the model — leave it open. If your writing is the product, closing it costs you nothing in rankings, and that is worth knowing before someone tells you that blocking AI crawlers will hurt your SEO. Through this token, it will not.

The one thing it will not do is speak for anyone else. Every operator publishes its own token and honours its own rules, so Google’s covers Google. Treating AI access as a single on/off switch is the most common mistake in this area. In practice it is a short list of separate decisions, one per operator, and it belongs in a policy document rather than in a robots.txt file somebody edits from memory.

How to tell whether any of it is working

This is the honest weak point, and pretending otherwise is how this field loses credibility. Measurement tooling for AI citation is immature next to classic search analytics.

The most reliable method available right now is unglamorous: take your actual target queries, run them against ChatGPT, Perplexity and Google AI Overviews, and record whether you are cited. Alongside that, watch your analytics for referral traffic from AI platforms, knowing it will undercount.

One warning that costs people real money. Do not measure this with a single check. These systems regenerate their answers constantly, and a screenshot taken on a Tuesday is closer to a coin flip than to a data point. Run a fixed prompt list repeatedly across several days and report the share of runs in which you appear — and freeze that list early, because changing it mid-measurement destroys the comparison. I wrote up the volatility data behind that advice separately, in How long until ChatGPT and Perplexity actually cite you?, if you want the numbers rather than the rule.

What I would actually do

Fix what the crawlers see first, because it is the cheapest work and the most common cause of failure. Then make your structure liftable: a direct answer near the top of each section, real headings, schema that states plainly who you are and what you sell. Then go and make your facts consistent everywhere else on the web, which is the part nobody enjoys and the part that compounds.

Then decide the licensing question deliberately, per operator, and write down why.

And set expectations on the way in: bounded gains, domain-dependent tactics, a 1% click-through on the citation itself, and measurement you have to build yourself. Everything in that sentence is from published research or the platforms’ own documentation, and every one of them is a thing a GEO pitch tends to leave out.

Read the full version

This is the condensed version. The full article covers the SEO-versus-GEO overlap in more detail, the schema types worth implementing first, and the questions clients actually ask — including whether anyone can guarantee a citation.

Generative Engine Optimization (GEO) Explained

Sources

The definition and the measured gains come from Aggarwal et al., GEO: Generative Engine Optimization (arXiv:2311.09735). The click-through figures are from the Pew Research Center panel study of 68,879 searches. The platform behaviour is documented by Google directly, in AI features and your website and Google crawlers and user-triggered fetchers, and the structured-data vocabulary is Schema.org. The read on what all this means for a smaller site is mine.

If you have been running a fixed prompt list against these engines for any length of time, I would like to know how stable your results have been — particularly whether the Google-Extended decision changed anything measurable for you either way.

Leave a Reply