Generative engine optimisation has quickly become the question on every ecommerce leader’s lips, and prompt tracking tools have not kept up with the answer. The dashboards show brand mentions, share of voice, and AI visibility scores across ChatGPT, Perplexity, Gemini, AI Overviews, and AI Mode, yet they cannot tell a team whether a website change actually made the business better. That gap is why more ecommerce teams are moving from passive monitoring to controlled experimentation, testing GEO changes the same way the industry learned to test SEO.
Why prompt tracking falls short
Prompt tracking is useful for a narrow set of jobs. It can show whether a brand appears in a sample of AI answers for a sample of prompts. It can flag early warning signs. It can describe how a brand is framed in certain contexts. It can help debug a specific visibility problem. It cannot, however, prove that a change to a website caused a measurable lift in business outcomes.
The underlying reason is structural. Prompts are personal, often long, and shaped by prior conversation. The same user may follow up with a question that changes the entire context. Different users get different answers. The universe of possible prompts is effectively infinite, which is why some practitioners describe this as a “search volume one” world: there is one prompt per query, and almost no repeat behaviour to anchor a measurement against. A dashboard sampling a small slice of that reality cannot substitute for controlled testing.
Why ecommerce is the right testing ground
ChatGPT and its peers will not ship the shoes, manufacture the product, or fulfil the order. They may help a shopper research, compare, and decide, but the retailer, marketplace, travel site, or brand still owns the transaction. That gives ecommerce teams a structural advantage in the AI discovery era. Discovery may look more like a mix of search engine and chatbot, and the journey may be more conversational, but the commercial question is familiar: will the customer buy from you?
Large ecommerce sites also have something most media and publishing sites do not: scalable templates. Product detail pages, product listing pages, category templates, internal search pages, faceted pages, buying guides, FAQs, and review blocks are all repeatable surfaces. Repeatable surfaces can be split into variant and control groups, changed, and measured. That makes ecommerce a practical place to run GEO experiments at scale.
How LLMs find fresh product information
AI systems draw on two broad information sources. The first is training data, the language, entities, relationships, and brand associations the model absorbed during training. Training data is largely fixed until the next training run, which makes it a slow lever. A team cannot walk into leadership with a strategy that amounts to waiting for the next model and hoping it likes the brand more.
The second source is retrieval, often described as retrieval-augmented generation or RAG. When a user asks a question, the model can pull in fresh information during the interaction, opening fifteen or fifty tabs in the background, reading across the web, and synthesising an answer. For ecommerce, this matters because many of the questions shoppers ask depend on data the model cannot have memorised: current stock, today’s price, active discounts, latest reviews, delivery options, product availability, and new launches. Product recommendations need freshness, and that freshness typically comes from retrieving live pages, feeds, or search results.
What fan-out queries change about optimisation
Behind a single user prompt there is often a swarm of background searches. The model may break the task into fan-out queries for product comparisons, reviews, pricing, availability, best options for a use case, brand reputation, delivery details, and other supporting information. The user sees one answer. Behind that answer there may have been many searches.
This shifts the optimisation question away from keyword-first thinking. In traditional search, a team asks which keyword it targets, where it ranks, and what the result looks like. In AI discovery, the hidden fan-out queries are where retrieval actually happens, and teams usually cannot see the full list. They may get clues, but they cannot treat the process as a clean keyword list. That is one reason controlled testing becomes more important than ever. A team can make a change, then measure whether that change moved LLM referrals, Google organic traffic, or the net business outcome, without needing perfect visibility into every hidden query.
How to write a testable GEO hypothesis
Traditional SEO hypotheses tend to work through one of three mechanisms: targeting new keywords, improving rankings for existing keywords, or changing the search result appearance to lift click-through. GEO has analogues, but the language changes. A GEO hypothesis might aim to target new fan-out queries, improve visibility for existing fan-out queries, influence the summary an LLM returns, or make a page, product, or brand easier for the model to recommend.
The fourth mechanism feels close to conversion rate optimisation, except the converter is partly the machine. The question becomes whether the page has given the model what it needs to confidently recommend the product: product detail, comparison language, reviews, freshness, structured data, key features, FAQs, delivery information, stock signals, or buying guidance. A weak GEO hypothesis says “this might help AI visibility.” A stronger hypothesis reads: adding clearer product suitability information to product detail pages may help models retrieve and recommend these products for more specific fan-out queries while also improving confidence in the AI-generated summary. That version is testable, and that is the point.
Where GEO testing happens on a large site
GEO testing on large ecommerce sites happens on the same scalable surfaces SEO testing already uses. Product detail pages, product listing pages, category templates, buying guide modules, comparison content, FAQs, review summaries, key feature summaries, internal linking modules, structured data, freshness indicators, product feed-aligned content, and availability and delivery information can all be split into variant and control groups. The mechanics look familiar to anyone who has run an SEO A/B test. The difference is the journey being measured.
In traditional search, a user might open several tabs, compare sources, read reviews, check products, and arrive at the site later. A lot of that research was visible across a trail of searches and visits. In AI discovery, more of that research happens inside the conversation. The model reads, compares, summarises, and narrows options before the user arrives. The site may only see the final click, which can be more valuable but is harder to interpret.
Why GEO and SEO can disagree
Many GEO changes are also plausible SEO changes. More useful content, better structure, fresher product information, clearer summaries, stronger internal links, and better structured data can all carry SEO hypotheses. That does not mean every GEO-positive change is SEO-positive. Most practical GEO work today still reaches AI systems through search-related retrieval, which creates overlap with SEO. Overlap is not sameness.
A change can help an LLM understand and summarise a page while hurting Google organic performance. A change can make a page richer for AI retrieval while making it bloated, duplicative, or less effective in traditional search. This is where single-channel measurement becomes dangerous. A team that looks only at LLM referrals, sees a positive result, and rolls the change out could quietly lose more Google organic traffic than it gained. That is the bigger risk with guessing what works in GEO: a visible win in one channel can hide a larger loss elsewhere.
What the Omio GEO test showed
SearchPilot partnered with Omio to share early GEO testing results with the wider industry, and the lessons go beyond confirming that GEO can be tested. The clearest finding was that GEO and SEO do not always move together. In one test, adding brand USPs increased LLM-driven traffic by roughly 18 percent. In another, adding structured key takeaways performed positively for LLM-driven traffic, but would likely have hurt Google organic sessions by about 6.5 percent. Omio chose not to roll that change out and developed follow-up iterations instead.
That is the practical value of testing. The AI result looked positive in isolation. The business result was not. Without measuring Google organic performance at the same time, the team could have shipped a net-negative change across a large site. This is also the strongest argument against treating GEO as a checklist. A tactic can be directionally plausible and still wrong for a specific site, page type, or business goal.
What prompt tracking can and cannot tell you
Prompt tracking has a role, and it is worth being precise about what that role is. The best comparison is rank tracking in traditional SEO. Rank tracking is useful for debugging. It can diagnose problems, show movement, surface early warning signs, and help teams understand where visibility may be shifting. It is not the same as business impact, and prompt tracking has the same limitation with extra complications.
Teams do not know the full set of prompts users are typing. Many prompts are unique. Answers are personalised. The model may use memory, previous conversations, location, and other context. A brand may appear for a prompt in one run and not another. A dashboard can only sample a small portion of reality. It can help teams form hypotheses, notice issues, and explain some AI visibility patterns. It should not be the main evidence that a GEO programme is working.
How to report GEO results to leadership
Executive attention on AI search has arrived faster than the industry’s ability to answer it. Leadership wants to know whether the team is ready, whether it is doing the right things, and how it can move faster. Those are reasonable questions, and they deserve better answers than another dashboard of AI visibility scores.
The answer that travels best is a measured one. Frame each GEO change as a hypothesis, test it against a control group, and report the joint outcome across LLM referrals and Google organic. Show the net business effect, not the channel effect. Be explicit about what was rolled out, what was held back, and what the next iteration will test. That gives leadership a programme it can govern rather than a checklist it has to take on faith.
FAQ
Why is prompt tracking not enough to measure GEO?
Prompt tracking tools sample a small slice of how AI systems actually answer queries, and prompts are personal, long, and shaped by prior conversation. The full universe of prompts is effectively infinite, so a dashboard can show whether a brand appeared in some answers without proving that a website change caused a business outcome.
Can GEO changes hurt Google organic traffic?
Yes. A change that makes a page richer for AI retrieval can also make it bloated, duplicative, or less effective in traditional search. In SearchPilot’s work with Omio, one GEO-positive change was projected to lift LLM-driven traffic while costing roughly 6.5 percent of Google organic sessions, so the team chose not to roll it out.
What surfaces on an ecommerce site can be GEO tested?
Repeatable templates can be split into variant and control groups and tested the same way SEO A/B tests are run. That includes product detail pages, product listing pages, category templates, buying guides, comparison content, FAQs, review and feature summaries, internal linking modules, structured data, freshness indicators, product feed-aligned content, and availability and delivery information.
This article summarizes reporting from searchpilot.com.