SEO best practices can lose traffic: what controlled testing on enterprise sites reveals

Cracked cliff edge with a glowing checklist arrow falling during SEO A/B testing

Written by

in

Controlled SEO A/B testing on large ecommerce and travel sites has shown that many long-standing best-practice recommendations can swing traffic in either direction, including negative ones. The lesson from years of split tests is straightforward: experience should generate hypotheses, not certainty.

That is the working premise behind the SearchPilot platform, which was spun out of the search agency Distilled in 2019 and now focuses exclusively on running controlled SEO experiments for enterprise teams.

Why SEO best practice is not the same as SEO proof

Enterprise SEO work is full of high-confidence recommendations that behave differently once they are tested. A technical audit finds missing markup. A title tag looks under-optimised. A template could include more keywords. The recommendation feels sensible, so it goes onto the roadmap. The missing step is evidence.

Once a change is tested at page-group level, some recommendations win, some lose, and some do nothing. The outcome often depends on the site, template, vertical, competitors, SERP layout, user behaviour, and timing. The point of a testing programme is not to make SEO expertise redundant. It is to turn that expertise into testable hypotheses with a clean mechanism and a measurable expectation.

How SearchPilot came out of Distilled

SearchPilot did not start as a standalone software company. It began as internal tooling inside Distilled, a global search agency that grew to roughly 60 employees with offices in London, Seattle, and New York, plus the SearchLove conference series, DistilledU training, and a content publishing operation.

Around the early 2010s the agency hit its first real growth wall, with revenue falling year over year for the first time. Out of that pressure came a sharper R&D push. The team, including the current CTO, built what became known as DistilledODN, a software-enabled capability for SEO consulting.

In 2019, during acquisition conversations with Brainlabs, the agency and software sides split. Brainlabs took the agency business. The software team spun out into a standalone company. The spin-out happened shortly before the pandemic hit in 2020, which forced an early question: what does the new company do better than anyone else? The answer became controlled SEO experimentation for large websites.

Why focus became the real business lesson

Running an agency and running a software company are not the same job. Distilled had juggled consulting, conferences, training, publishing, R&D, international expansion, and software at once, which was both exciting and chaotic. The spin-out narrowed the mission to a single core capability: helping enterprise teams run SEO tests at scale.

That focus shapes everything downstream. The platform is built for large ecommerce, retail, marketplace, and travel sites with enough pages, enough traffic, and enough commercial value to make controlled testing worthwhile. The same lesson applies to a testing programme: prioritising a small number of strong hypotheses matters more than generating a long backlog, and building clean variants with enough traffic sensitivity gives the business something it can act on.

How SEO A/B testing actually works

SEO A/B testing is structurally different from traditional CRO testing. In CRO, users are split between a control and a variant experience, and behaviour is measured. In SEO, the search engine crawler is part of the system, so users cannot be bucketed without creating cloaking risk or breaking the test.

The fix is to split pages, not users. A statistically balanced group of pages is chosen as the control. A comparable group is chosen as the variant. The change is applied server-side on variant pages only, so it is visible to users, to Googlebot, and to LLM-related crawlers alike. There is no separate crawler-only version. The page experience is identical for every visitor to that page.

That page-level approach is why large ecommerce and travel catalogues, with many pages receiving meaningful organic traffic, are such a strong fit for the model.

What makes a site testable

Two ingredients matter: traffic and pages. A rough rule of thumb is around 30,000 organic sessions per month, or about 1,000 per day, distributed across a useful number of pages. Two pages are not enough. Dozens can sometimes work. Hundreds or thousands are better.

Conversions matter to the business, but they are often too noisy to be the primary metric for a test. Revenue is affected by discounts, competitor pricing, promotions, seasonality, and macro conditions. Organic traffic is usually a cleaner primary metric for detecting the effect of a page change. The business can then translate the expected traffic lift into revenue using its own commercial models.

What makes a strong SEO hypothesis

An idea is not a hypothesis. Adding FAQ content to category pages is an idea. The hypothesis is the mechanism and the expected effect: adding concise FAQ content will help category pages rank for additional long-tail queries and improve relevance for existing ones, increasing organic sessions to the tested page group.

Weak hypotheses produce weak tests. Some failures are not because the idea was wrong but because the change was too small to produce a measurable result, or because the expected mechanism was unclear. There is even a working concept of an underpowered-idea detector: a way to flag test ideas that are unlikely to move the needle before engineering time is spent on them.

Alt attributes are a useful example. Adding alt attributes to images is good practice for accessibility, usability, compliance, and sometimes image search. Controlled tests have not produced evidence that changing alt attributes moves standard organic SEO traffic in either direction. That does not make alt attributes unimportant. It makes them a weak candidate for a traffic-focused SEO A/B test.

Why test backlogs grow faster than teams can run them

Large SEO teams rarely run out of test ideas. Ideas come from in-house SEOs, agencies, product teams, content teams, merchandising teams, technical audits, competitor analysis, internal politics, and old recommendations that have been waiting for engineering resource. The challenge is not generation. It is prioritisation.

Helpers usually judge ideas on two axes: how likely they are to move the needle, and how easy they are to build. The goal is the overlap, meaning changes that are both commercially meaningful and practical to implement.

The larger blocker is often organisational. Some teams cannot test above-the-fold on product pages because product owns that space. Some cannot change templates because engineering capacity is scarce. Some cannot touch certain content modules because brand or merchandising controls them. When test velocity is low or win rates disappoint, the problem is not always the quality of the SEO ideas. It may be that the team is only allowed to test low-impact parts of the page. Enterprise SEO experimentation is as much an operating model as a tool.

Why title tags are still dangerous to change

Title tag tests are among the most likely to produce very large positive results and among the most likely to produce very large negative results. Title tags affect two things at once: rankings and click-through rate. Google rewrites some titles, but not all of them, so the wording still often contributes to the search snippet and therefore to whether the result is clicked.

Old recommendations to load title tags with target keywords now look higher-variance than many SEOs treated them at the time. Paid search analysts have long known that one advert can dramatically outperform another because the copy is better, not because it contains a different number of keywords. Organic snippets work the same way, as the test in the section above showed.

One title tag test produced an initial decline of more than 20% in organic traffic. The team iterated on the same underlying idea and found a version that produced a large positive result. The lesson was not that title tags work or do not work. It was that implementation matters.

What a breadcrumb schema test taught the team

A breadcrumb schema fix that reduced traffic is one of the more memorable test results. SEOs usually ship breadcrumb markup fixes without much hesitation. The change is technical housekeeping. The test produced a negative result.

The lesson is not that breadcrumb schema is bad. It is that search result changes can affect user behaviour in ways that are hard to predict. Structured data changes how a result appears, and a technically cleaner search result is not always a more clickable one; in the breadcrumb test, the rich result shifted user expectations in a way that reduced clicks.

What this means for SEO teams today

Controlled testing reframes the role of SEO expertise. Pattern recognition and experience still matter, but they feed a prioritised experiment queue rather than a fixed list of changes. Best-practice recommendations become hypotheses with a stated mechanism, a measured effect size, and a clear go or no-go signal.

For enterprises with enough traffic and template depth, the upside is faster learning and fewer quiet losses. For smaller sites, the same discipline still applies, even if the testing infrastructure has to be lighter: form a hypothesis, define a primary metric, build a clean control, and read the result rather than the assumption.

FAQ

What is SEO A/B testing?

SEO A/B testing is a controlled experiment where a statistically balanced group of pages is held as a control while a comparable group receives a change. The change is applied server-side so every visitor, including Googlebot and LLM crawlers, sees the same variant page. Results are compared between the groups over the test window.

Why can standard SEO best practices lose traffic?

Recommendations that look obvious, such as fixing breadcrumb schema or rewriting a title tag, can change ranking signals, snippet appearance, and click-through rate at the same time. A technically cleaner result is not always a more clickable one. Controlled tests have shown both large positive and large negative outcomes from the same type of change.

What does a site need to run controlled SEO tests?

Enough organic traffic and enough pages to form meaningful control and variant groups. A rough rule of thumb is around 30,000 organic sessions per month distributed across dozens to thousands of pages, which is why large ecommerce, retail, marketplace, and travel sites are the strongest fit.

Try the site audit tool

The SEOScanPro site audit report

The site audit tool runs a full technical audit of a site and shows the measured result behind every check. Open the site audit tool.


This article summarizes reporting from searchpilot.com.