Google On Large Product Aggregator Websites Rankings

Product pages streaming through an industrial machine representing a large product aggregator indexed poorly

Written by

in

What Happened

A large product aggregator site was crawled by Googlebot 1.17 million times over several weeks, yet only four URLs from the site appeared in Google’s index, and the homepage was not one of them. The owner reported the situation, and a Google Search Advocate responded by pointing to two distinct issues: a technical problem with redirects and the structural quality challenges that come with running a large product aggregator built on programmatic SEO.

The case is a useful window into how Google evaluates sites that generate hundreds of thousands of URLs through templates, feeds, and automation, and what site owners can do to move past the crawl-but-don’t-index pattern.

The Crawl Versus Index Gap on the Aggregator Site

Search Console data shared in the forum thread showed a sharp spike in crawl activity on the site, while the page indexing report listed just four indexed pages out of roughly 180,000 URLs. Crawling and indexing are separate stages in Google’s pipeline. Crawling is the fetch step, where Googlebot retrieves a URL. Indexing is the decision step, where Google evaluates whether the page is worth storing and serving in search results.

A high crawl count with a near-zero index count is a clear signal that Google is willing to fetch the pages but is not convinced they deserve a place in the search index. The aggregator in question had not earned that trust, and the reasons fall into the two categories the Google Search Advocate highlighted.

Issue One: A Technical Problem With Redirects

The first issue was mechanical. Every URL that Google had managed to index from the site redirected back to the homepage. That pattern suggests the site was issuing redirects at scale, either through misconfigured rules, templated redirect logic, or a content management system that pointed large groups of URLs to a single destination.

When most of the indexed URLs on a site resolve to the same page, search engines cannot build a meaningful map of unique content. Each redirect collapses what looks like a distinct page into the homepage, which removes any reason to keep the original URL in the index. Resolving this requires auditing the redirect rules, identifying which URL groups are being collapsed, and ensuring that genuinely distinct pages return a 200 status code rather than a redirect chain.

Issue Two: Programmatic SEO Quality at Scale

The second issue is harder to fix because it is about content quality, not configuration. Large product aggregators typically combine data feeds, templates, and automation to publish product pages at scale. A catalog of 180,000 URLs is not unusual for this model, and that scale is exactly where quality problems appear.

The Google Search Advocate framed the problem directly: a large product aggregator is not trivial to build or maintain well, and the result is a flood of URLs that often look auto-generated. Convincing search engines that those pages are worth indexing, and convincing users that they are worth visiting, is a high bar. Programmatic SEO is treated with caution inside Google’s guidance because the pattern is frequently associated with thin content, duplicated templates, and pages that add little beyond a rearranged product feed.

For aggregators, the practical path forward involves several checks: confirming that each product page has unique, substantive content beyond the product feed itself; verifying that internal linking treats pages as distinct destinations rather than as interchangeable nodes; and reviewing whether the template structure creates near-duplicate pages that Google’s systems would reasonably collapse.

Why Google Is Wary of Programmatic Aggregators

Programmatic SEO is not banned, and many legitimate affiliate and comparison sites rely on it. The challenge is that the same template that lets a site publish 180,000 product pages can also produce a sea of pages with interchangeable content. Google’s systems are designed to detect that pattern and to devalue pages that do not add original information.

Aggregators that succeed at indexation tend to invest in differentiating content per product or category, enriching template pages with editorial content, user reviews, or unique data, and pruning URL groups that cannot meet a quality bar. Crawl budget becomes a practical constraint as well: when Googlebot spends time on low-value pages, fewer resources remain for the pages that actually deserve indexing.

What Site Owners Should Do

For a site caught in the crawl-but-don’t-index pattern, the diagnostic order matters. Start with the technical layer: confirm that indexed URLs return 200 status codes, that redirect chains are clean, and that the sitemap accurately reflects the URLs the site wants indexed. Then move to the content layer: review a sample of the auto-generated pages and ask whether each one offers something a search user could not get from a dozen other pages on the same site.

For sites built on programmatic SEO, technical and on-page SEO audits are the fastest way to surface redirect loops, duplicate templates, and thin content at scale. SEOScanPro runs a full technical audit of a site and shows the measured result behind every check, including indexability and redirect issues.

FAQ

Why would Google crawl a site 1.17 million times but index only 4 pages?

Googlebot and the index are separate stages. A high crawl count with almost no indexed pages usually means Google is fetching the URLs but rejecting them at the indexing step, often because of technical problems such as widespread redirects to the homepage, thin or auto-generated content, or quality issues associated with large-scale programmatic SEO.

What is a large product aggregator in Google’s view?

A large product aggregator is a site that combines data feeds, templates, and automation to publish product or listing pages at scale, often in the tens or hundreds of thousands. Google treats these sites with caution because the same template that enables scale frequently produces pages with little unique value.

Does Google penalize programmatic SEO?

Programmatic SEO is not banned, but Google has consistently signaled that auto-generated, low-differentiator pages are hard to index and rank. Aggregators that succeed invest in per-page unique content, clean internal linking, and pruning of low-value URL groups.

Try the site audit tool

The SEOScanPro site audit report

The site audit tool runs a full technical audit of a site and shows the measured result behind every check. Open the site audit tool.


This article summarizes reporting from seroundtable.com.