You can generate five thousand pages this afternoon. A template, a spreadsheet, a model, and a loop.
That is the problem, not the achievement. Google can tell, the March 2026 core update was built to tell, and sites that shipped exactly that lost 60 to 90 percent of their rankings almost overnight.
Programmatic SEO still works. Comparison tools with live pricing, directories with verified listings, and guides built on real inventory all continue to rank. The difference between those and the sites that got wiped is not the technique. It is a bar you can count, and most 2026 guides never state it.
What Google actually banned
Read the policy wording carefully, because the thing people assume it says is not what it says.
Google defines scaled content abuse as "when many pages are generated for the primary purpose of manipulating Search rankings and not helping users".
There is no mention of AI. No mention of templates. No mention of automation. The policy is about purpose and value, and it applies identically to a thousand pages written by a thousand freelancers.
That cuts both ways and the second half is the useful half. Generating pages at scale is not itself a violation. Generating pages that do not help anyone is, whether a model or a person produced them.
So the question to ask of your pipeline is never "will Google know this was automated". It will, and it does not care. The question is whether each page does something for the person who lands on it.
The policy does not mention AI, templates or automation. It is about whether the page helps anyone. That is a much harder bar and a much fairer one.
The bar you can count
Here is the number the generic guides leave out.
Three or more unique data points per page, enforced at the template level.
Unique means genuinely different between one page and the next, and genuinely useful. A city name substituted into four sentences is one data point wearing four hats. A price, a capacity, a verified opening time and a photo are four.
The test that keeps you honest: strip out everything the template shares across all pages. Is what remains worth a page on its own? If the residue is a name and a postcode, you have a database row, not a page.
Sites that kept their rankings through March 2026 were built on real structured data: directories with verified listings, comparison tools reading live pricing, guides backed by actual inventory. Sites that lost 50 to 80 percent had hundreds or thousands of pages published without editorial oversight, and the near-identical ones lost 60 to 90 percent.
The pipeline
Five stages. The middle one is the one everybody skips.
| Stage | What happens | Fails when |
|---|---|---|
| 1. Data | Assemble the source. Real, current, and yours to publish | You are scraping someone else''s facts and adding nothing |
| 2. Template | One layout, with slots for the variable data and a fixed spine | The variable slots are all adjectives rather than facts |
| 3. Enrichment | A human-written intro and outro per cluster, not per page | You skip it because it does not scale, which is the point |
| 4. Review gate | Automated checks, then a sample read by a person | The gate exists but nothing is ever rejected by it |
| 5. Publish and prune | Ship, measure, then noindex or delete what underperforms | You publish everything and never look again |
Stage three deserves the attention. A human-curated intro and outro at the cluster level is the compromise that actually works: you write one genuinely good introduction for "accountants in Delhi NCR" rather than one for each of 400 individual accountants. It scales to the number of clusters, not the number of pages, which is usually a hundredfold difference.
Stage five is the one people find hardest. Pages that get no impressions after three months are not neutral, they are dilution. Noindex them.
The checks to run before a page goes live
Automate these as a gate in the build, so a page cannot ship without passing. This is the part that turns a policy into a pipeline.
- Unique data point count. Reject anything under three. Count fields that differ from the cluster median, not fields that merely exist.
- Near-duplicate detection against siblings. Compare each page to the others in its cluster. If body similarity is above roughly 90 percent, the template is doing all the work.
- Minimum content that is not boilerplate. Measure the page minus the shared spine. A page whose unique portion is two sentences is thin regardless of total word count.
- A named user task. Each template should answer one specific question a person typed. Write it down. If nobody on the team can state it, the template should not exist.
- An internal link to a hand-written pillar. Every generated page points up to something a person actually wrote. This is both a ranking signal and an honesty check: if there is no pillar worth linking to, the cluster is probably not a real topic.
- Crawl control. Robots rules and noindex on thin variants, so the ones that do not clear the bar never enter the index at all. Cheaper than removing them later.
- Real HTML in the response. If your pages render client-side, a crawler that does not run JavaScript sees an empty shell. Check the served HTML, not the browser.
That last one catches more sites than it should. It is invisible until you curl the page.
Pages that get no impressions after three months are not neutral. They are dilution, and the fix is noindex, not patience.
What survives
The pattern across the survivors is consistent, and it is not subtle.
Live data beats generated prose. A comparison page reading current prices is useful every time it loads. A page describing those prices in paragraphs is stale the day it ships.
A real task beats a keyword. "Accountants in Delhi NCR who file GST returns" is a task. "Best accountants Delhi" is a keyword with a page attached.
Fewer, denser pages beat more, thinner ones. Four hundred pages with real data outrank four thousand with a name swapped in, and they cost less to maintain.
Somebody owns the cluster. The surviving sites have a person who reads a sample every month and prunes. The wiped ones had a cron job.
Common mistakes
Assuming AI generation is the violation. It is not. Google''s policy names purpose, not method. Hand-written thin pages die the same way.
Counting words instead of facts. A thousand words of template boilerplate is thinner than two hundred words of real data. Measure the unique portion.
Building the review gate and never rejecting anything. A gate that passes everything is documentation, not a control. If nothing has ever been rejected, your threshold is wrong.
Writing the intro per page. It does not scale and you will stop doing it in week two. Write it per cluster, properly, once.
Never pruning. Publishing is the easy half. The sites that recovered from March 2026 did it by fixing or deindexing thin pages, not by adding more.
Skipping the pillar. Generated pages that link only to each other look exactly like what they are. Every cluster needs something hand-written at the top of it.
Key takeaways
- Google defines scaled content abuse by purpose, not method. AI, templates and automation are not mentioned in the policy.
- The countable bar is three or more unique data points per page, enforced in the template rather than checked by eye.
- Strip the shared template. If what remains is not worth a page, it is a database row.
- Sites with near-identical templated pages lost 60 to 90 percent of rankings in March 2026. Sites built on verified, live data kept theirs.
- Write the intro and outro per cluster, not per page. It scales to the number of clusters, which is the whole point.
- Gate on unique data count, sibling similarity above 90 percent, non-boilerplate length, a named user task, and an internal link to a hand-written pillar.
- Prune. Pages with no impressions after three months should be noindexed, not left to dilute the site.
If you want the pipeline built and gated rather than the theory, automated content systems are one of the four things we do. And for the AI-search half of this, our post on what the crawler logs actually show covers which signals AI engines genuinely read, which is a shorter list than most guides claim.


