Skip to content
Google Search

Google's Spam Policy Stopped Asking How the Page Was Made

Sighting.aiAugust 27, 202611 min read

Google's spam policies page defines the practice most arguments about AI-written content are actually about, in one sentence: "Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users." The sentence directly after it is where the answer lives. "This abusive practice is typically focused on creating large amounts of unoriginal content that provides little to no value to users, no matter how it's created."

No matter how it's created. Read on 27 August 2026, with a last-updated stamp of 15 May 2026, that page contains no rule about whether a person or a model produced the words.

People get this wrong in both directions, and the same clause corrects both. A writer who drafts with a model and worries the page is now tainted is reading a rule that is not there. A publisher who takes "AI is allowed" as clearance to generate four thousand near-identical pages is reading the same rule with the second half deleted.

Does Google penalise a page because AI wrote it?

No. The position has been stated in a first-party document since February 2023 and never withdrawn. From the FAQ attached to "Google Search's guidance about AI-generated content", posted 8 February 2023 by Danny Sullivan and Chris Nelson: "Appropriate use of AI or automation is not against our guidelines. This means that it is not used to generate content primarily to manipulate search rankings, which is against our spam policies."

Note the shape of that answer. It is conditional, and the condition is purpose, not process. The same post makes the point from the other end: "Using AI doesn't give content any special gains. It's just content. If it is useful, helpful, original, and satisfies aspects of E-E-A-T, it might do well in Search. If it doesn't, it might not."

It also says why production method was rejected as an axis: "about 10 years ago, there were understandable concerns about a rise in mass-produced yet human-generated content. No one would have thought it reasonable for us to declare a ban on all human-generated content in response." The tool changed; what made those pages bad did not.

What does Google actually penalise?

Purpose and value, assessed across pages rather than within one. Three of the named policies catch what people usually mean by "AI spam", and it is worth seeing exactly what each makes the test.

| Policy | The violation, as defined | What is explicitly not the test | | --- | --- | --- | | Scaled content abuse | Many pages generated "for the primary purpose of manipulating search rankings and not helping users" | How it was made — "no matter how it's created" | | Site reputation abuse | Third-party content published "mainly because of that host's already-established ranking signals" | Using third-party content at all — "Having third-party content alone isn't a violation" | | Expired domain abuse | An expired domain "purchased and repurposed primarily to manipulate search rankings by hosting content that provides little to no value" | The domain's age — "It's fine to use an old domain name for a new, original site that's designed to serve people first" |

That last quotation is from the March 2024 announcement; the rest are from the policy page. What the three share is a page whose reason for existing is a ranking signal it did not earn. A model can execute any of them faster. None is defined by one.

Inside the scaled content abuse section, exactly one listed example names AI: "Using generative AI tools or other similar tools to generate many pages without adding value for users". Both halves are load-bearing. The examples beside it — scraping feeds, stitching content from different pages, spreading output across sites to hide its scale — describe the same behaviour with no model in the sentence at all.

Why did Google's policy stop naming automation?

Because Google decided production method did not separate the cases well enough to enforce on. Scaled content abuse is not an addition to the rulebook; it is a replacement. The 5 March 2024 post announcing it says the new policy "builds on our previous spam policy about automatically-generated content", and applies "no matter whether content is produced through automation, human efforts, or some combination of human and automated processes."

Read it as an admission. The old policy required Google to establish how a page was made; the new one does not, and the same post gives the reason: it was expanded "to account for more sophisticated scaled content creation methods where it isn't always clear whether low quality content was created purely through automation." A criterion you cannot reliably evaluate is a criterion you stop writing into policy — and dropping it widened enforcement, not narrowed it.

The same edit happened again eight months later. Clarifying site reputation abuse on 19 November 2024, Google deleted the process defence the original wording allowed — "no amount of first-party involvement alters the fundamental third-party nature of the content" — and added a line that cuts both ways at once: "we don't simply take a site's claims about how the content was produced at face value". You cannot be convicted for using a model, and you cannot be acquitted by describing your editorial workflow.

Do I need to do anything special to appear in AI Overviews or AI Mode?

No, and Google states it flatly. From "AI features and your website", stamped 10 December 2025: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." The same page closes off the specific thing people keep building: "You don't need to create new machine readable files, AI text files, or markup to appear in these features."

The one requirement is the ordinary one: a page "must be indexed and eligible to be shown in Google Search with a snippet". Note which direction the AI traffic runs in the policy, though. The spam page's opening definition now includes "attempting to manipulate generative AI responses in Google Search". Gaming the generated answer is spam; appearing in it is not a separate discipline.

Do I have to disclose that I used AI?

For Search, no requirement — only a recommendation conditioned on what a reader would expect. For Merchant Center product data, a requirement, and a machine-readable one.

The Search guidance, on the "Creating helpful, reliable, people-first content" page, hands the judgement back to you: "AI or automation disclosures are useful for content where someone might think 'How was this created?' Consider adding these when it would be reasonably expected."

Merchant Center is the exception, and the only place in these documents where an AI label is mandatory. "All images created using generative AI must contain meta data indicating that the image was AI-generated by using the IPTC DigitalSourceType TrainedAlgorithmicMedia metadata tag." Generated titles and descriptions must be declared in the `structured_title` and `structured_description` attributes, carrying a `digital_source_type` of `trained_algorithmic_media`.

That is a self-declaration, not a detection. Google asks the merchant to say so; it does not offer to work it out. The vocabulary is IPTC's DigitalSourceType — the same vocabulary C2PA Content Credentials carry, and the same terms our image lane reads in `inference/app/engines/rungs/c2pa_rung.py`, where `trainedAlgorithmicMedia` alone means wholly generated. Google names the failure mode in its own instruction: "Don't remove embedded metadata tags such as the IPTC DigitalSourceType property from images created using generative AI tools". A declaration lasts exactly as long as nobody strips it, which is why a missing tag is absence of evidence and never evidence of absence.

How do Google's own quality raters handle AI-written pages?

They are instructed not to decide the question. The General Guidelines dated 11 September 2025, section 4.6.6: "the use of Generative AI tools alone does not determine the level of effort or Page Quality rating." Section 4.6.5, on scaled content abuse: "Even if you are unsure of the method of creation, e.g. whether or not the page is created using generative AI tools, you should still use the Lowest rating when you strongly suspect scaled content abuse after looking at several pages on the website."

Two things sit in that instruction. The rating survives not knowing how the page was made. And the unit of judgement is stated out loud — "after looking at several pages on the website". One document is not the evidence.

The nearest thing to an AI-spotting instruction in the document's 182 pages is a check for chatbot preamble somebody forgot to delete: raters are told paraphrased content is likely to have "words like 'As an AI language model'". Google supplies its own caveat on how much this matters: its guidance on using generative AI content, stamped 10 December 2025, says the rater guidelines "are not a guide to ranking first in Google" and that raters' "ratings don't directly influence ranking".

Can an AI detector tell you whether Google will penalise your page?

No, and the reason is structural rather than a question of how good the detector is. Every criterion in the three policies above is a property of something a per-document classifier never sees: how many near-identical pages sit on the same domain, why they were published, whose ranking signals they borrow. Our text engine reads a passage of at least a hundred words and returns one of two verdicts, `ai` or `uncertain`. Purpose is not among its inputs, and neither is the site.

Nor is there a Google signal to compare against. In every Google document cited above — read in full for this post — no AI-detection signal is named anywhere. The closest is the 2023 FAQ on SpamBrain: systems that "analyze patterns and signals to help us identify spam content, however it is produced." The object of that sentence is spam. Its last four words say what is not being identified.

Our own instrument is no substitute, and its numbers say why. On the corpus in `evidence/DETECTION-STATUS.md` it identifies 12 of 18 AI-written passages — 18 documents from a single generator, the weakest figure we publish — and wrongly accuses 1 of 496 human documents, all pre-2022, with no non-native English slice at all. The text lane also cannot return a human verdict at any score, by design, so a clean scan of your draft is not evidence a person wrote it, let alone evidence about a ranking system nobody outside Google can query.

The honest use for a detector here is narrower than the one people want. It can tell you a supplier handed you machine-written copy they invoiced as original. It cannot tell you what Google will do with the page.

What is worth checking before you publish?

Google publishes its own self-assessment questions, and they are closer to a usable checklist than anything a scan produces. From the creating-helpful-content page, warning signs for search-engine-first content:

  • "Are you using extensive automation to produce content on many topics?"
  • "Are you mainly summarizing what others have to say without adding much value?"
  • "Does your content leave readers feeling like they need to search again to get better information from other sources?"
  • "Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"

Look at what the first one measures: not whether automation was used, but how much of it and across how many topics. Google's generative-AI optimisation guide compresses the same test into a phrase, advising against publishing what "could easily be produced by a generative AI model" — not a prohibition on models but a test of whether the page needed you. Its own example pair contrasts "7 Tips for First-Time Homebuyers" with "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line". Only one of those requires having been somewhere.

Every question on that list can be answered by the person publishing the page. None can be answered by looking at the page's text.

Google SearchContent PolicyDetection Limits