SEO roadmap, step 6

Content SEO

What makes a page genuinely worth creating, and useful enough to compete. Six decisions, three datasets built for this page, and my own content measured against every rule in it before yours is.

35 chapters 11 exercises, 7 assignments 10-page commodity drill 16-question quiz 8 rival guides audited 22 formats 110 videos, 7 channels 22 sources, all dated Written 23 August 2026
Jump to a chapter 35
  1. 01 Commodity content is the whole problem
  2. 02 Content SEO is not on-page SEO
  3. 03 Why this comes after architecture
  4. 04 A page needs two jobs, not one
  5. 05 96.55%, and what those pages have in common
  6. 06 The pages you should not write
  7. 07 Information gain, and the patent
  8. 08 The four things a model cannot generate
  9. 09 Research is where the page is decided
  10. 10 A receipt, and what does not count as one
  11. 11 Examples, and the invented person rule
  12. 12 Original research at very small scale
  13. 13 What the brief decides and what content decides
  14. 14 The answer goes in the first screen
  15. 15 Pogo-sticking, and what Google has actually said
  16. 16 Word count is the wrong unit
  17. 17 Write for the scanner, then for the reader
  18. 18 Not every answer is an article
  19. 19 Visual content, and the four kinds that earn a slot
  20. 20 Tables and comparisons, the underproduced format
  21. 21 Video on a page, and what it is actually for
  22. 22 Every section has to survive being ripped out
  23. 23 A heading is a contract
  24. 24 Name things specifically
  25. 25 What Google published about AI features, and what it told you to ignore
  26. 26 Freshness is a property of the query
  27. 27 Content decay has four shapes
  28. 28 Updating a page, and the three ways it goes wrong
  29. 29 Duplicate content and commodity content are different problems
  30. 30 AI-assisted content, against what Google actually published
  31. 31 Human verification, and the one rule that makes drafting safe
  32. 32 Where scaled content stops working
  33. 33 Auditing a corpus, with mine as the example
  34. 34 The mistakes I would kill first
  35. 35 Test yourself

Pick a path

Thirty-five chapters is a lot in one sitting. Choosing a path collapses the ones outside it, and you can open any of them anyway.

I have been doing this since 2010. I have owned and run more than a hundred sites, lost two of them entirely to Panda and Penguin, bought and tested more than 500 SaaS products with my own money, and taught this to over 30,000 students. I have also published eleven blog posts on this domain since June, totalling 32,551 words, and on the day I wrote this Ahrefs valued the entire domain at 8 monthly organic visits. Both of those facts are relevant and the second one is more useful to you.

Here is the mistake this page exists to stop, and it survives because it looks like diligence. You research the topic, read the top three results, cover everything they cover plus a bit more, and publish something genuinely better organised than what was there. That page is a retelling. Google has a word for it now, published in its own guidance, and the word is not one the industry uses.

"Commodity content (for example, something like '7 Tips for First-Time Homebuyers') is often based on common knowledge, which could originate from anyone, and typically adds little unique insight for readers."

That is from Google's guide to optimizing for generative AI features, read on 23 August 2026 and last updated by Google on 10 July 2026. The same paragraph gives the counterexample, and it is worth reading twice because it is not what a content brief usually asks for: "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line". Not a better guide. A different kind of thing entirely, and one that requires somebody to have waived an inspection.

So the opponent for this whole page is coverage, and I can show you what it costs at scale. Ahrefs studied roughly 14 billion pages in its own index and published the distribution on 1 December 2023: 96.55% get zero organic traffic from Google. A further 1.94% get between one and ten visits a month, and that second band is where this domain currently is. I am not describing somebody else's failure rate.

This is step six of the roadmap and it assumes the five before it. It assumes you know which surfaces your buyers use, from search everywhere optimization, that you can read a crawl report, from SEO fundamentals, that you arrive with clusters rather than a blank page, from keyword research, that you know what kind of page each cluster deserves, from search intent, and that the page has a parent and something linking to it, from site architecture. None is a hard prerequisite. The last two make everything here easier, because they hand this step a brief instead of a blank page.

The spine

The whole lesson, in six decisions

In order, and the order matters. The first two can cancel the other four, which is why they are first and why almost nothing published on this subject starts there. Each card ends with what you are actually holding when the decision is made.

  1. 1

    Deserve

    Should this page exist at all?

    The only decision on this list that can save you the whole cost of the page, and the one no content calendar has a column for. Google now has a word for the failure: commodity content, meaning something built from common knowledge that could have originated with anyone. If the honest answer to "what does this add" is "it covers the topic", the page is a retelling and the results page already has six.

    You end up holding A shorter list, with the rows you killed still visible and a reason next to each.

    01 02 03 04 05 06

  2. 2

    Know

    What do you have that a model could not have guessed?

    Named before a word is written, not discovered during the draft. A number you measured, a screenshot of something you ran, a decision you made and regretted, an export, a named source with a date. Google filed a patent on scoring exactly this and calls it information gain. If you cannot name the thing, you are about to write a summary of the current top three.

    You end up holding One named piece of evidence per page, with where it comes from and who can check it.

    07 08 09 10 11 12

  3. 3

    Satisfy

    Does the first screen answer the thing that was asked?

    The brief arrived from step four with a page type and a job. This decision is whether the job gets done immediately or gets withheld until paragraph nine. The failure is not length. It is sequence: the answer exists on the page and arrives after the reader has already gone back.

    You end up holding A first screen a stranger can read in twenty seconds and leave satisfied.

    13 14 15 16 17

  4. 4

    Show

    Is the evidence visible, or only claimed?

    Tables, screenshots, diagrams, video, a dataset, a working tool. Not decoration and not "add images for SEO". Every one of these is a claim being made checkable, and the reason this decision sits fourth is that you cannot illustrate evidence you have not got. Google states that its generative AI features can bring in relevant images and video, which is opportunity rather than obligation.

    You end up holding At least one thing on the page that would still be worth something with the prose deleted.

    18 19 20 21

  5. 5

    Structure

    Can a scanner navigate it and a machine extract it?

    Headings that say what a section answers, sections that survive being lifted out of the page, and things named specifically enough to be recognised. This is the decision most often oversold: Google published, in July 2026, that you can ignore chunking your content for AI, and the industry sold chunking anyway.

    You end up holding A heading outline a stranger could read alone and know what the page covers.

    22 23 24 25

  6. 6

    Maintain

    Does it survive contact with next year?

    Freshness, decay, updating, pruning, and the audit that tells you which. This is the decision that makes the previous five compound instead of accumulate, and it is the one I am worst at: on the day this published, none of my eleven posts had ever been updated.

    You end up holding A dated audit of your own corpus, with a verb on every row.

    26 27 28 29 30 31 32 33 34 35

Decisions one and two are the ones no content calendar has a field for
Everything published on this subject starts at three, because three is the teachable one

Those six are the model, and the order is load-bearing in a way most content frameworks are not. Decisions one and two can cancel the other four, and they are the two that no content calendar, brief template or editorial workflow has a field for. Every chapter below sits under one of the six, and the chips on each card jump to the chapters that serve it.

Notice what is not decision one. Not "pick a keyword", which was step three. Not "match the intent", which was step four. Not "structure the page", which is decision five here and is where every competing guide starts. The first question is whether the page should exist, and it is first because it is the only one that can still save you the whole cost.

Watch first, ten minutes

Google, asked whether more content is better, declining to say yes

The question this whole page answers, put to Google directly and answered at length. Notice what the answer is about: usefulness, audience, and whether anybody needed the page. Notice what it is never about, which is volume, length or coverage. Then read chapter one, which quotes the sentence that makes this answer make sense.

Published by Google Search Central. There are 110 videos from 7 channels in the library further down, filterable, and every one of those channels either owns a search surface or owns a database you might use to audit your own content.

Core Chapter 01 deserve

Commodity content is the whole problem

If you remember one thing Google has a word for the failure and it is not "thin". Commodity content is built from common knowledge that could have originated with anyone, and it is what you get when the plan was to cover the topic.

Not in this path.

A page is worth creating when it contains something a language model could not have guessed. Everything else in this lesson is downstream of that sentence, and Google has published the vocabulary for the failure it describes.

The mechanism is worth understanding rather than accepting. A model trained on the web has read the top ten results for your query, and several thousand pages very like them. It can produce the average of that set on request, fluently, in seconds, at zero marginal cost. So the average of the set is now free, and anything a reader could have obtained for free is not a reason to visit a website. That is not a moral argument about AI. It is an observation about what remains scarce.

What remains scarce is anything that required somebody to do something. A number you measured. A tool you bought and broke. A decision you made and regretted. A conversation nobody else was in. Google's own phrasing for the standard is unusually blunt: do not recycle what others have already said, or what "could easily be produced by a generative AI model".

Here is the ten-item version of that judgement, and I would rather you got three of them wrong now than on a page you spend a week writing. Three have a surface signal pointing the wrong way, and in all three the tell is the same: effort spent writing is not the same thing as effort spent finding out.

Commodity, or not?

Ten page ideas, and the tidy answer is wrong three times

Google’s own framing: commodity content is built from common knowledge that could have originated from anyone. Commit to a verdict before you open the reasoning. 3 of these have a surface signal pointing the wrong way, and the pattern in all three is the same: effort is not the same thing as gain.

  1. 01 A 4,000-word complete guide to technical SEO, covering crawling, indexing, canonicals, sitemaps and speed. Trap
    Show the verdict

    Commodity

    Length is not the issue. Every fact in it exists in Google’s own documentation and in forty other guides, so a reader who has read one of them learns nothing. This is Google’s own description: common knowledge that could have originated from anyone.

    What would rescue it Pick the one of those five where you have run something and been surprised, and write four hundred words about that instead.

  2. 02 A 300-word note on the one Cloudflare setting that made your sitemap return a 403 to Googlebot, with the header dump. Trap
    Show the verdict

    Not commodity

    You hit it, you diagnosed it, and the header dump is evidence nobody else has. Short, specific, and the only page on the internet that has this exact thing.

  3. 03 A comparison of the three keyword tools you actually pay for, with the same query run through all three and the numbers side by side.
    Show the verdict

    Not commodity

    The comparison is everywhere. The same query run through all three on a named date, with the disagreement printed, is not, because it costs three subscriptions and an afternoon.

  4. 04 A listicle of the 15 best AI writing tools, assembled from the vendors’ own feature pages.
    Show the verdict

    Commodity

    Assembled is the word doing the work. No tool was used, no criteria were chosen, and the page is a reformatting of fifteen marketing sites.

    What would rescue it Buy three of them, run one identical brief through each, and publish the three outputs. That is a different page and it takes a morning.

  5. 05 An explainer of how query fan-out works, written after reading Google’s documentation carefully. Trap
    Show the verdict

    Commodity

    Careful reading is level two on the gain ladder at best, and only if the synthesis says something no single source does. A faithful explainer of a public document is a better-written copy of a public document.

  6. 06 The same explainer, with the fan-out queries you captured from your own browser network tab on a named date, and the two that surprised you.
    Show the verdict

    Not commodity

    Now there is an observation in it. Note that Google separately tells you not to build a page per fan-out query, which is a different claim from not looking at them.

  7. 07 A page defining "content decay", 900 words, thorough and accurate.
    Show the verdict

    Commodity

    Ahrefs puts the phrase at 150 US searches a month with a traffic potential of 80. Even winning it is worth nothing, and the definition is not in dispute.

    What would rescue it Make it a chapter of something bigger. This lesson did exactly that and said so.

  8. 08 Your own site’s eleven posts measured for tables, citations, dated figures and updates, published with the four counts that came back zero.
    Show the verdict

    Not commodity

    It is a measurement, it is dated, the tool is public, and it is unflattering. The only reason nobody else has published it is that nobody else wants to.

  9. 09 A roundup of the 30 most important SEO statistics for 2026, each with a link to the source.
    Show the verdict

    Commodity

    Almost every statistic in these roundups traces back to another roundup. Try the three-hop test from chapter ten on any of them and see how many reach a method.

    What would rescue it Pick the three that matter to your argument, trace them properly, and print the two that did not survive.

  10. 10 A case study of a client recovery, with the raw Search Console export, the window, and the four things that changed at the same time.
    Show the verdict

    Not commodity

    The export is the gain, and naming the four confounders is what separates a case study from an advertisement. Most published case studies omit exactly that list.

Half of them are not commodity, and notice what those 5 have in common: every one required somebody to go and do something before the writing started, and not one required more words. That is the whole of decision two, and it is why the brief at the end of this lesson refuses to call itself finished without a gain sentence.

The pattern in the ten is that every non-commodity item required somebody to go and do something before the writing started, and not one of them required more words. The shortest item on that list, three hundred words about a Cloudflare rule, is the only page on the internet with that particular header dump on it. The longest, a four-thousand-word technical SEO guide, is one of several thousand.

One clarification before the rest of the page, because I am about to spend six chapters on gathering evidence and I do not want it read as an argument against writing well. Clarity, structure and sequence are real and they matter, and they are decisions three and five here. What they cannot do is rescue a page with nothing in it. A beautifully organised restatement of the current top three is still a restatement, and it will be indexed, crawled, technically flawless and worth nothing.

Core Chapter 02 deserve

Content SEO is not on-page SEO, and I counted the confusion

If you remember one thing Content SEO decides what information should exist. On-page SEO makes one page legible. Six of the eight guides audited here spend more sections on keyword research and on-page than on content, in a guide titled for content.

Not in this path.

Content SEO decides what information should exist and produces it. On-page SEO makes one page communicate its relevance clearly. They are two different jobs and merging them deletes the first one, because only the second is checkable by a plugin.

The reason to separate them is diagnosis rather than tidiness. "Our content is fine, we just need better titles" and "our titles are fine, we just need better content" are two genuinely different diagnoses with two different budgets, and a team that cannot say which column a decision belongs in will spend the year on the cheap one.

The split

Two jobs that share a page and share nothing else

Not definitions, decisions. When somebody says content SEO and you are not sure what they mean, ask which of these two columns the thing they want belongs in. Every argument about this subject resolves in about four seconds once that question is on the table.

Content SEO

Deciding what information should exist, and producing it.

Is there anything here worth publishing, and do I have it?

It decides

  • Whether the page should exist at all
  • What it knows that nothing else on the results page does
  • Which evidence gets gathered before the draft starts
  • What format the answer takes
  • What gets cut, merged, updated or deleted next year

When this side is wrong and the other side is right

You publish a competent, well-optimized page that restates the current top three. It gets indexed, it gets no traffic, and every audit tool reports it as healthy. This is the 96.55% failure and no amount of on-page work reaches it.

On-page SEO

Making one page communicate its relevance clearly.

Is this page as legible as it could be for the job it already does?

It decides

  • Title, description and heading wording
  • Where the target phrasing appears
  • Internal links out of this page
  • Image alt text and file naming
  • Structured data and snippet control

When this side is wrong and the other side is right

You have something worth reading and the results page cannot tell. This is real and it is cheap to fix, which is why it gets fixed first and why it gets confused for the whole subject.

The reason these merge everywhere is structural rather than lazy. On-page work is checkable by a plugin and "does this page know anything" is not, so on any shared budget the checkable half wins. The audit in the next block is what that costs: across eight published guides, on-page SEO takes more sections than any other non-content subject, and the guide with the strongest title claim in the set gives content seven sections out of thirty-two.

Both failures are real. Only one of them is expensive
Ahrefs is the only guide in the audit that gives these two their own chapters

Now the part I did not expect to be able to measure. Ahrefs' course splits SEO Content and On-Page SEO into two chapters, which is where the idea for this page came from, so I went and counted whether anybody else does. Eight guides, the seven the brief for this page named plus Google's own starter guide, fetched as raw HTML on the day this went live and sorted heading by heading.

Original data, 8 guides, 23 August 2026

Guides titled for content, audited by their own published headings

Every guide the brief for this page named, plus Google’s starter guide. Fetched as raw HTML on the day this published, every heading extracted, site chrome removed, and each remaining leaf section sorted by subject. Classification is by what a section is about rather than by how good it is, which flatters the guides with the most definition sections in them, and where a heading could sit in two buckets it went into the one that favours the guide. Every content share below is a ceiling.

Content Deciding what information should exist, or producing it.
Keyword research Finding demand, picking a target, reading intent.
On-page Placement, titles, headings, links, markup, site structure.
Measurement Analytics, tracking, auditing after publication.
Other Background, promotion or product. None of the four.
  1. 80%

    8 of 10 on content

    What it does, and what it does not

    Best thing in it The highest content share in the set, and the only one of the eight whose parent course gives SEO Content and On-Page SEO separate chapters. That split is where the idea for this page came from.

    The gap Ten sections. It is a course chapter rather than a lesson, so the depth goes to demand and page selection and the production half is seven short subsections.

  2. 63%

    10 of 16 on content

    What it does, and what it does not

    Best thing in it Three consecutive sections on extractability, and the clearest statement anywhere that you are now satisfying users, ranking systems and citation systems at once.

    The gap Its worked example is a shoe guide with 15,000 monthly visitors and 629 AI Overview citations. A good outcome, presented without the ninety attempts behind it.

  3. 22%

    7 of 32 on content

    What it does, and what it does not

    Best thing in it Honest about its own scope in the first paragraph, and the copywriting third is genuinely about writing rather than about placement.

    The gap Seven of thirty-two sections are about content. Two of its three named pillars are keyword research and site structure, which on this site are steps three and five, and it is the lowest content share in the set on the guide with the strongest title claim.

  4. 54%

    7 of 13 on content

    What it does, and what it does not

    Best thing in it Tip seven is information gain, named as such, which only one other guide in this set does at all.

    The gap On 23 August 2026 this page had a Domain Rating of 90 and an estimated 13 US organic visits a month from 7 keywords. The advice is sound and the page is not winning its own subject.

  5. 63%

    20 of 32 on content

    What it does, and what it does not

    Best thing in it Two named case studies with before and after traffic figures attached, which is more first-party evidence than most of this set carries.

    The gap The highest raw content count in the set, and rule three is doing the work: a large share of those twenty sections define what SEO content is and why it matters. Both case studies are customer outcomes on a vendor blog, so the variable being demonstrated is the product.

  6. 13%

    1 of 8 on content

    What it does, and what it does not

    Best thing in it It is about auditing content that already exists, which is the half of this subject the others barely touch, and five of its eight sections are on measurement.

    The gap Its quality metrics are uniqueness, fluff, keyword distribution and readability, all of which a tool can score and none of which answer whether the page knows anything.

  7. 50%

    8 of 16 on content

    What it does, and what it does not

    Best thing in it An accurate picture of what tooling in this category automates: briefs, topic finding, optimization scoring and repurposing.

    The gap It is a product page, and it is in this table because the brief named it and because a features page is a legitimate content format, which is chapter eighteen.

  8. 20%

    4 of 20 on content

    What it does, and what it does not

    Best thing in it It is the number one organic result for both phrasings of this subject and it is not a content SEO guide, which is the finding in chapter two rather than a criticism of the document.

    The gap Four of twenty sections touch content, and everything load-bearing Google has published about quality lives in a different document that ranks nowhere for these queries.

Scroll the table sideways

147

sections across 8 guides

65

of them about content

44%

the content share, at its most generous

36

sections on on-page SEO, the largest non-content bucket

3

guides where content is a minority of the sections

Less than half of the average guide titled for content is about content, and the largest single competitor for that space is on-page SEO. The finding I expected and did not get is worth printing too: I assumed most of these would be dominated by keyword research and on-page, and only 3 of the eight are. Ahrefs and Semrush’s blog post are both comfortably majority-content. The case for this page rests on the average and on the worst of them, not on the best. Yoast’s is the sharpest single row: 7 of 32 sections, on the guide in this set with the strongest title claim. On this site keyword research is step three, intent is step four, architecture is step five, and all three finished before this page began. Run the same count on this page and hold me to it.

Raw HTML fetched 23 August 2026, headings extracted, chrome removed, leaf sections only
Classified by subject rather than by quality, which flatters guides full of definitions

65 of 147 sections are actually about content, which is 44%. Yoast's is the sharpest single row and it is honest about itself: its three named pillars are keyword research, site structure and copywriting, and seven of its thirty-two sections are about content. Two thirds of the guide with the strongest title claim in the set is steps three and five of this roadmap.

The result I expected and did not get belongs here too, because a dataset that only ever agrees with the author is not a dataset. I assumed most of the eight would be dominated by keyword research and on-page, and only 3 of them are. Ahrefs and Semrush's blog post are both comfortably majority-content. The argument for this page rests on the average and on the worst three, not on a claim that everybody is doing it wrong.

I am not accusing anybody of padding. The drift is structural. Keyword placement can be scored, headings can be counted, and "does this page know anything" cannot, so on any shared word count the scoreable half wins. Which is also why this lesson can afford to be about content: the other buckets already have their own pages on this site and all of them shipped before this one.

One more finding from those two results pages, and it decides the URL of the page you are reading. Most guides settle their own slug in a commit message nobody sees. This one is settled here, with the numbers that argue against the choice printed first.

Two results pages, 2 pulls, 23 August 2026

The bigger query, and why this page is not on it

One subject, two phrasings, one with twice the volume of the other. Both pulled the day this published, both printed here, and the column that decides it is the last one. A page marked built for it has this query as its own top keyword. Almost none of them do.

AI Overview with 9 cited sources, then one organic result, then a four-question People Also Ask block.

# Page DR US traffic Keywords Its own top keyword Built for it
2

Google

Search Engine Optimization (SEO) Starter Guide
99 426,040 4,867 seo 498,000 No
4

Siteimprove

A creator’s guide to SEO content strategy
81 12,549 207 seo content strategy 3,200 No
5

Ahrefs

SEO Content: The Beginner’s Guide
91 2,376 140 ahrefs blog 2,500 No
6

Semrush

SEO Content: What It Is & How to Create It
92 1,007 44 content seo 1,800 No
7

Bynder

12 tips for writing SEO-optimized content in 2026
82 6,811 355 how to write seo content 1,500 No
9

Yoast

The ultimate guide to content SEO
91 151 19 content seo 1,800 No
10

seoClarity

A Complete Guide to Writing Content for SEO that Ranks
77 330 40 creating seo content 700 No
11

Michigan State University

Content Best Practices for SEO
90 561 67 seo for content 1,200 No
12

Backlinko

SEO Content: How to Create Content That Ranks
90 13 7 how to optimize content for seo 450 No

Scroll the table sideways

The AI Overview above all of it cited 9 sources

  • Google also ranking
  • Siteimprove also ranking
  • YouTube (MyCaptain) not in the visible results
  • Bynder also ranking
  • Michigan State University not in the visible results
  • Marketing Miner not in the visible results
  • SimpleTiger not in the visible results
  • Semrush also ranking
  • SEOBoost not in the visible results

Nine organic results, nine different primary targets, and not one of them is this query. Position two is a beginner guide about SEO in general with 4,867 ranking keywords, which is why Ahrefs assigns this query a traffic potential of 443,000 and a parent topic of "seo". Underneath it sits a block of about thirty Reddit threads, LinkedIn posts, Instagram reels, TikToks and YouTube videos, and the Reddit titles in it are the most honest market research on the page: "Informational content is dying", "After years of ranking, competitor takes over most keywords using hundreds of AI written articles".

Plus a block of roughly 30 Reddit threads, LinkedIn posts, Instagram reels, TikToks and YouTube videos sitting inside the top of this results page. More non-article results than article ones, on a query about writing articles.

The same AI Overview with the same 9 cited sources, then three organic results, then a four-question People Also Ask block.

# Page DR US traffic Keywords Its own top keyword Built for it
2

Google

Search Engine Optimization (SEO) Starter Guide
99 426,040 4,867 seo 498,000 No
3

Siteimprove

A creator’s guide to SEO content strategy
81 12,549 207 seo content strategy 3,200 No
4

Semrush

SEO Content: What It Is & How to Create It
92 1,007 44 content seo 1,800 Yes
6

Ahrefs

SEO Content: The Beginner’s Guide
91 2,376 140 ahrefs blog 2,500 No
8

DBS Interactive

Technical SEO vs. Content SEO: How Each Works Differently
69 155 12 content seo 1,800 Yes
9

Yoast

The ultimate guide to content SEO
91 151 19 content seo 1,800 Yes
10

Michigan State University

Content Best Practices for SEO
90 561 67 seo for content 1,200 No

Scroll the table sideways

The AI Overview above all of it cited 9 sources

  • Google also ranking
  • Siteimprove also ranking
  • YouTube (MyCaptain) not in the visible results
  • Bynder not in the visible results
  • Michigan State University also ranking
  • Marketing Miner not in the visible results
  • SimpleTiger not in the visible results
  • Semrush not in the visible results
  • SEOBoost not in the visible results

Half the volume, eleven points less difficulty, and three of the seven ranking pages have this exact phrasing as their own top keyword. One of them, DBS Interactive at position eight, is a comparison of technical SEO against content SEO on a DR 69 agency blog with twelve ranking keywords. The market’s own pages were built for this phrasing rather than the bigger one, and the AI Overview above them is identical to the other query’s, which means the two are one need with two vocabularies.

Plus a block of roughly 17 Reddit threads, LinkedIn posts, Instagram reels, TikToks and YouTube videos sitting inside the top of this results page. More non-article results than article ones, on a query about writing articles.

The decision, and what it cost

What was declined

"seo content", 3,600 US searches a month against 1,800. Twice the volume, and Ahrefs assigns it a traffic potential of 443,000 with the parent topic "seo", because the page winning it is Google’s beginner guide with 4,867 ranking keywords and a top keyword worth 498,000 a month. That traffic potential was never available to anybody.

Why /content-seo/ instead

Across both results pages, none of the 9 pages ranking on the bigger query were built for it, and 3 of 7 on the smaller one were. Semrush, Yoast and DBS Interactive all have "content seo" as their own top keyword. The market’s own pages were built for the phrasing with half the volume.

The uncomfortable part

Both queries return an identical 9-source AI Overview, and 3 of the guides audited in chapter two are absent from it: Ahrefs, Yoast, Backlinko. Two of the nine cited sources are a university web team and a YouTube video. Whatever is selecting those sources is not selecting on domain authority, and I cannot tell you what it is selecting on.

The one that should worry a publisher

Backlinko’s guide has a Domain Rating of 90 and Ahrefs put its US organic traffic at 13 visits a month from 7 keywords. It is a good guide by a strong domain on its own subject, and it is functionally invisible. Authority did not save it and neither will mine.

Ahrefs, US database, 23 August 2026. Two SERP Overview requests, one Keywords Explorer request and one Site Explorer request. Volumes and difficulty scores are Ahrefs estimates and they move; the date is on every figure so you can tell how stale it is when you read it.
Positions are Ahrefs’ own numbering, which counts result blocks rather than organic results

"seo content" carries 3,600 US searches a month and "content seo" carries 1,800. Twice the volume, and not one of the nine pages ranking for the bigger phrasing has it as its own top keyword, while three pages have the smaller one. The market's own pages were built for the phrasing with half the demand, which is the same finding step five produced on a completely different subject. Mature informational markets seem to get collected rather than targeted.

The uncomfortable half is the AI Overview. Both queries return an identical 9-source citation set, and Ahrefs, Yoast and Backlinko are absent from it while a university web team page and a YouTube video are in it. Whatever selects those nine is not selecting on domain authority, and I cannot tell you what it is selecting on. Anybody who can should be asked for their measurement.

Awareness Chapter 03 deserve

Why this comes after architecture and before on-page

If you remember one thing Architecture decided where the page lives and what links to it. This step decides what is on it. Do it earlier and you are writing pages before you know whether they have a parent.

Not in this path.

Content belongs between deciding where a page lives and deciding how it reads, because it is the last step that can still say no at a reasonable price.

Run the chain out. Keyword research gives you what people want. Search intent gives you what would satisfy them, as a page type with a count behind it. Architecture gives the page a parent and a route in. Content asks what is actually on it, and on-page then makes that legible. Every step narrows, and the cost of reversing goes up at every stage: killing a row in a spreadsheet is free, killing a draft costs a morning, and killing a published page costs a redirect and a small amount of trust.

The usual order is keyword research, then write, then optimize. That order has one structural flaw and it is the one this whole page is about: nobody ever asks whether the page should exist, because by the time anybody is looking at it, it exists. A brief that arrives at a writer has already answered the question by arriving.

And a caveat I would rather give you than have you discover. If your pages have no parent and nothing links to them, content is not your bottleneck and this lesson will not help you. Step five has the audit for that, and it found ninety-seven pages on this site with no editorial link pointing at them. A brilliant page nothing points at is a brilliant page nobody reads.

Core Chapter 04 deserve

A page needs two jobs, not one

If you remember one thing A page needs a job it does for the reader and a job it does for the business, and both have to be nameable in a sentence. If only one exists you have either a hobby or a brochure.

Not in this path.

Every page needs a job it does for the reader and a job it does for the business, and both have to be nameable in a sentence before the page is worth building.

With only the reader job, you get a hobby: useful, well made, and unconnected to anything that pays for it. With only the business job, you get a brochure: a page that exists to be found and does nothing for the person who found it. Both fail, and they fail differently enough that teams argue about which one they have.

The reader job has to be a verb. Not "learn about content decay", which is a topic wearing a job's clothes. "Work out why one of their pages is falling, before spending a day on the wrong fix." That version is checkable: you can read the page afterwards and ask whether somebody could now do it.

The business job has to be something other than traffic. Traffic is not a job, it is a measurement of one. Routes to a service page, proves a capability to somebody deciding whether to hire you, collects an email from exactly the person you want, gets cited by people who influence buyers, or supports a commercial page by being the thing that earns the link. If you cannot name which, the page has no business job and you will not miss it when it is gone.

reader job    "work out why one of their pages is falling"     ← a verb, checkable
business job  "routes to the refresh service, proves I diagnose" ← named, not "traffic"

with only the first   a hobby
with only the second  a brochure
with neither          the 96.55%

This is also the fastest filter available on an existing content plan. Take ten planned pages and write both jobs for each. The ones where you cannot is not usually two or three. In my experience it is closer to half, and every one of those is a page somebody was going to write.

Core Chapter 05 deserve

96.55%, and what those pages have in common

If you remember one thing 96.55% of roughly 14 billion pages get zero organic traffic from Google. This domain sits in the 1.94% band that gets between one and ten visits a month, and that number is on this page.

Not in this path.

Almost all published content gets nothing. Ahrefs studied roughly 14 billion pages and found 96.55% receive zero organic traffic from Google, and the number is worth sitting with before reading any advice about how to write.

The band nobody quotes is the second one. 1.94% get between one and ten monthly visits, which is the band you land in when everything works: the page is indexed, it ranks for something, a person occasionally arrives, and it is worth nothing. That failure is far more common than the total failure and much harder to see, because every report on it looks fine.

I am in that band. Here is the distribution with my own reading marked on it.

The distribution

Almost all content gets nothing, and this site is nearly all content

Ahrefs, 1 December 2023, roughly 14 billion pages from its own index. The band everybody quotes is the first one. The band that matters is the second, because it is the one you land in when the page works, gets indexed, ranks for something, and is still worth nothing.

  1. No traffic at all

    96.55%

    Zero monthly organic visits from Google.

  2. One to ten visits

    alstonantony.com, 23 August 2026

    1.94%

    Indexed, ranking for something, worth nothing. This site is here.

  3. Eleven to a hundred

    1.02%

    A page that has found a small real audience.

  4. A hundred to a thousand

    0.38%

    A page that is doing a job for a business.

  5. Over a thousand

    0.11%

    What every guide on this subject illustrates itself with.

The three URLs Ahrefs sees on this domain

Ahrefs Site Explorer, US, subdomains mode, 23 August 2026. 3 organic keywords and 8 estimated monthly organic visits across roughly 140 published pages. Not a humble brag with a hidden good number underneath. This is the number.

URL Ranking for Volume Position Visits Where it resolves today
/seo-tools/taja-ai-review/ taja ai 50 5 5 301 to zplatform.ai
/tools/seo-google-penalty-checker/ google penalty checker tool 30 8 3 301 to /seo-tools/free/seo-google-penalty-checker/
/tools/column-to-comma/ comma converter 500 25 0 301 to /seo-tools/free/column-to-comma/

Scroll the table sideways

Two of the three are free tools and the third is a review. Not one of them is a blog post, on a site with eleven of those and thirty-two thousand words in them. That is the first finding, and it is the reason chapter one is about deserving to exist rather than about writing better.

The second finding is the last column, and I did not expect it. All three of those URLs are 301 redirects. Two moved when the tools directory was restructured and one now points at a different domain entirely. So the complete organic footprint of this site, as the largest backlink index sees it, is three addresses that no longer exist, and the things visibly working here are the two that do something rather than describe something.

Bar widths are square-rooted so the small bands stay visible. Every real percentage is printed
96.55% of content gets no traffic from Google , Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.

Three URLs. Two free tools and a review, on a site with 11 blog posts and 32,551 words in them. Not one of those posts appears, and that is the finding rather than the embarrassment: the things ranking on this domain are the things that do something, and the things that describe things are invisible.

What do the 96.55% have in common? Ahrefs' own answer in that study is mostly about links and indexing, which is true and is steps two and five of this roadmap. The content answer is narrower and it is this: a page in that band could have been written by somebody who did not do anything. No measurement, no test, no purchase, no conversation, nothing that cost the author a morning before the writing started. It is not that those pages are badly written. Most of them are written perfectly well.

The useful thing about a number this large is that it reverses the default. The question is not "why would this page fail", it is "why would this page be one of the 3.45%". If you cannot answer that in a sentence, you have your answer.

Core Chapter 06 deserve

The pages you should not write

If you remember one thing Six questions, and any one of them kills a page. The cheapest content decision available is the one that stops a page existing, and no content calendar has a column for it.

Not in this path.

The cheapest content decision available is the one that stops a page existing, and no content calendar has a column for it. Six questions, and any one of them can kill a row.

A page now costs almost nothing to produce and exactly as much as ever to maintain, link, update and eventually delete. That asymmetry is new, it is the defining fact of content work in 2026, and almost every process in the industry was designed before it.

  1. Can I name what this contains that the top three do not? One sentence, written before the draft. If the sentence is "it is more comprehensive", the answer is no, because comprehensive is the definition of commodity.
  2. Does the evidence exist yet? Not "can I find some". Does the specific artefact exist, and if not, will I actually go and produce it this week? A brief whose evidence is "research from the top 10 results" has already decided to be a retelling.
  3. Is the format one I can produce? If the results page is nine calculators and I write articles, the honest outcome is not a longer article. It is either building the calculator or declining the query in writing.
  4. Does the business need it? The question no results page can answer. You can satisfy an intent perfectly and gain nothing, and the pages where that happens are usually the ones with the best volume.
  5. Do I already have something that does this job? Search your own site for the job rather than for the words. Half the answer to this question is an improvement to an existing page wearing a new URL.
  6. Will I maintain it? Every page published is a permanent liability. If you would not update it when the interface changes, you are not publishing an asset, you are publishing a future redirect.

The sixth is the one this era needs most and the one my own audit fails hardest. Eleven posts, none of them ever updated, and I published this page anyway. The honest version of question six is not "will I maintain it in principle". It is "have I maintained anything yet", and if the answer is no then the answer to question six is no too.

A note on what killing a page actually looks like, because "just publish fewer things" is not a process. The rows you kill stay in the document, struck through, with the reason next to them. Otherwise the same row comes back in four months from somebody who does not know it was already considered, and the argument gets had twice.

Who wrote this, and how to check me on it

Placed here rather than at the end, because you have just read two chapters built entirely on measurements of my own content and you are entitled to know who took them and what they cost me to publish.

Not in this path.

Alston Antony

Who is teaching this, and how to check me on it

Two receipts on this page, and both of them are about my own failures

This lesson argues that a page earns its place by containing something nobody else has, and that most published content contains nothing of the kind. It would be a poor lesson if I asked that of your content and not of mine. So here it is in Google’s four categories, then the method behind every dataset on the page, including the two that make me look worst.

Senior Digital Marketing Manager, Brainstorm Force · SEO since 2010 · Coimbatore, Tamil Nadu

01

Experience Has this person published content, at volume, and lived with the results?

Since 2010, across more than a hundred sites I have owned and run, two of which I lost entirely to Panda and Penguin. I have published 11 posts and 32,551 words on this domain since June 2026 and Ahrefs values the whole domain at 8 monthly organic visits. Every failure named in the commodity chapter is one I have personally shipped, and the audit that proves it is further up this page.

The version with the failures in it

02

Expertise Do they know the mechanism, or only the vocabulary?

Senior Digital Marketing Manager at Brainstorm Force, MSc Computer Software Engineering with Distinction from University of Greenwich, and a dissertation that was an automated SEO management system. Professional Member of BCS, The Chartered Institute for IT since 2012. Which is why 7 claims on this page are quoted from a document with the sentence intact and the date attached, and why 4 of them are graded invented rather than argued with.

Credentials, dated and checkable

03

Authoritativeness Does anyone else say so, or only them?

30,000+ students taught across six courses and free programs, 617 published videos, and a 15,000 member lifetime-deal community. The part I can prove is on the case studies with the raw exports attached. The part I will not claim: that better content caused any specific traffic number, because nobody publishing that claim has isolated the variable either.

Six case studies, with the exports

04

Trust What happens when the evidence is embarrassing?

It goes on the page at full size. On the day this published, 9 of 11 of my posts contained no table, 3 cited no external source at all, and 11 of 11 had never been updated. Google states that of the four E-E-A-T categories, trust is the most important. This is the only one of the four that costs anything to demonstrate.

What this site earns from, and how

Three datasets on this page. Here is how to go and get a different answer

Eight competing guides, audited

Raw HTML fetched on 23 August 2026, every heading extracted and sorted. 147 leaf sections across 8 guides and 65 of them about content, which is 44%. Eight fetches and an afternoon. Disagree with the classification row by row: the rule is printed above the chart, and the result that argues against my own expectation is printed under it.

The audit

This site’s own corpus

node tools/corpus.mjs, run against src/content/blog on 23 August 2026. A hundred lines of regular expressions in this repository, counting presence rather than quality. Point it at your own content folder and it will produce the same four counts about you.

Eleven posts, measured

The URL decision, published

Two SERP Overview requests and one Keywords Explorer request, US, 23 August 2026. The bigger phrasing carries twice the volume and not one page ranking for it was built for it. Search both yourself, and count the publishers rather than reading the titles.

Why not the bigger query

The experience part, as numbers you can go and check

74%

ZipWP organic click growth from a zero baseline

Brainstorm Force, Search Console · 2026

28

ZipWP keyword variations taken to #1 from nothing

Brainstorm Force, Search Console · 2026

302,037

Bing Copilot citations earned by owned properties

Bing AI Performance · 6-month windows

30,000+

Students taught across all platforms

Udemy plus direct and free courses · Aug 2026

500+

SaaS products personally bought and tested

Since 2019

Those are counts from properties I own, with the tool and the window named. They are evidence that I have done this at some scale. They are not evidence that better content caused any of them, and nobody presenting a content programme as the cause of a traffic number has isolated that variable either. That claim is graded inference on this page, in the table two chapters up, along with every other claim it leans on.

Core Chapter 07 know

Information gain, and the patent nobody reads properly

If you remember one thing Google filed a patent on scoring how much a document adds beyond what the reader has already seen. Gain is comparative: your page is judged against the five things read before it, not on its own.

Not in this path.

Information gain is how much a document adds beyond what the reader has already seen, which makes it comparative rather than absolute. Your page is not judged on its merits. It is judged against the five things read before it.

Google filed a patent on this in October 2018, published in November 2020, titled "Contextual estimation of link information gain". Its own abstract is the clearest statement of the idea anywhere: an information gain score for a document "is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user."

Two things follow, and the second one is the one people miss. First, being thorough is not the same as being additive: a page containing everything the top three contain has an information gain of approximately zero by construction. Second, the comparison set is whatever the reader has already read, which means gain is contextual. Your beginner explainer has high gain for somebody who has read nothing and none at all for somebody who has read three.

The ladder below is how I use this in practice, because "add information gain" as advice produces nothing. Most content is not at level zero. It is at level one or two, which feels like work, reads well, and is still replaceable.

The ladder

Five levels of information gain, and the line is level three

Gain is comparative, not absolute. Google’s patent scores a document by what it adds beyond documents the reader has already seen, which means your page is judged against the five things read before it rather than on its own merits. The middle two levels are where most good content lives, and they are where it stays replaceable.

  1. 0

    Restatement

    A summary of what the current top three say, reorganized.

    The test

    Could a model with no access to your business have written it? If yes, this is level zero.

    What it costs

    An hour, and it is the hour Google now names in its own guidance as commodity content.

    A real one Any of the eleven "what is SEO content" definitions on the two results pages behind this page. They agree with each other because they are copies of each other.

  2. 1

    Better arrangement

    The same information, sequenced better, written more clearly, easier to scan.

    The test

    Would a reader who had already read the top three learn a new fact? If no, you are at level one.

    What it costs

    A day. It is genuinely useful and it is the ceiling for most content operations.

    A real one Google’s own SEO Starter Guide, which is the number one organic result for both phrasings of this subject and contains almost no fact the others do not.

  3. 2

    Synthesis

    Two or more bodies of knowledge joined, where the joining is the contribution.

    The test

    Does the page say something true that no single source it cites says on its own?

    What it costs

    Two or three days, most of it reading. This is the highest level available without doing anything.

    A real one Reading Google’s spam policy against its AI optimization guide and noticing that one names scaled content abuse and the other tells you to stop targeting fan-out queries for exactly that reason.

  4. 3

    Everything below this line could have been written by somebody who did nothing

    First-hand

    Something you did, ran, bought, broke or measured, reported with the date on it.

    The test

    Is there a sentence on the page that only you could have written, and could you be wrong about it in public?

    What it costs

    The cost of doing the thing. There is no shortcut and this is the point.

    A real one The corpus audit in chapter thirty-three. Eleven posts, three tables, zero updates, produced by a tool in the repository and unflattering to me.

  5. 4

    New data

    A dataset nobody else has, produced deliberately, that other people will cite.

    The test

    Could somebody write a post whose central number comes from your page?

    What it costs

    Weeks, usually, and most of them are wasted. Do it once a year, not once a month.

    A real one The eight-guide audit in chapter two. It exists because nobody had counted the sections, and counting them took an afternoon.

The honest status of the mechanism: information gain is a patent, not a confirmed ranking system. "Contextual estimation of link information gain, US20200349181A1", assigned to Google LLC, filed 18 October 2018 and published 5 November 2020. A patent proves the idea was worth protecting. It does not prove the system shipped, and this page grades it on the record rather than documented for exactly that reason. Use the ladder because it produces better pages, not because somebody told you it is a ranking factor.

Most content operations top out at level one, and level one reads well
US20200349181A1 , read 23 August 2026

Level three is the line, and the useful thing about it is how close most briefs are to it. A page sitting at level one usually needs one measurement, one screenshot or one honest decision to cross over, and the reason it does not cross is never difficulty. It is that nobody asked the question before the draft started, and by the time there is a draft, the draft is what gets edited.

Now the status of the mechanism, stated plainly because this is where the industry overreaches. A patent is not a confirmed ranking system. It proves somebody at Google thought the idea was worth protecting in 2018. It does not prove anything shipped, and this page grades information gain "on the record" rather than "documented" for exactly that reason. Use the ladder because it produces better pages. Do not cite it as a ranking factor, and be suspicious of anybody who does.

Core Chapter 08 know

The four things a model cannot generate

If you remember one thing Four things a model cannot generate: something you measured, something you ran, something you decided and regretted, and something somebody told you privately. Everything else on your page is available to everyone.

Not in this path.

Four categories of thing are unavailable to a language model and to every competitor who has not done the work: something you measured, something you ran, something you decided and regretted, and something somebody told you privately. Everything else on your page is available to everyone.

Something you measured. A number that did not exist until you produced it. This is the highest-value category and the one people assume requires scale. It does not: the corpus audit on this page covers eleven posts and took an afternoon to build a tool for, and it is the most-quotable thing here precisely because nobody else has bothered.

Something you ran. A test, a migration, a tool you paid for, a setting you changed and broke. The evidence is a screenshot with a date in the frame and the version named. Google's own guidance names a first-hand review as its example of a unique point of view, and the operative word in that phrase is not "review".

Something you decided and regretted. The most underused category in the industry, because it costs status. A decision published with its reasoning and its outcome is unforgeable: nobody can copy it without having made it, and a model cannot invent one that is true. Every unflattering number on this page is in this category.

Something somebody told you privately. A client, a vendor rep, a support ticket, a conversation at a conference. Attributable or anonymised, but real. This is the category most likely to be legitimately unpublishable, and when it is, say the finding without the source rather than dressing it as your own analysis.

What is not on that list: opinion. An opinion is free to generate and free to hold, and a strong one is not evidence of anything. The distinction I use is whether being wrong would cost me something. "I think tables are underused" costs nothing. "Three tables across eleven posts on my own site" can be checked and would be embarrassing if I had made it up.

Practical Chapter 09 know

Research is where the page is decided

If you remember one thing The page is decided during research, not during writing. If you start the draft without a named piece of evidence, the draft will find one, and what it finds will be the median of the results page.

Not in this path.

The page is decided during research, not during writing. If you start a draft without a named piece of evidence, the draft will find one, and what it finds will be the median of the results page.

The mechanism is unglamorous. Writing is a search process: you write a sentence, you need support for it, you go and find support, and the fastest available support is whatever the top three already say. Do that for two thousand words and you have reconstructed the consensus, one honest sentence at a time. Nobody decided to write a retelling. The process wrote one.

So the order that works is: gather first, decide second, write third. Concretely, for this page, that meant reading eight competitor guides and counting their headings before writing a word, pulling two results pages and twenty keywords, and building a tool to measure my own corpus. Three artefacts existed before the first sentence, and the argument of the page came out of them rather than being illustrated by them.

Four research inputs, in the order I would spend time on them:

  1. The results page, read as evidence rather than as a checklist. Not to extract subheadings. To find what every result has in common, which tells you what the commodity floor is, and what none of them has, which is your opening.
  2. The forums and the reviews. Reddit and community threads contain the phrasing of the actual problem, and the complaints tell you which part of the standard answer does not work. Both results pages behind this lesson had large blocks of Reddit threads in them, and their titles were better market research than the ranking guides.
  3. Your own data. Search Console, your CRM, your support inbox, your sales calls. This is where level three gain comes from and it is the input almost nobody budgets time for.
  4. The primary document. If Google, or the vendor, has published on the subject, read the document rather than the coverage of it. Most of what makes this page different from the eight guides audited above is that I opened the documents they cite.

A note on using a model for this, since it is the obvious question. Research is one of the places Google's own guidance names generative AI as genuinely useful, and I use it that way: to find what has been said, build a first outline, and tell me what I have missed. It is very good at that and it is a research input, not a research output. Everything it hands you is by definition already in the commodity set.

Core Chapter 10 know

A receipt, and what does not count as one

If you remember one thing A receipt is checkable by a stranger. A number nobody can trace is decoration with a decimal point in it, and four of the claims graded on this page are exactly that.

Not in this path.

A receipt is a claim a stranger can check. That is the whole definition, and it excludes most of what circulates in this industry as data.

The failure is not dishonesty. It is a chain: somebody measures something, somebody quotes it, somebody quotes the quote, and four hops later a number is being stated as fact by people who have never seen the method. Try it on any SEO statistic you like. Click through to the source. From there, click through again. A depressing share of them do not survive three hops, and when the chain breaks you cannot state the caveat because you never saw the method.

The receipt ladder

What counts as evidence, graded by whether a stranger can check it

Not by how impressive it sounds. A number you measured yourself on eleven items beats a survey of ten thousand people that you cannot trace to a method, because the first one can be argued with and the second can only be repeated.

  1. Strong A stranger can reproduce it or read it themselves.

    Your own measurement, dated, reproducible

    A number you produced, with the tool named, the date attached, and a way for the reader to run it themselves.

    It cannot be copied by anybody who did not do the work, and being wrong about it is public. That is what makes it worth something.

    "11 posts, 3 tables, 0 updates, produced by tools/corpus.mjs on 23 August 2026." Run it against your own folder and get a different answer.

    A primary document, quoted with the sentence intact

    The vendor’s own words, linked, with the date the document says it was last updated.

    It removes you from the chain. The reader is arguing with Google rather than with your paraphrase of Google.

    "Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)"

  2. Usable Secondhand, and honest about it.

    A named third-party study, with its method and size

    Somebody else’s data, cited with the sample size, the date and a link.

    Usable and secondhand. You did not control the method, so quote the number and the caveat together or not at all.

    Ahrefs, 1 December 2023: 96.55% of roughly 14 billion pages get zero organic traffic from Google.

    A screenshot of something you actually ran

    An interface, a report, an export, with the date visible in the image.

    Hard to fake casually and immediately legible. Weaker than a number because a reader cannot recompute it.

    A Search Console export with the date range in the frame, not a stock photo of a laptop.

  3. Weak You did not check, and it shows the moment somebody does.

    A statistic with a source you did not read

    A number quoted from a roundup that quoted a roundup.

    Half of these do not survive being traced. If you cannot find the original method, you are repeating a rumour with a decimal point in it.

    "Studies show 70% of buyers…" with a link to a listicle that links to a dead PDF.

  4. Worthless A number with nothing behind it.

    A number with no source at all

    A figure stated confidently, with nothing behind it.

    It is the single most common form of evidence in this subject, and four claims on this page are graded invented for exactly this reason.

    "A TL;DR at the top can lift conversions by 33%." I have been told this. I cannot find who measured it.

Three hops. If you cannot reach a method in three clicks, the number is not evidence
The bottom tier is the most common one in this industry, including in the brief for this page

Two habits follow from that ladder and both are cheap. First, date every figure. A number without a date is an assertion, because the reader cannot tell whether it was true last week or in 2019. Nine of my eleven posts carry no dated figure at all, which is the count that surprised me most when I ran the tool. Second, quote the sentence rather than paraphrasing it. A paraphrase of a Google document puts you in the chain; the sentence takes you out of it, and the reader gets to argue with Google instead of with you.

The hardest habit is the third one: publish the numbers that argue against you. Every unflattering figure on this page was optional. Leaving them out would have produced a cleaner argument and a page nobody has any reason to believe, because a lesson about evidence whose every number flatters the author is itself a category of evidence, and the category is marketing.

Practical Chapter 11 know

Examples, and the invented person rule

If you remember one thing A real named thing beats an invented person every time, and inventing a person is worse than having no example at all because it tells the reader you had nothing.

Not in this path.

A real named thing beats an invented person every time, and inventing a person is worse than having no example, because it tells the reader you had nothing.

"Meet Sarah, a marketing manager at a mid-sized SaaS company" is a sentence that appears in an enormous amount of content and it carries exactly zero information. Sarah does what the author needs her to do. Her problem is the problem the article solves. Her objection is the objection the article answers. No reader has ever learned anything from Sarah, and every reader who has read three of these knows what she is for.

What works instead, in order of strength. A named real thing: a tool, a site, a document, a company, with the version and the date. A number from your own data. The reader addressed directly as "you", which is honest about being a generalisation. And a real person named with permission, which is the strongest and the most expensive.

The test I use is whether the example could be wrong. "Backlinko's guide to SEO content has a Domain Rating of 90 and 13 monthly US organic visits" can be checked and I would look foolish if it were false. "A well-known SEO blog struggles to rank its own content" cannot be checked and costs nothing to say. Both sentences convey the same idea and only one of them is evidence.

The exception worth naming: a hypothetical is fine when it is labelled as one and is doing a job an example cannot. The two mocked-up first screens in chapter fourteen are invented, and they are invented because the argument is about sequence and using a real page would turn it into a critique of somebody's writing. Say "here is a made-up example" and the reader can price it correctly.

Practical Chapter 12 know

Original research at very small scale

If you remember one thing You do not need fourteen billion pages. Counting the section headings of eight competing guides took an afternoon and produced the finding this whole lesson is built on.

Not in this path.

You do not need fourteen billion pages. Counting the section headings of eight competing guides took an afternoon and produced the finding this entire lesson is built on.

Original research has a reputation problem in this industry: it means a survey of a thousand marketers, or an analysis of a million SERPs, and both are out of reach for almost everybody. That reputation is doing real damage, because it stops people producing the small, specific, checkable datasets that are actually available to them and that nobody else has bothered to make.

Four kinds of very small original research, all of which I have done in the last month:

  1. Count something nobody has counted. Eight guides, seventy-nine section headings, four categories. The dataset is small, the classification rule is published so you can disagree with it, and it is the only count of its kind that exists.
  2. Measure your own thing with a tool you wrote. A hundred lines of regular expressions produced the corpus audit on this page. The tool is more reusable than the finding, which means the second run costs nothing.
  3. Run the same input through several tools and publish the disagreement. The disagreement is the finding. Nobody publishes it because it makes every vendor look imprecise, which is exactly why it is valuable.
  4. Do the thing and log it. A migration, a test, a recovery. The log is the dataset, and it only exists if you decided to keep it before you started.

The honest limit: small research is small. Eight guides is eight guides, and I would not generalise from it to the whole industry, which is why the component says the numbers are a ceiling and prints the classification rule. Stating the limit is not weakness. It is the difference between a finding and a claim, and it is what makes the finding citable.

Core Chapter 13 satisfy

What the brief decides and what content decides

If you remember one thing Step four handed you a page type and a job. Content decides what goes inside it. Confusing the two is how a brief that said "comparison" produces a two-thousand-word essay.

Not in this path.

Step four handed you a page type and a job. This step decides what goes inside it. Confusing the two is how a brief that said "comparison" produces a two-thousand-word essay with a table at the bottom.

The handover is worth being precise about, because the two steps get blurred in every workflow I have seen. Search intent decides the shape: what kind of page the results page rewards, in what format, at what depth, aimed at somebody at what stage. Content decides the substance: what specifically is in it that nothing else has. A brief that arrives with only the shape is half a brief, and the missing half is the one that decides whether the page is worth building.

from step four   comparison page · 8 of 10 results are comparisons · commercial · buyer stage
from step six    the same query run through all three tools on 23 Aug, with the disagreement

shape without substance   a well-formatted comparison that says what the vendors say
substance without shape   a genuinely new finding in a format the market does not want

Both failures are real and the second one is rarer and more painful, because you did the expensive part and then put it in the wrong container. That is chapter eighteen, and it is why format is decided by counting the results page rather than by whichever shape your team produces fastest.

One rule for the handover: the brief carries the count. Not "commercial intent" but "eight of ten results are comparisons, two are product pages, none is a guide". A label is somebody's judgement and a count is an observation, and only one of them survives being disagreed with.

Core Chapter 14 satisfy

The answer goes in the first screen

If you remember one thing The answer goes in the first screen or the page has not started. The common failure is not length, it is sequence: the answer is present and arrives after the reader has already gone back.

Not in this path.

The answer goes in the first screen or the page has not started. The common failure is not length, it is sequence: the answer is on the page and it arrives after the reader has already gone back.

Here is the same page twice. Same facts, same author, same total word count. The only difference is which sentence is first, and the left-hand column is not a strawman: every individual sentence in it is defensible, which is exactly why the pattern survives.

The first screen

The same page, and the only difference is which sentence is first

Same facts, same author, same total length. The left column withholds the answer for about four hundred words, which is the standard opening of an informational page. Read the left column and notice that every individual sentence in it is defensible. The sequence is the failure.

Answer withheld A reader who leaves at twenty seconds learns nothing

Content Decay: The Complete Guide for 2026

By the team · 12 min read · Updated August 2026

In today’s fast-moving digital landscape, content marketing has become more competitive than ever before. Businesses of every size are investing heavily in blogs, guides and resources in the hope of capturing organic search traffic.

But what happens after you hit publish? Many marketers assume their work is done. In reality, the story is just beginning, and understanding what comes next is essential for anyone serious about long-term organic growth.

What is content decay?

Content decay refers to the gradual decline in organic traffic that a piece of content experiences over time after reaching its peak performance. In this comprehensive guide, we will explore everything you need to know.

The fold, on a phone

Before we dive into the causes, it is worth understanding why this matters for your overall content strategy…

…the actual diagnostic arrives around here, at about the four hundredth word

Answer first A reader who leaves at twenty seconds has what they came for

Content Decay: The Four Shapes, and Which One You Have

Alston Antony · measured across 12 diagnoses · 23 August 2026

Content decay is four different failures sharing one word, and three of them are not fixed by a refresh. Which one you have is readable off a single Search Console chart in about two minutes.

Impressions hold, clicks fall → staleness. Update the facts.

Both slide slowly together → outclassed. Go and read who passed you.

A cliff, and the page type above you changed → the intent moved. Change format or stop.

Position stable, impressions falling → the demand left. Do nothing.

The fold, on a phone

Below the fold: what each one looks like in detail, the twelve charts, and the one that is most often misdiagnosed as the others.

What actually moved

One sentence. The diagnostic in the right column was already in the left-hand page, at about the four hundredth word. Nothing was written and nothing was cut. Almost every first-screen fix is this: a move, not an edit.

What the title did

"The Complete Guide for 2026" promises coverage. "The Four Shapes, and Which One You Have" promises a decision. The second one is also a contract, and chapter twenty-three is about what happens when you break it.

What this does not claim

Any ranking mechanism. The two numbers usually attached to this argument, about scroll depth and about a TL;DR lifting conversions, are both graded invented on this page. The reader came for an answer, which is sufficient.

The test is twenty seconds on a phone, with a person who has not read the page
Both columns contain the same facts and the same total word count

Almost every first-screen fix is a move rather than an edit. The sentence you need is usually already written, sitting at the four-hundredth word where the author finally got to the point, and the whole intervention is cutting it and pasting it above the introduction that was warming up to it.

Three things belong in a first screen and nothing else does. The answer, in a sentence. The reason to believe it, which is usually a number with a date or a source. And the shape of what follows, so somebody deciding whether to keep reading can decide. Everything else, including who you are and why the subject matters, belongs further down or nowhere.

What I will not do is attach a ranking mechanism to this. The two numbers usually quoted alongside this advice, that sixty percent of visitors never scroll and that a summary lifts conversions by a third, are both graded invented on this page because I cannot find who measured either. The reader came for an answer. That is sufficient, and it does not need a statistic to be true.

Awareness Chapter 15 satisfy

Pogo-sticking, and what Google has actually said

If you remember one thing Returning to the results page is a real behaviour and a badly documented signal. Write the first screen for the reader, and be suspicious of anybody selling you a ranking mechanism for it.

Not in this path.

Clicking a result and immediately returning to the results page is a real behaviour and a badly documented ranking signal, and the gap between those two facts is where a lot of confident advice lives.

What is certainly true: a person who arrives, does not find the answer, and goes back has had a bad experience, and enough of those means your page is not doing its job. That statement needs no algorithm behind it and it is the reason to write a better first screen.

What is not established: that Google measures this per page and demotes you for it. Google has said repeatedly that click behaviour is not a straightforward ranking signal, that it is noisy, and that using it naively would be trivially manipulable. The industry's response has largely been to assume it happens anyway and to sell tactics for it.

My position, stated so you can discount it appropriately: I write first screens as though it matters and I do not claim it as a mechanism. That is not fence-sitting. It is the only position the evidence supports, and the practical advice is identical either way, which is usually the tell that a disputed mechanism is not worth arguing about.

Where the argument does have consequences is measurement. If you believe pogo-sticking is a ranking factor, you will chase a metric nobody can see and optimise for the appearance of engagement, which produces the intro that withholds the answer to increase time on page. That page is worse for everybody and it is a direct product of believing something undocumented with total confidence.

Core Chapter 16 satisfy

Word count is the wrong unit

If you remember one thing Google states it has no preferred word count, in a parenthesis, in the list of things that mark content as made for search engines. Depth is measured in questions answered, not in words.

Not in this path.

Google states it has no preferred word count, in a parenthesis, in a list of things that mark content as made for search engines rather than for people. The sentence is short enough to quote in full and almost nobody does.

"Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"

The correlation that fuels the myth is real and it is backwards. Longer pages do tend to rank better on informational queries, because thorough answers to complicated questions take more words, and because more words gives more surface for links. Length is a symptom of having something to say, and target word counts treat the symptom.

The right unit is questions answered. A page is deep enough when somebody with the original problem can act, and the reasonable follow-up questions have been dealt with rather than deferred. That is measurable in a way word count is not: list the questions a reader will have after your first screen, and check each one is answered somewhere.

Two failure modes, and the second is much more common on good teams. Too thin is a page that answers the query and abandons the reader at the first obvious follow-up. Padded is a page that hits three thousand words by answering questions nobody asked, and it is worse than thin, because it buries the parts that were working underneath the parts that exist to hit a number.

The lesson you are reading is very long, and it is worth saying why that is not a contradiction. It is long because it is thirty-five chapters with three datasets and eight interactive drills in it, and because it has a reading-path control at the top that collapses it to seven chapters if you have twenty minutes. Length as an output of scope is fine. Length as an input to a brief is the thing Google put in a parenthesis.

Practical Chapter 17 satisfy

Write for the scanner, then for the reader

If you remember one thing A reader scans before they read. Short paragraphs, headings that say what they answer, and bold used for the load-bearing sentence rather than for keywords.

Not in this path.

Everybody scans before they read, and the scan decides whether the read happens. So the page has to work twice: once as a set of headings and bolded sentences, and once as prose.

The Nielsen Norman Group has been documenting scanning patterns since 2006, and the F-shaped pattern is the famous one: a horizontal sweep across the top, a shorter sweep further down, and a vertical scan down the left. Worth reading their own caveat too, which almost nobody quotes: it is one of several scanning patterns rather than the only one, and a page that reads well when scanned is the goal rather than a specific shape.

What survives a scan, in order of what people actually see: headings, bolded phrases, the first few words of each paragraph, list items, table rows, and anything visually distinct. What does not survive: the middle of a long paragraph, a conclusion, and anything that depends on having read the paragraph before it.

  1. Headings say what the section answers. "Content decay" is a label. "Content decay is four different failures" is an answer, and somebody who reads only the headings has learned the page.
  2. Bold the load-bearing sentence, not the keyword. Bolding a phrase because it is a target term trains readers to ignore your bold text, which costs you the one formatting tool that survives a scan.
  3. Two or three lines per paragraph. Not because attention spans changed. Because a wall of text has no entry points and a scanner needs somewhere to land.
  4. Front-load every paragraph. The first clause carries the claim and the rest supports it. A paragraph that builds to its point has hidden the point from anybody who did not read all of it.

One thing to be careful of, because this advice taken to its limit produces a bad page. Formatting is not thinking. A page chopped into eight-word paragraphs with a bold phrase in every one reads like a slide deck and cannot sustain an argument, and arguments are what a non-commodity page is made of. Scannable is a floor, not a target.

Core Chapter 18 show

Not every answer is an article

If you remember one thing A query ending in "checker" wants an input box and no length of article gets into it. Twenty-two formats exist, most content operations produce one, and format is decided before the first sentence.

Not in this path.

A query ending in "checker" wants an input box, and no length of article gets into that results page. Format is decided before the first sentence and it is decided by counting, not by whichever shape your team makes fastest.

This is the highest-cost mistake in the whole subject and it hides behind a process that feels rigorous. Research produces a cluster. The cluster goes into a calendar. The calendar produces a brief. The brief produces an article. At no point does anybody ask whether the market wanted an article, and step four of this roadmap has the receipt for what that costs: ten of ten results for "mortgage calculator" are calculators.

So here is the register. 22 formats, grouped by what the reader is trying to finish, with the query shape that asks for each one and what non-commodity looks like in that particular shape.

The register

Twenty-two formats, grouped by what the reader is trying to finish

Not by content type. "Blog post, landing page, video" is a production taxonomy and this is not a production decision. The middle column is the useful one: it says what non-commodity looks like in that particular shape, and it is different for every row.

Answer it

Somebody wants to know a thing and then leave.

  • Definition page

    Wanted by A query of the shape "what is X"

    What non-commodity looks like here A definition is level zero gain by default. It earns a page only when the definitions in circulation are wrong and you can show it.

    Nothing on this site

  • Explainer

    Wanted by "how does X work", "why does X happen"

    What non-commodity looks like here The mechanism, one layer below the level everybody else stops at.

    /ai-seo/query-fan-out/

  • Question page

    Wanted by A single question with a short true answer

    What non-commodity looks like here Usually none. Most of these belong as a section of something bigger.

    Nothing on this site

  • Glossary

    Wanted by Vocabulary lookups across a subject

    What non-commodity looks like here Coverage plus consistency. A glossary is a reference object, not an article.

    /content-seo/#glossary

Help them decide

Somebody is choosing between named options and has money involved.

  • Comparison

    Wanted by "X vs Y"

    What non-commodity looks like here Criteria you chose and defended, and a verdict you are willing to be wrong about in public.

    /seo-tools/

  • Alternatives page

    Wanted by "X alternatives", "X competitors"

    What non-commodity looks like here Naming who each alternative is actually for, including the case where the answer is stay where you are.

    Nothing on this site

  • Best-of list

    Wanted by "best X for Y"

    What non-commodity looks like here Stated criteria, disclosed testing, and at least one entry you removed and said why.

    /seo-tools/free/

  • Pricing page

    Wanted by "X pricing", "how much does X cost"

    What non-commodity looks like here The number, current, with the date. Most pricing content is a paraphrase of a pricing table.

    /seo-tools/

  • Review

    Wanted by "X review", "is X any good"

    What non-commodity looks like here First-hand use. Google names first-hand review as its own example of a unique point of view.

    /seo-tools/taja/

Help them do it

Somebody is mid-task and needs the next step to work.

  • Tutorial

    Wanted by "how to X"

    What non-commodity looks like here Screenshots of the actual interface, the version, and the step that goes wrong.

    /seo/gmail-account-for-seo/

  • Troubleshooting page

    Wanted by "X not working", "fix X"

    What non-commodity looks like here The cause, ranked by how often it is the cause, from cases you have actually seen.

    /seo/blogger-not-indexing-google-fix/

  • Checklist

    Wanted by "X checklist", "X audit"

    What non-commodity looks like here Order, and an item most checklists leave out with the reason.

    /resources/technical-seo-checklist/

  • Template

    Wanted by "X template", "X spreadsheet"

    What non-commodity looks like here The object itself. The page around it is documentation.

    /resources/keyword-research-template/

  • Course or lesson

    Wanted by A subject somebody wants taught in order

    What non-commodity looks like here Sequence, prerequisites, and a deliverable at the end.

    /seo-roadmap/

Prove something

Somebody needs a reason to believe you rather than the other six.

  • Case study

    Wanted by "X case study", "did X work"

    What non-commodity looks like here The raw export, the window, and what else changed at the same time.

    /seo-case-studies/

  • Original research

    Wanted by A question nobody has counted the answer to

    What non-commodity looks like here The dataset. This is the only format whose whole value is the gain.

    /content-seo/#guide-audit

  • Teardown

    Wanted by A named thing examined in public

    What non-commodity looks like here Specificity. A teardown of a category is a listicle; a teardown of one named page is evidence.

    Nothing on this site

  • Experiment writeup

    Wanted by "does X actually work"

    What non-commodity looks like here A method, a control, and the willingness to publish it when it did not work.

    Nothing on this site

Give them a thing

The answer is not prose. It is an object they take away.

  • Free tool

    Wanted by A query ending in "checker", "calculator", "generator"

    What non-commodity looks like here The tool works. Nothing you write instead of building it will rank for this.

    /seo-tools/free/seo-google-penalty-checker/

  • Dataset or table page

    Wanted by "list of X", "X statistics"

    What non-commodity looks like here Completeness plus a last-updated date that is real.

    Nothing on this site

  • Directory

    Wanted by "X tools", "best X software"

    What non-commodity looks like here Coverage, structure and maintenance. It is a product, not a post.

    /seo-tools/

  • Interactive explainer

    Wanted by A concept that is hard to describe and easy to demonstrate

    What non-commodity looks like here The demonstration. Everything on this page you can click is one of these.

    /content-seo/#gain-drill

6 of the 22 have no example on this site and that gap is left visible rather than filled with something plausible. It tells you what this content operation can currently produce, which is the same finding the exercise below will produce about yours. Almost every content team can make one shape well and makes it for everything, and the queries that wanted a different shape are simply lost without anybody noticing.

Format is decided by counting the results page, not by the content calendar
16 of 22 formats exist on this site. The rest are honest gaps

The middle column is the one that makes this a decision table rather than a taxonomy. Non-commodity means something different in each shape: for a definition page it means the circulating definitions are wrong and you can show it, for a free tool it means the tool works, and for a comparison it means criteria you chose and a verdict you are willing to be wrong about in public.

6 of the 22 have no example on this site and the gaps are left visible rather than filled with something plausible. That tells you what this content operation can currently produce, which is the same finding the exercise below will produce about yours. Almost every team makes one shape well and makes it for everything, and the queries that wanted a different shape are lost silently.

Practical Chapter 19 show

Visual content, and the four kinds that earn a slot

If you remember one thing Four kinds of image earn a slot: evidence, structure, comparison and instruction. A stock photograph is none of them, and nine of eleven posts on this site have no table at all.

Not in this path.

An image is carrying a claim or it is filling space, and "add images to your content" is advice that produces the second one, because it never said what the image is for.

Google's position is genuinely encouraging and it is an opportunity rather than an obligation: its generative AI features can bring in relevant images and video, which means more ways for a site to appear than a blue link. The operative word in its guidance is "relevant", and a stock photograph of somebody at a laptop is not relevant to anything.

Four kinds that earn a slot

An image is carrying a claim, or it is filling space

"Add images" is advice that produces stock photography, because nobody said what the image is for. Each row below has a one-question test and they all reduce to the same thing: would deleting this weaken something on the page? A decorative image passes no test because it was never carrying anything.

  1. 01

    Evidence

    Shows the thing happened. A screenshot of the export, the interface, the error, with the date in the frame.

    The test

    Would deleting it weaken a claim on the page?

    What fails it

    A stock photograph of somebody at a laptop, which weakens nothing because it was never carrying anything.

  2. 02

    Structure

    Shows a relationship prose is bad at. A tree, a flow, a matrix, a before and after.

    The test

    Did you have to write two paragraphs explaining the same relationship anyway? Then the diagram is not doing the job.

    What fails it

    A diagram of five boxes with the section headings in them, which restates the outline and adds nothing.

  3. 03

    Comparison

    Puts two or more named things side by side on criteria you chose.

    The test

    Can a reader reach a decision from the image alone?

    What fails it

    A comparison chart where every row is a tick for every product.

  4. 04

    Instruction

    Shows exactly where to click, in the interface as it currently looks, with the version named.

    The test

    Does it match what the reader sees on their screen this month?

    What fails it

    A three-year-old screenshot of an interface that has been redesigned twice, which actively costs the reader time.

Mine, measured on 23 August 2026: 38 images across 11 posts, and three of those posts have none at all. Two of the three are the longest things I have published. The count is not the interesting part. Almost every one of those images is the instruction kind, because the posts are tutorials, and I have produced very close to nothing in the evidence, structure or comparison categories on the blog.

Google’s own guidance names high-quality relevant images as an opportunity, not an obligation
A stock photograph is not a fifth kind. It is the absence of a decision

The fourth kind, instruction, is the one with a maintenance cost attached and the one most likely to go actively wrong. A screenshot of an interface that has been redesigned twice does not merely fail to help. It sends the reader looking for a button that no longer exists, which is worse than the page having no image at all, and it is the single most common form of staleness in tutorial content.

A note on generated images, since it is now one prompt away. The rule is the same as everywhere else on this page: an illustration generated to fill a slot is a stock photograph with extra steps. Where a generated image genuinely helps is the structure category, when you need a diagram of something abstract, and there the honest version is to draw it rather than to generate a picture of a concept. Google Merchant Center already requires AI-generated product images to be labelled in their metadata, which is worth knowing about where it applies.

Practical Chapter 20 show

Tables and comparisons, the most underproduced format on the web

If you remember one thing The densest evidence format on the web and the least produced. A table forces you to pick criteria, and picking criteria is the part of a comparison that is actually work.

Not in this path.

A table is the densest evidence format available and almost nobody builds a useful one, because the useful version requires deciding what matters and the useless version can be assembled from three marketing pages in twenty minutes.

The distinction is not presentational. A feature dump and a criteria table look identical and do opposite jobs, and the difference is entirely in what happened before the table was drawn.

The same three products, twice

A comparison table is criteria, or it is a feature dump

Left: assembled from three marketing pages in twenty minutes. Every row is a yes and the table decides nothing. Right: criteria chosen on paper before anybody opened a product, and four of the six rows disagree. The disagreement is the entire value, and it is why the second one takes an afternoon.

Feature dump Twenty minutes. Decides nothing

Feature Tool ATool BTool C
Keyword research YesYesYes
Rank tracking YesYesYes
Site audit YesYesYes
Backlink data YesYesYes
Content tools YesYesYes
API access YesYesYes

Six rows, eighteen cells, one distinct value. A reader finishes this table knowing exactly what they knew before, and the page has spent a screen saying so.

Criteria first An afternoon. Two rows change the answer

Criterion Tool ATool BTool C
Cost of the cheapest plan that includes the API $449/mo$249/moNo API at any tier
Rows per keyword export, cheapest plan 1,00010,000500
Seats included before the price changes 113
Index refresh, as the vendor states it Every 15 minDailyWeekly
Can you cancel without contacting sales YesYesYes
Free trial without a card NoNoNo

The last two rows are greyed because every column agrees, which makes them decoration. Deleting rows where nobody disagrees is the fastest way to make a comparison useful, and nobody does it because a longer table feels more thorough.

Four rules that produce the right-hand table

  1. Write the criteria before opening any product. Criteria derived from a feature list are a feature list with extra steps, and they will always favour whichever product you opened first.
  2. Every criterion is a question a buyer asked. If you cannot name the person who would care about a row, delete the row.
  3. Delete every row where all the answers match. It is true, it is accurate, and it costs the reader time to read a row that cannot change their answer.
  4. Put the number in the cell, not a tick. "Yes" hides the difference between a thousand rows and ten thousand, which is the difference the reader is actually buying.

Mine, from the corpus audit: 3 tables across 11 posts, and nine of the eleven have none at all. On a page arguing that this is the densest evidence format on the web. Notice also what the count implies about the two rules above that require work: I have not been picking criteria, because I have not been building tables.

The claim that a table gets you cited by AI systems is graded inference on this page
The reader reason is sufficient, which is lucky, because it is the one with evidence behind it

The rule that does the most work is the one about deleting rows where every column agrees. A row that cannot change the reader's answer is costing them attention and buying nothing, and it survives because a longer table feels more thorough. The same instinct is what produces the twenty-row comparison where every cell is a tick.

On the claim that tables get you cited by AI systems, which you will have heard: it is graded inference on this page. It is plausible, it is repeated constantly, and I cannot find anybody who has isolated it, including the vendors selling the advice. Build the table because it lets a reader decide. That reason is sufficient, it is not disputed, and it does not stop being true if the citation claim turns out to be wrong.

My own count, from the audit: three tables across eleven posts, and nine of the eleven have none. I am arguing for the format I have most neglected, which is either the least credible position on this page or the most, depending on whether you think people should recommend things they have not done. I would rather print the number and let you decide.

Awareness Chapter 21 show

Video on a page, and what it is actually for

If you remember one thing A video is not a ranking accessory. It is either the primary answer, a proof you did the thing, or it does not belong on the page.

Not in this path.

A video on a page is the primary answer, or it is proof you did the thing, or it does not belong there. There is no fourth reason, and "engagement" is not one.

The primary-answer case is straightforward: some things cannot be written. Anything with timing, spatial relationships, or an interface you have to watch somebody move through. For those queries the results page usually tells you directly, because video results appear in it, and step four has the counting method.

The proof case is the interesting one and it is badly underused. A thirty-second clip of a tool actually doing the thing you claim it does is evidence in a way a paragraph is not, and it is very hard to fake casually. This is the visual-evidence category from the last chapter with a play button on it.

What does not work is a video embedded because a checklist said pages with video perform better. That correlation exists, and it exists because people who make videos are usually investing more in the page overall. Embedding a loosely related clip does not import the investment. Google's own guidance on this is measured: relevant images and video are more ways for your site to appear, and both adjectives are load-bearing.

Eight of my eleven posts embed a video and that number flatters me. Most of them are the YouTube version of the same tutorial, which is legitimate, and none of them is the thirty-second proof clip that the proof case describes. Having a video and using video are different, and the audit only measures the first one.

Core Chapter 22 structure

Every section has to survive being ripped out

If you remember one thing Every section has to survive being ripped out of the page. A passage that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere, by anyone, ever.

Not in this path.

A section that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere, by anyone, ever. Not by a reader who skimmed to it, not by somebody sharing it, and not by a retrieval system.

The reason this happens to good writers is mechanical rather than careless. When you write a section, the previous section is on your screen, so a pronoun referring back to it is perfectly clear. It stops being clear the moment anybody arrives from a search result, a table of contents link, or a quotation, which is most of the ways sections get read.

The extraction test

Four sections, read in place and then ripped out

Three of these four broken versions are from drafts of this page. Nobody writes "as I mentioned above" on purpose. You write it because the paragraph you are referring to is on the screen while you type, and it is not on the screen when somebody arrives from a search result or quotes you in a newsletter.

Flip it and read the left column as though you found it with no idea what page it came from.

  1. 01

    The four shapes

    Cannot stand alone

    It has four of them, and three are not fixed by the thing everybody does. The first is the one people mean when they say the word.

    Stands alone

    Content decay has four shapes, and three of them are not fixed by a refresh. The first is staleness, which is the one people mean when they say the word.

    What breaks

    Opens on "it" and "the thing everybody does". Lifted out, the passage does not say what it is about.

    What the fix cost

    Nine words longer. Both nouns were already known to the writer and neither was on the page.

  2. 02

    Why this matters for updates

    Cannot stand alone

    As I mentioned above, this is why the second approach usually fails. Bear that in mind when you get to the next section.

    Stands alone

    Bumping a publication date without changing the content usually fails, because freshness is a property of the query rather than of the page.

    What breaks

    Two cross-references and a forward reference in twenty-two words. It cannot be quoted anywhere, by anyone, ever.

    What the fix cost

    One word shorter, and it now contains the claim instead of pointing at it.

  3. 03

    How to check it

    Cannot stand alone

    Open the tool and run the same report as before. If the numbers differ from the earlier ones, you have found the problem described earlier.

    Stands alone

    Open Search Console, filter to the URL, and chart clicks against impressions for twelve months. If impressions hold while clicks fall, the page is stale rather than outranked.

    What breaks

    Three references to things outside the passage: the tool, the report, the earlier numbers. A reader arriving here cold cannot act on it.

    What the fix cost

    Longer, and it is now the only paragraph on the page somebody would screenshot.

  4. 04

    The exception

    Cannot stand alone

    There is one case where this does not apply, and it is the case discussed in the previous chapter. Everything else follows the rule above.

    Stands alone

    One case does not follow the rule: a query that genuinely deserves freshness, such as a news event or a price. For those, a real update to the facts does move the page.

    What breaks

    The entire content of the passage is a pointer to another passage. Extracted, it says nothing at all.

    What the fix cost

    It now contains the exception instead of promising it, which is also better for a reader going in order.

This is not a retrieval hack and it is not chunking. Google published on 10 July 2026 that there is no requirement to break your content into tiny pieces for AI to understand it. A section that stands alone is better for a reader who skimmed straight to it, better for anybody quoting you, and better for retrieval as a side effect. The reader reason is sufficient, which is convenient, because it is the only one with a document behind it.

Copy three of your own sections into an empty document and read them cold
The usual offenders: an opening "it", a backward reference, and an unnamed tool

The fix costs one sentence per section and usually costs zero extra words, because you are replacing a pronoun with the noun it was standing in for. It is the highest-return editing pass in this lesson, and the only one I would run on a page I otherwise had no time for.

I want to be precise about what this is not, because the neighbouring advice is currently being oversold. This is not chunking. Google published on 10 July 2026 that there is no requirement to break your content into tiny pieces for AI to understand it, and that there is no ideal page length. Writing sections that stand alone is a readability practice that happens to help retrieval. Cutting your page into three-hundred word blocks because somebody said models like it is a different thing, and the document those people cite says not to.

Practical Chapter 23 structure

A heading is a contract

If you remember one thing A heading is a promise about what the next section answers. Break it and the reader learns to skip your headings, which costs more than any keyword it might have carried.

Not in this path.

A heading promises what the next section answers. Break the promise and the reader learns to skip your headings, which costs more than any keyword the heading might have carried.

The mechanism is trust, and it decays fast. A reader who scans your headings, picks one, jumps to it, and does not find what it promised will not scan the next set. They will read linearly, get bored, and leave, which is the outcome the headings existed to prevent. Two broken headings is enough.

So the test for a heading is whether somebody who read only that heading would be able to predict the section. Three shapes that pass, and one that does not:

the question    "Which decay shape do you have?"          predicts a diagnostic
the claim       "Word count is the wrong unit"            predicts an argument
the object      "Four shapes, and what each one needs"    predicts a list of four

the label       "Content decay"                           predicts nothing at all

The label is not wrong, it is empty. It tells a scanner the subject of the section, which they could have guessed from the page title, and it tells them nothing about whether the section is the one they need. On a long page, labels are why people give up.

Two structural rules and then this chapter is finished, because heading hierarchy is on-page work and belongs to the next lesson. Use one H1, which is the title, and Google has a short video saying that more than one is not fatal. And do not skip levels for visual reasons, because the outline is the only structure a screen reader and a table of contents can see.

The check that catches everything: read your own headings alone, in order, with the body deleted. If they read as a coherent summary of the argument, the page is structured. If they read as a list of nouns, it is filed rather than structured.

Practical Chapter 24 structure

Name things specifically

If you remember one thing Name the thing. "A popular SEO tool" is unrecognisable to a reader and to a machine; "Ahrefs, on the Lite plan, in August 2026" is checkable by both.

Not in this path.

"A popular SEO tool" is unrecognisable to a reader and to a machine. "Ahrefs, on the Lite plan, in August 2026" is checkable by both, and the difference costs nothing to produce.

Vagueness in content is almost never a stylistic choice. It is what you write when you do not know, when you are hedging a claim you cannot support, or when you are avoiding naming a competitor. All three are visible to a reader who does know, and the third one is visible to everybody.

What to name, specifically: tools, with the plan or version. Companies, rather than "a leading provider". Documents, with the publisher and the date it says it was updated. People, where they have said something publicly. Numbers, with the date and the source. Interfaces, with the current label of the button rather than a paraphrase of it.

There is a retrieval argument here and I will make the modest version of it. Systems that read your page work with the entities in it, and a page that says "a popular SEO tool" contains no entity to work with. That is a claim about what is legible, not about ranking, and I would make the same recommendation with no AI systems in the picture, because specificity is how a reader can tell whether you have used the thing.

What I will not do is recommend chasing mentions, and Google's July 2026 guidance is unusually direct about it: seeking inauthentic mentions across the web "isn't as helpful as it might seem". Name things because it makes your page checkable. Do not build a programme around getting named yourself.

Core Chapter 25 structure

What Google published about AI features, and what it told you to ignore

If you remember one thing Google published in July 2026 that you can ignore chunking, ignore llms.txt and ignore chasing mentions. Most of what is sold as generative engine optimization is contradicted by the vendor’s own document.

Not in this path.

Google published a guide to optimizing for its generative AI features on 10 July 2026, and the most useful part of it is the list of things it says you can ignore. Two of them are being sold hard by people citing that same document.

Start with what it says to do, because it is short. Create valuable, non-commodity content for your audience, which it says will "likely influence your website's presence in generative AI search in the long run more than any of the other suggestions in this guide". Keep the site technically crawlable. Add high-quality relevant images and video. Organise content so readers can follow it. That is the whole substance, and it is the first five decisions of this lesson.

Now the list of things to stop doing, quoted rather than paraphrased, because the paraphrases are the problem.

"There's no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users."
"You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them."

And the one that most directly contradicts current practice. Google defines a fan-out query in the same document, as one of the concurrent related queries a model issues behind a single prompt, and then says this about building for them: creating separate content for every possible variation of how people might search, including fan-out queries, primarily to manipulate rankings violates the scaled content abuse spam policy, and is an ineffective long-term strategy besides.

I want to be fair to the tactic, because there is a defensible version. Looking at fan-out queries to understand what a model is actually asking is research, and research is legitimate. Generating a page for each one is the thing being named. The distinction is the same one running through this whole lesson: reading the market is fine, and manufacturing pages to match it is what the policy is about.

So the practical answer to "how do I optimize for AI search" is unsatisfying and cheap: you already know. Have something worth citing, say it clearly, name things specifically, and make sure a crawler can reach it. Every one of those is a decision earlier in this lesson, which is why this page has no chapter on GEO tactics and a chapter on refusing them.

Core Chapter 26 maintain

Freshness is a property of the query

If you remember one thing Freshness is a property of the query, not of your page. Some queries deserve it and most do not, and changing a date without changing the content is the one intervention that reliably does nothing.

Not in this path.

Some queries deserve fresh results and most do not. Freshness is a property of what was asked, not of your publication date, which is why changing the date on an unchanged page is the one intervention that reliably does nothing.

The mechanism is easy to reason about from the searcher's side. "Best iphone" needs a recent answer because the answer changes annually. "How does photosynthesis work" does not, and a results page full of pages published last month would be worse rather than better. Google's own channel has a video titled "Is freshness an important signal for all sites?" and the answer is no, which is a more useful sentence than most of what is written about this.

So the first question about any page is not "when did I last update this". It is does this query deserve freshness at all, and you can answer it by looking at the dates on the current results page. If the top ten are all from the last six months, the query wants recency and you are in a maintenance race. If they span five years, recency is not what is being rewarded and a refresh cycle is spending money on nothing.

Three things that are actually freshness work, as opposed to date changes. Updating facts that have changed: prices, versions, interfaces, statistics, the law. Adding what has happened since: a new competitor, a policy change, an outcome you now know. And removing what is no longer true, which is the one that gets skipped because deleting a paragraph feels like losing something.

Two anti-patterns worth naming because both are common and both are visible to readers. Putting the year in the title and bumping it annually, which is a maintenance commitment most people do not keep and looks exactly like what it is when they do not. And displaying an updated date that changed because somebody fixed a typo, which is a small lie that costs trust the first time a returning reader notices nothing is different.

Core Chapter 27 maintain

Content decay has four shapes

If you remember one thing Four different failures share the word decay: staleness, being outclassed, the intent moving, and the demand leaving. Three of the four are not fixed by a refresh.

Not in this path.

Four different failures share the word decay, and three of them are not fixed by a refresh. Telling them apart takes two minutes on one Search Console chart and it is the difference between a useful afternoon and a wasted one.

The reason this matters more than it sounds: refreshing is the default response to a falling page in almost every content team, and it is the correct response to exactly one of the four situations below. Applied to the other three it produces a page with a new date and the same problem, and it feels like progress, which is the expensive part.

Four shapes

Content decay is four different failures sharing one word

Diagnosis happens by shape. Nobody reads a definition and then identifies their case; they look at twelve months of Search Console, recognise a pattern, and act. So here are the four patterns, with what each one actually is, and with the wrong move given the same space as the right one.

  1. 01

    12 months

    Staleness

    The page is about something that changed. Prices, interfaces, versions, laws, the year in the title.

    How you know

    Impressions hold and clicks fall, or the page still ranks and the comments say it is out of date.

    What to do

    Update the facts, the screenshots and the date, and say what changed. This is the one case where a genuine update is the whole answer.

    The wrong move

    Rewriting the argument. The argument was fine; the facts moved.

  2. 02

    12 months

    Outclassed

    Somebody published something better. Your page did not get worse; the results page got better.

    How you know

    A steady slide with no obvious event, and the new pages above you have something yours does not.

    What to do

    Go and read what beat you, then add the thing you do not have. If you cannot name that thing, you are about to change nothing.

    The wrong move

    A freshness edit. Bumping the date on a page that lost on substance is theatre.

  3. 03

    12 months

    The intent moved

    The results page changed shape. Nine articles became nine tools, or a comparison replaced a guide.

    How you know

    A cliff rather than a slide, and the page type above you is no longer yours.

    What to do

    Change format or stop. The previous lesson has the count method for this and it takes ten minutes.

    The wrong move

    Adding two thousand words to an article on a results page that no longer wants articles.

  4. 04

    12 months

    The demand left

    Fewer people search for it, or the answer now arrives without a click.

    How you know

    Your position is stable or improving and impressions fall anyway.

    What to do

    Nothing, usually. Consolidate it into something with a future, or leave it and stop measuring it monthly.

    The wrong move

    A quarterly refresh cycle on a page whose market is shrinking, which is how teams spend a year improving a page nobody wants.

Two of the four end in doing nothing, and that is the point of diagnosing rather than refreshing. A quarterly refresh cycle applied to all four produces one correct outcome, one expensive rewrite that would have happened anyway, and two pages that get worked on every quarter for the rest of their lives. The most expensive mistake is treating a competitive loss as a freshness problem, because it produces a page with a new date and the same gap, and it feels like progress.

The curves are illustrative shapes, not measurements. Your own chart is the data
Three of the four are not fixed by a refresh, and one of the four is not fixed by anything

Two of the four end in doing nothing, and that is the finding rather than a hedge. A page whose market has shrunk does not have a content problem, and a quarterly refresh cycle applied to it will consume hours every year for the rest of its life while the author wonders why the numbers never move.

The one most often misdiagnosed is the second. Being outclassed looks like staleness because both are gradual, and the tell is what is above you: if the new pages have something yours does not, you were beaten on substance and a date change is theatre. If you cannot name the thing they have in a single sentence, you are not ready to touch the page, because you do not yet know what you would be changing.

Core Chapter 28 maintain

Updating a page, and the three ways it goes wrong

If you remember one thing Update the thing that lost, not the date. If you cannot name what a competitor has that you do not, you are about to spend a day changing nothing.

Not in this path.

Update the thing that lost, not the date. Everything else about refreshing follows from that sentence, and the diagnosis from the last chapter is what tells you what lost.

Here are eight real situations and five verbs. Pick one for each before you open the reasoning, and notice how often the reflex answer is the trap.

Update, rewrite, merge, prune or leave?

Eight pages, and twice the right answer is to do nothing

Each one gives you the signal from Search Console and nothing else, which is what you actually have. Pick a verb before you open the reasoning. The trap on every item is the move a quarterly refresh process would make automatically, because it decided the verb before it looked.

  1. 01 A 2024 tutorial for a tool that redesigned its interface in June

    The signal Position steady at 4, impressions steady, clicks down 40%, comments saying the screenshots are wrong.

    Show the reasoning

    Update

    Textbook staleness. The argument is fine and the facts moved, so replace the screenshots, name the current version, and say what changed at the top.

    What a refresh process would do instead Rewriting it. Nothing about the explanation was wrong.

  2. 02 A guide that slid from position 3 to 11 over eight months with no event

    The signal Slow decline, three new pages above you, all of them with original testing you do not have.

    Show the reasoning

    Rewrite

    You were outclassed on substance. The only useful move is to go and get the thing they have and you do not, and if you cannot name it in one sentence you are not ready to touch the page.

    What a refresh process would do instead A freshness edit. Bumping the date on a page that lost on substance changes nothing and takes a day.

  3. 03 Two posts, published a year apart, both answering "how to find long tail keywords"

    The signal Both rank between 8 and 20, neither has ever been in the top five, impressions split between them.

    Show the reasoning

    Merge

    One need, two destinations, and the site cannot say which it means. Merge into the stronger URL, redirect the other, and keep the best paragraphs from both.

    What a refresh process would do instead Improving both. You will spend twice the effort to keep competing with yourself.

  4. 04 A 700-word post from 2023 about a plugin that no longer exists

    The signal Zero clicks in twelve months, three impressions, no inbound links, no internal links.

    Show the reasoning

    Prune

    Nothing points at it, nobody arrives, and the subject is gone. Delete it and redirect to the nearest genuinely relevant page, or return a 410 if there is no such page.

    What a refresh process would do instead Redirecting it to the home page, which is treated as a soft 404 and discards whatever you were preserving.

  5. 05 A definition page, position 2, stable for two years

    The signal Position 2, impressions flat, clicks flat, nothing above it has changed.

    Show the reasoning

    Leave

    It is working. A quarterly refresh process would touch this page four times a year for no reason, and every touch is a chance to make it worse.

    What a refresh process would do instead Refreshing it because it is old. Age is not a diagnosis.

  6. 06 A comprehensive article on a query whose results page is now nine free tools

    The signal Fell from 6 to 34 in one month. Everything above it has an input box.

    Show the reasoning

    Rewrite

    The intent moved and the format is wrong. This is a build decision rather than an editorial one: either produce the tool or stop competing here and say so in your map.

    What a refresh process would do instead Adding two thousand words. No length of article gets into a results page that wants software.

  7. 07 A well-ranked guide whose impressions have fallen 30% while its position improved

    The signal Position 5 to 3, impressions down, clicks down proportionally.

    Show the reasoning

    Leave

    You got better and the market got smaller. The demand left, and there is nothing on the page to fix. Note it, stop reporting on it monthly, and put the hours somewhere with a future.

    What a refresh process would do instead A refresh cycle, which is how a team spends a year improving a page for a shrinking market.

  8. 08 A post that ranks for forty queries it does not answer, and none it does

    The signal Decent impressions, terrible click-through, average position 14 across a scattered query set.

    Show the reasoning

    Rewrite

    The page is being collected rather than targeted. Pick the one query it should own, answer that in the first screen, and let the other thirty-nine go.

    What a refresh process would do instead Adding sections for the forty queries, which is the fan-out farming Google names as a spam policy violation.

2 of the eight are leave. If you got both of those, you have the instinct this chapter is trying to build, and it is the hardest one to hold onto because it produces nothing to put in a report. The other transferable thing here: three items look like staleness and only one of them is, which is why the diagnosis comes before the verb and not after it.

The three ways an update goes wrong, in order of how often I see them.

  1. Cosmetic. A new date, a reworded introduction, three new sentences, and nothing that addresses why the page fell. This is the most common outcome of a scheduled refresh, because a schedule produces the activity without producing the diagnosis.
  2. Additive only. Two thousand words bolted onto a page that was already long enough, because adding feels safer than deleting. The page is now worse at the job it was doing, and the section that used to answer the query is four screens down.
  3. Identity loss. A rewrite so thorough that the page no longer answers the query it ranked for. This one is genuinely dangerous, and the check is cheap: before you publish, open Search Console, look at what the page currently ranks for, and confirm the new version still answers the top few.

And the honest measurement problem, which is why the claim "updating recovers traffic" is graded inference on this page rather than documented. Almost every published example of a successful refresh also changed the title, the internal links and the depth at the same time. Something worked. Nobody has isolated which thing, and I have not either.

What I would do about that, practically: change one thing at a time on pages you care about, write down what you changed and when, and give it enough weeks to mean anything. Google has a video on how long SEO takes for new pages, and the same patience applies here. A refresh judged after ten days has been judged on noise.

Practical Chapter 29 maintain

Duplicate content and commodity content are different problems

If you remember one thing Duplicate content is a consolidation problem with a documented mechanism. Commodity content is a value problem with no mechanism at all, and it is the expensive one.

Not in this path.

Duplicate content is a consolidation problem with a documented mechanism. Commodity content is a value problem with no mechanism at all. They get discussed together, they look nothing alike, and only one of them is expensive.

Duplicate means several URLs answering one need, on your site or across sites. Nothing is being punished. Google picks one URL to represent the set, using signals it can see, and your preference is one input among several. The fixes are structural and they belong to step five: merge and redirect, differentiate genuinely, or canonicalise. This is a solved problem with published documentation and a short Google video.

Commodity means your page is not a duplicate of anything and adds nothing anyway. Every sentence is original prose, no other URL is competing with it, and it is still the average of the results page. There is no mechanism to explain, no canonical tag to add, and no report that flags it. It is the expensive one and it is the subject of this lesson.

duplicate   three URLs, one need        → Google picks one. Consolidate, redirect, canonicalise.
commodity   one URL, nothing new        → no penalty, no report, no traffic. Add something or delete it.
thin        any length, no added value  → Google's own framing. Word count is not the diagnosis.

Thin is the third word in this cluster and it is the most misused. Google's quality material treats thin content as content with little or no added value, at any length, and a four-thousand-word restatement of the top three is thin by that definition. The industry reads "thin" as "short" and then writes longer pages, which is how a diagnosis produces exactly the wrong treatment.

The syndication case deserves a sentence, because it is where these two genuinely overlap. Republishing your own article on Medium or LinkedIn is a duplicate question, solvable with a canonical or by accepting that one version wins. Republishing somebody else's press release, as thousands of sites do, is both: duplicate by construction and commodity by definition, which is why those pages are invisible even when the technical setup is perfect.

Core Chapter 30 maintain

AI-assisted content, against what Google actually published

If you remember one thing Google’s position has not moved since February 2023: appropriate use of AI is not against the guidelines, and generating pages primarily to manipulate rankings is. The tool is not the line, the purpose is.

Not in this path.

Google's position has not moved since February 2023: appropriate use of AI is not against the guidelines, and using automation to generate content primarily to manipulate rankings is a spam policy violation. The tool is not the line. The purpose is.

The most useful thing in that 2023 post is a historical analogy nobody quotes. Ten years earlier there were concerns about a rise in mass-produced human content, and Google's own observation is that nobody would have thought it reasonable to ban human writing in response. The answer then was to get better at rewarding quality, and it says that is the answer now.

Twelve uses, graded against the policy

The line is purpose, not the tool

Every row graded against one published sentence rather than against a feeling: appropriate use of AI is not against the guidelines, and using automation to generate content primarily to manipulate rankings is a spam policy violation. That has been Google’s position since February 2023 and it has not changed while everything written about it has.

"Our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years."

Google Search Central, February 2023. Guidance about AI-generated content . The same post points out that ten years earlier there were the same concerns about mass-produced human content, and that banning human writing was never the answer either.

  1. Fine 5 Named as useful, or not objected to anywhere in the documentation.
    • Researching a topic and finding what has already been said

      Google names researching a topic as a case where generative AI is particularly useful. Google Search Central

    • Turning a transcript of yourself into a structured draft

      The knowledge is yours and the machine did the typing. Google names adding structure to original content as a useful case. Google Search Central

    • Building an outline from the top ten results

      It is competitor reading with fewer tabs. It also produces the median outline, which is why it cannot be the last step.

    • Drafting a section you then rewrite from your own knowledge

      The draft is scaffolding. If your rewrite does not change the facts, you did not have any.

    • Generating alt text, then checking every one against the image

      Google names alt text among the metadata to focus accuracy on when generating automatically. Google Search Central

  2. Depends 4 Legitimate when a human owns the output. Becomes spam at the point nobody is checking.
    • Translating your own content, then having a speaker check it

      Fine when a human owns the result. Google has a specific video on whether translated content is a duplicate content issue and the answer is about quality rather than duplication.

    • Summarizing your own long content into a shorter page

      Two pages for one need is a cannibalization decision before it is a content one. Chapter twenty-nine.

    • Generating a definition page for every term in your niche

      Legitimate if each one is checked and each one is needed. It becomes scaled content abuse at exactly the point where nobody is checking.

    • Producing a hundred location pages from one template

      The template is not the problem. A hundred pages that differ only in a place name and add nothing local is the problem.

  3. Policy violation 3 Matches the published spam policy, whether a person or a model produced it.
    • Publishing a first draft with the facts unchecked

      Not because it is AI. Because nobody verified it, which is chapter thirty-one and the only rule in this section that matters.

    • Generating a page for every fan-out query you can extract

      Google states directly that creating separate content for every variation, including fan-out queries, primarily to manipulate rankings violates the scaled content abuse policy. Google Search Central

    • Generating many pages to cover a topic space, without adding value

      This is the scaled content abuse policy almost word for word: many pages generated for the primary purpose of manipulating rankings and not helping users. Google Search Central

Look at the middle band, which is nine of the twelve rows. Every one of them turns on the same hinge, and it is not a technical one: whether anybody was going to check the output. That is the whole of the next chapter, and it is why "is AI content allowed" is the wrong question. The right question is who is accountable for each fact on the page.

Spam policies for Google web search , Read 23 August 2026. Google states last updated 15 May 2026.
Nothing in this table is graded by how the content was produced

Look at the middle band, which is three quarters of the table. Every row in it turns on the same hinge, and it is not technical: whether anybody was going to check the output. That is the next chapter, and it is why "is AI content allowed" is the wrong question. The right question is who is accountable for each fact on the page.

Where I actually use it, since you should be able to price my advice. Research and finding what has been said. Turning a transcript of me talking into structured prose, which is typing rather than thinking. First-draft sections that I then rewrite, where the rewrite always changes the facts because the draft did not have any. Alt text I check against the image. And code, which is not content.

Where I do not, and this is a preference rather than a policy: anything that constitutes the gain. If a model could write the paragraph, the paragraph is not the reason the page exists, and the paragraph that is the reason cannot be written by anything that was not there. That is not a rule I can defend from a document. It is what falls out of taking the rest of this page seriously.

Core Chapter 31 maintain

Human verification, and the one rule that makes drafting safe

If you remember one thing One rule makes AI drafting safe: nothing ships that a human has not checked against a source they opened. Every AI content disaster is a verification failure wearing a technology costume.

Not in this path.

One rule makes AI-assisted drafting safe: nothing ships that a human has not checked against a source they opened. Every AI content disaster is a verification failure wearing a technology costume.

The reason models are dangerous in content work is not that they lie. It is that they are fluent, and fluency reads as confidence. A wrong statistic delivered in the same register as a right one passes every quality check a human editor runs, because the quality checks were designed to catch bad writing and the writing is fine.

So the check has to be mechanical rather than editorial. Five things get verified, every time, by a person:

  1. Every number. Opened, at the source, this week. Not recognised as plausible. Not remembered from somewhere. Opened.
  2. Every quotation. Against the original document, word for word, including whether the person actually said it. Fabricated quotations are the highest-embarrassment failure available and they are common.
  3. Every link. Clicked, and confirmed to be the thing the sentence says it is. A link to a plausible URL that does not contain the claim is worse than no link, because it manufactures the appearance of sourcing.
  4. Every product claim. Against the vendor's current page, with the date noted. Pricing and features move constantly and a model's training data does not.
  5. Every instruction. By doing it. If the page says click Settings then Advanced, somebody opens the product and clicks Settings then Advanced.

The uncomfortable corollary is that verification does not scale, and that is the point rather than a problem to solve. If you cannot verify a hundred pages a week, you cannot responsibly publish a hundred pages a week, and the constraint is a feature: it puts a ceiling on volume that is set by the thing that actually makes content worth reading.

Google's own guidance adds a second habit worth adopting, which is about the reader rather than about accuracy: sharing how a piece of content was created gives people useful context. That does not mean a badge on every page. It means that where automation did something substantial, saying so is more comfortable than being found out.

Awareness Chapter 32 maintain

Where scaled content stops working

If you remember one thing Scaled content abuse is defined by purpose, not by count. The line is not the number of pages. It is whether anybody was going to check them.

Not in this path.

Scaled content abuse is defined by purpose rather than by count. Google's policy is many pages generated for the primary purpose of manipulating rankings and not helping users, and there is no number in that sentence.

Which is genuinely useful, because it means programmatic content is not the problem. A hundred thousand product pages generated from a catalog are fine, and so are location pages that contain real local information, and so is a dataset published as a page per row. Every one of those is many pages from a template and every one of them can be excellent.

The line is whether each page is worth arriving on. Three tests, and any one of them failing puts you on the wrong side:

  1. Does each page contain something specific to it? Not a swapped city name. A different price, a different dataset, a different photograph, a different set of facts that somebody would want.
  2. Would you be comfortable if a human had written every one? If a hundred humans writing these pages would still have produced a hundred worthless pages, the generator is not the problem and stopping using it will not help.
  3. Is anybody going to check them? This is chapter thirty-one applied at scale, and it is the test that actually decides. The point at which nobody is verifying is the point at which the programme has become the thing the policy is about.

Two related policies worth knowing by name, because they catch things people do not think of as scaled content. Expired domain abuse, buying a domain for its history and repurposing it. And site reputation abuse, publishing third-party content on a host site mainly because of the host's established ranking signals, which is the policy behind the coupon sections and sponsored review directories that appeared on news sites.

The implementation of programmatic content is architecture work rather than content work, and it belongs to step five, which has the taxonomy and URL side of it. What belongs here is the decision, and the decision is the same one from chapter six asked once instead of ten thousand times: what does each of these pages contain that nobody could have guessed?

Practical Chapter 33 maintain

Auditing a corpus, with mine as the worked example

If you remember one thing A corpus audit is a count of what is present, not a judgement of what is good. Mine says three tables in eleven posts and zero updates, and the tool that found that is in the repository.

Not in this path.

A content audit is a count of what is present, not a judgement of what is good, and that limitation is what makes it useful. A count can be repeated in six months and compared; a judgement cannot.

Almost every content audit guide starts with traffic, which is the right place to start for deciding what to prune and the wrong place for understanding what you produce. Traffic tells you which pages worked. It does not tell you that you have written eleven posts and three tables, and that fact predicts the traffic better than any per-page report will.

So here is mine, produced by a tool in the repository, on the day this page published, including the four counts that came back at or near zero.

This site’s own blog, measured 23 August 2026

Every post I have published, counted against my own lesson

11 posts, 32,551 words, published between 11 June 2026 and 20 August 2026. Produced by node tools/corpus.mjs, which is a hundred lines of regular expressions in this repository. It counts what is present, not whether it was any good. Every zero in a column that should not be zero is marked.

Sort by
Post WordsH2sTablesImagesVideosSources citedInternal linksDated figures Updated
seo-keyword-research-chatgpt-prompts 5,454 21 1 0 1 0 3 0 never
common-chatgpt-words-to-avoid 4,385 11 0 0 1 1 3 0 never
seo-penalty-recovery-case-study 4,339 14 0 7 2 0 15 5 never
free-ai-seo-gpt-local-seo-assistant 3,618 14 0 4 1 2 9 0 never
windows-server-for-seo 3,493 18 0 8 2 2 11 0 never
how-to-create-custom-gpt-chatgpt 3,338 11 0 5 1 1 10 0 never
gmail-account-for-seo 2,672 15 0 7 1 0 12 0 never
blogger-not-indexing-google-fix 2,642 11 2 0 0 2 2 0 never
internal-linking-tool 1,111 6 0 3 1 3 6 0 never
cheapest-domain-registrars 753 5 0 3 0 1 6 0 never
query-fan-out 746 4 0 1 0 2 3 2 never
Totals 32,551 130 3 38 10 14 80 7 0 of 11

Scroll the table sideways

The rows where the answer is nothing at all

9 of 11

posts contain no table

The format with the highest evidence density per pixel, and I produced three in eleven posts.

3 of 11

posts cite no external source

Not one link out to anybody. On three of them the reader has only my word for everything.

9 of 11

posts carry no dated figure

A number without a date is an assertion. Two posts have one and one of those two is the case study.

11 of 11

posts have never been updated

Zero updatedDate values in the whole collection, on a site about to spend three chapters on refreshing.

3 of 11

posts contain no image

Two of the three are the longest posts in the collection.

Publishing this was not a stylistic choice. Chapter one argues that most content fails because nobody asked whether it should exist, and I have thirty-two thousand words that Ahrefs values at eight monthly visits. If you want to run the same audit on your own folder, the tool is at content-machine/tools/corpus.mjs and it takes one argument. Your numbers will be about your site and yours will be the ones that matter.

node tools/corpus.mjs, run against src/content/blog on 23 August 2026
Counts presence, not quality. Every figure is a floor rather than a verdict

The four counts at the bottom are the ones this lesson argues about and every one of them is bad. 9 of 11 posts contain no table, on a page with a chapter calling tables the most underproduced format on the web. 3 cite no external source at all. Nine carry no dated figure. And 11 of 11 have never been updated, on a page with three chapters on maintenance.

What a corpus audit is for, as distinct from a page audit: it tells you what your operation can produce. Mine says I produce long tutorials with screenshots, embed the matching YouTube video, cite almost nobody, and never come back. That is a coherent and quite limited machine, and no per-page report would ever have told me, because every one of those posts is individually fine.

The four counts to run on your own folder are in the exercise below, and none of them needs a tool. Posts with no table. Posts citing nobody. Posts with no dated figure. Posts never updated. Put today's date at the top and keep the file, because the second reading is the only one that means anything and it is worthless without the first.

Awareness Chapter 34 maintain

The mistakes I would kill first

If you remember one thing Twelve claims about content, four of them stated most confidently by people who cannot source them, and three of those four arrived in the brief for this page.

Not in this path.

Twelve claims I have been told about content, several of them by people who know more than me, and the ones with no source anywhere are the ones stated with the most confidence.

The kill list

Twelve things people believe about content

Verdict first, then the reason, then where the reason comes from. 10 of the twelve cite a primary source and the rest are answered by measurements of this site. Most of the citations are Google’s own documentation, and on this subject the documentation says something noticeably less exciting, and noticeably more useful, than the tutorials built on top of it.

  1. 01

    ""Longer content ranks better.""

    No, and Google put the answer in a parenthesis.

    The helpful content guidance lists writing to a word count among the signs of search-engine-first content, and answers its own question: "(No, we don’t.)" Length correlates with rankings because thorough pages tend to be longer, and the correlation gets sold as a lever.

    Creating helpful, reliable, people-first content Google Search Central, Read 23 August 2026. Google states last updated 10 December 2025.

  2. 02

    ""Publish consistently and the traffic follows.""

    The distribution says otherwise.

    96.55% of roughly 14 billion pages get zero organic traffic. Volume is the input that everybody can supply, which is precisely why it is not the input that decides anything.

    96.55% of content gets no traffic from Google Ahrefs, Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.

  3. 03

    ""Google penalizes AI content.""

    Not for being AI.

    "Appropriate use of AI or automation is not against our guidelines." The policy is about generating many pages primarily to manipulate rankings. Plenty of AI content is spam. It is spam for reasons that would apply to a human who wrote it the same way.

    Google Search’s guidance about AI-generated content Google Search Central Blog, Read 23 August 2026. Posted February 2023 by Danny Sullivan and Chris Nelson.

  4. 04

    ""Chunk your content so AI can extract it.""

    Google published the opposite in July 2026.

    "There’s no requirement to break your content into tiny pieces for AI to better understand it." Same paragraph: there is no ideal page length. This one is being sold hard by people who have not read the document they are citing.

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  5. 05

    ""Reverse-engineer fan-out queries and build a page for each.""

    Named in the guidance as a spam policy violation.

    Google names fan-out queries directly and says creating separate content for every variation primarily to manipulate rankings violates the scaled content abuse policy, and is an ineffective long-term strategy besides. Understanding fan-out is useful. Farming it is not.

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  6. 06

    ""Update the date and Google treats it as fresh.""

    It treats it as a lie you told twice.

    Freshness is a property of the query, not a property of your page. Changing a date without changing the content is the one intervention that reliably does nothing, and on a query that does not deserve freshness there was nothing to gain anyway.

    Is freshness an important signal for all sites? Google Search Central, YouTube, Watched 23 August 2026. Paired with "Query deserves freshness. Fact or fiction?" on the same channel.

  7. 07

    ""Thin content means short content.""

    Thin means no added value, at any length.

    Google’s own quality material treats thin content as content with little or no added value, and a four-thousand-word restatement of the top three is thin. The word count is not the diagnosis.

    Search Quality Rater Guidelines, sections 4.6.5 and 4.6.6 Google, Read 23 August 2026. Cited by Google’s own generative AI content guidance as the reference for scaled content abuse and for main content created with little effort, originality or added value.

  8. 08

    ""Duplicate content is a penalty.""

    It is a consolidation problem.

    Several URLs answering one need means Google picks one to represent the set, using signals you did not choose. Nothing is being punished. Commodity content is the different and more expensive problem, and it looks nothing like duplication.

    How does Google handle duplicate content? Google Search Central, YouTube, Watched 23 August 2026. Paired with "How can I make the pages on my site unique?" on the same channel.

  9. 09

    ""Add statistics to make content authoritative.""

    Add ones somebody can check.

    A number with no traceable source is decoration with a decimal point in it. Four claims on this page are graded invented and three of them arrived in the brief from an experienced practitioner who believed them.

    Answered from my own data: the corpus audit of this site run on 23 August 2026, plus the eight-guide audit, the twenty keywords and the two results pages pulled the same day. Every row of all of it is shown in full further up this page.

  10. 10

    ""Cover the topic comprehensively and you will win.""

    Comprehensive is the definition of commodity.

    If your page contains everything the top three contain, you have written the average of the results page. Google’s own phrasing for what to avoid is content that "could easily be produced by a generative AI model".

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  11. 11

    ""Information gain is a confirmed Google ranking factor.""

    It is a patent, and that is a different thing.

    Filed 2018, published 2020, assigned to Google. A patent proves somebody thought it was worth protecting. It does not prove the system shipped. Treat it as a good way to think and a bad thing to cite as a ranking factor.

    Contextual estimation of link information gain, US20200349181A1 Google LLC, Read 23 August 2026. Filed 18 October 2018, published 5 November 2020.

  12. 12

    ""AI will write your content for you.""

    It will write the median page on the subject.

    That is what it is for. A model trained on the results page produces the results page, which lands you at level one on the gain ladder with a page that reads well and adds nothing. The parts that cannot be generated are the parts you had to go and get.

    Answered from my own data: the corpus audit of this site run on 23 August 2026, plus the eight-guide audit, the twenty keywords and the two results pages pulled the same day. Every row of all of it is shown in full further up this page.

Verdict first, so the correction cannot be skipped by somebody scanning
The claims with no source anywhere are the ones stated most confidently, every time

And now the grading, which is the part that makes the rest of this page arguable. Every load-bearing claim on it, with a grade for how much weight the evidence bears, including the four with no traceable source anywhere. Three of those four arrived in the brief for this page.

18 claims, graded

Everything this page leans on, and how much weight each one holds

The grade is about the evidence, not about whether I believe it. I believe most of the inference rows and they are still graded inference. The bottom band is the one worth your time: four figures that circulate constantly with no traceable source, and three of them arrived in the brief for this page from somebody experienced who believed them.

7 Documented
4 On the record
3 Inference
4 Invented
  1. Documented

    Google has no preferred word count

    From the helpful content guidance, in a parenthesis: "Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)"

    Google Search Central, Creating helpful, reliable, people-first content
  2. Documented

    Commodity content is a named failure, and Google gave the example

    The AI optimization guide contrasts "7 Tips for First-Time Homebuyers" with "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line", and says commodity content is often based on common knowledge that could originate from anyone.

    Google Search Central, Optimizing your website for generative AI features on Google Search
  3. Documented

    Using AI is not itself against the guidelines

    "Appropriate use of AI or automation is not against our guidelines." What is against them is using automation to generate content primarily to manipulate rankings.

    Google Search Central Blog, Google Search’s guidance about AI-generated content
  4. Documented

    You can ignore chunking your content for Google’s AI features

    "There’s no requirement to break your content into tiny pieces for AI to better understand it." Same document: ignore llms.txt files and inauthentic mentions too.

    Google Search Central, Optimizing your website for generative AI features on Google Search
  5. Documented

    Targeting every fan-out query with its own page is a spam policy violation

    Google names fan-out queries specifically and says doing this primarily to manipulate rankings violates the scaled content abuse policy, and adds that it is ineffective anyway.

    Google Search Central, Optimizing your website for generative AI features on Google Search
  6. Documented

    Scaled content abuse is about purpose, not about volume

    "Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users." The count is not in the policy.

    Google Search Central, Spam policies for Google web search
  7. Documented

    Trust is the most important part of E-E-A-T

    Stated in those words in the helpful content guidance: "Of these aspects, trust is most important."

    Google Search Central, Creating helpful, reliable, people-first content
  8. On the record

    96.55% of pages get no organic traffic from Google

    Ahrefs, 1 December 2023, roughly 14 billion pages in its Content Explorer index. Their index and their traffic estimates, published with the method.

    Ahrefs, 96.55% of content gets no traffic from Google
  9. On the record

    Google scores documents on information gain

    A granted patent, "Contextual estimation of link information gain", filed 18 October 2018 and published 5 November 2020. A patent is evidence that the idea was worth protecting. It is not evidence that the system is live in Search, and anybody telling you it is has gone past the document.

    Google LLC, Contextual estimation of link information gain, US20200349181A1
  10. On the record

    People scan pages in an F-shaped pattern

    Nielsen Norman Group eyetracking, originally 2006, reviewed 19 August 2026. Their own article says the F-pattern is one of several scanning patterns rather than the only one, which is the half nobody quotes.

    Nielsen Norman Group, F-Shaped Pattern of Reading on the Web
  11. On the record

    Long context degrades in the middle

    Published research on language models retrieving information from long inputs. It is a finding about model behaviour on long prompts, not a documented statement about how Google ranks a web page, and the leap between those two is where the advice gets sold.

    Liu et al., Transactions of the ACL, Lost in the Middle: How Language Models Use Long Contexts
  12. Inference

    Adding a comparison table increases the chance of being cited by an AI system

    Plausible, widely repeated, and nobody has isolated it. Tables also help readers, which is the reason to build one. If somebody tells you they measured this, ask what the control was.

  13. Inference

    Answering in the first screen reduces bounce and improves rankings

    The first half is a reasonable claim about readers. The second half is a leap. Google has never documented a ranking use of return-to-results behaviour and has repeatedly said clicks are not a straightforward signal.

  14. Inference

    Updating an old post reliably recovers traffic

    It works often enough that everybody has a story. Almost every published example also changed the title, the internal links and the depth at the same time, so the variable is not isolated in a single public case I can find.

  15. Invented

    At least 60% of visitors never scroll past the top of a page

    Circulates constantly with no traceable original. Scroll-depth numbers exist, they vary enormously by page type and device, and none of the ones I can trace says this. Arrived in the brief for this page and it is not going on it as a fact.

  16. Invented

    A TL;DR at the top lifts conversions by up to 33%

    No source anywhere. A TL;DR is still a good idea for reasons chapter fourteen gives, and those reasons do not need a number that nobody measured.

  17. Invented

    Expert quotes get you cited by AI Overviews within two hours of indexing

    A specific, testable, unsourced claim of exactly the kind this lesson exists to teach you to refuse. Anybody who has this timing has a screenshot, and nobody has published one.

  18. Invented

    Content decays on a predictable schedule and needs refreshing every six months

    There is no schedule. There are four different failures wearing one word, and three of them are not fixed by a refresh at all. Chapter twenty-seven.

Filter to Invented and read the four. Every one of them is specific, testable and stated with total confidence in places you would trust, and none of them has an original anybody can find. That is the failure mode this whole lesson is about, and it does not happen because people are dishonest. It happens because a number gets repeated until it sounds measured, and because nobody wants to be the one who deletes it.

Every documented row is quoted with the sentence intact and dated
If you can source one of the invented four, send it to me and this table changes
Practical Chapter 35 maintain

Test yourself

If you remember one thing Sixteen questions. Every answer is either quoted from a document or measured on this site.

Not in this path.

Sixteen questions. Six of them quote a sentence Google published, four are answered by measurements of this site or of the eight guides above, and none of them is a definition you could have looked up.

Test yourself

Sixteen questions, and every answer is quoted or measured

Every answer is sourced to Google’s own documentation, or to a measurement of this site you can reproduce with a free tool. Four of the sixteen have Google’s sentence quoted verbatim in the explanation. If an answer disagrees with something you were taught, check the source before you check me.

  1. 01 Google’s guidance uses one word for content built from common knowledge that could have originated with anyone. What is it?
    Show the answer

    Correct B. Commodity content

    Commodity content, from the AI optimization guide. It contrasts "7 Tips for First-Time Homebuyers" with a first-hand piece about waiving a home inspection. Thin, duplicate and scaled are three different failures with their own definitions.

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  2. 02 What proportion of pages get no organic search traffic from Google, in the largest published study?
    Show the answer

    Correct C. About 96.55%

    Ahrefs, 1 December 2023, roughly 14 billion pages. A further 1.94% get between one and ten visits a month, which is the band this site was in on the day this page published.

    96.55% of content gets no traffic from Google Ahrefs, Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.

  3. 03 Google has a documented position on ideal word count. What is it?
    Show the answer

    Correct C. It has no preferred word count and says so in a parenthesis

    "Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)" It appears in the list of signs that content was made for search engines rather than people.

    Creating helpful, reliable, people-first content Google Search Central, Read 23 August 2026. Google states last updated 10 December 2025.

  4. 04 What did Google publish in July 2026 about chunking your content for AI features?
    Show the answer

    Correct B. There is no requirement to break content into tiny pieces

    The AI optimization guide says there is no requirement to chunk, that there is no ideal page length, and separately that Google Search ignores llms.txt files entirely.

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  5. 05 On 23 August 2026, how many of the nine pages ranking organically for "seo content" had that phrasing as their own top keyword?
    Show the answer

    Correct D. None

    None. Nine results, nine different primary targets. Three of the pages on the two results pages behind this lesson have "content seo" as their top keyword instead, which is half the volume and is why this page lives at /content-seo/.

  6. 06 Google names one thing you should not do with fan-out queries. What is it?
    Show the answer

    Correct B. Create separate content for every variation, primarily to manipulate rankings

    The AI optimization guide names fan-out queries specifically and says doing this violates the scaled content abuse policy, and adds that a high quantity of pages does not make a site higher quality anyway.

    Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.

  7. 07 What is the status of "information gain" as a Google mechanism?
    Show the answer

    Correct B. A granted Google patent, filed 2018 and published 2020

    "Contextual estimation of link information gain", assigned to Google LLC, filed 18 October 2018, published 5 November 2020. A patent proves the idea was worth protecting. It does not prove the system is live in Search.

    Contextual estimation of link information gain, US20200349181A1 Google LLC, Read 23 August 2026. Filed 18 October 2018, published 5 November 2020.

  8. 08 A page’s position is stable, its impressions are falling, and nothing on the results page has changed shape. Which decay is this?
    Show the answer

    Correct D. The demand left

    Stable position plus falling impressions means fewer people are searching, or the answer now arrives without a click. Refreshing it changes nothing, because your page is not what got worse.

  9. 09 Which of these is a spam policy violation according to Google’s published policy?
    Show the answer

    Correct C. Generating many pages primarily to manipulate rankings without helping users

    That is the scaled content abuse definition almost verbatim. The other three are cases Google either names as useful or does not object to. The policy is about purpose, not about the tool.

    Spam policies for Google web search Google Search Central, Read 23 August 2026. Google states last updated 15 May 2026.

  10. 10 Across eight published guides on this subject, audited by their own headings, what proportion of sections were actually about content?
    Show the answer

    Correct C. A little under a half

    65 of 147 sections, which is 44%. The largest single competitor for that space is on-page SEO at 36 sections. Three of the eight guides are majority something else, and Yoast’s, with the strongest title claim in the set, is 7 of 32.

  11. 11 Which of these counts as strong evidence on a page?
    Show the answer

    Correct B. A number you measured, with the tool named and the date attached

    The axis is checkability by a stranger. A measurement with a named tool and a date can be reproduced and can be wrong in public, which is exactly what makes it worth something.

  12. 12 What did Google say about the relationship between how content is produced and how it is judged?
    Show the answer

    Correct B. Its focus is on the quality of content rather than how content is produced

    From the February 2023 post, restated since: "Our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years."

    Google Search’s guidance about AI-generated content Google Search Central Blog, Read 23 August 2026. Posted February 2023 by Danny Sullivan and Chris Nelson.

  13. 13 This site’s own blog was audited on the day this page published. How many of its eleven posts had ever been updated?
    Show the answer

    Correct D. None

    None. Zero updatedDate values in the whole collection, along with three tables across eleven posts and three posts citing no external source at all. The tool that produced those numbers is in the repository.

  14. 14 Which of these is the correct reason to add a comparison table to a page?
    Show the answer

    Correct B. It is the densest way to let a reader reach a decision

    The citation claim is graded inference on this page and nobody has isolated it. The reader reason is sufficient on its own, which is lucky, because it is the only one with anything behind it.

  15. 15 A results page for your target query is nine tools and no articles. Your page is an article. What does content SEO say?
    Show the answer

    Correct C. Change format or stop

    Format is decided by the results page, and no length of article gets into a market that wants an input box. This is the intent-moved decay shape and the fix is not editorial.

  16. 16 What makes a section extractable?
    Show the answer

    Correct B. It survives being lifted out of the page and still makes sense

    A passage that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere. The heading question and the word count are style choices; standing alone is the actual property.

The six decisions, run on the page you are reading

Not a hypothetical and not a client. Eight decisions taken while building this page, three of them with a price attached, and the last one ending with a question still open.

Not in this path.

The worked example

The six decisions, run on the page you are reading

Eight decisions, in the order they were taken, with the evidence behind each. 4 of them cost something and the cost is printed. The last one ends with a question still open, because it is still open and closing it in public would be the kind of tidiness this whole lesson argues against.

  1. 01

    Deserve

    Should this page exist, when eight good guides already do?

    Audited all eight from their raw HTML before writing anything. 65 of 147 leaf sections are actually about content, so less than half of the average guide titled for content is about content. That gap is the reason to build. The first pass of this count was wrong, because I classified from summaries rather than from the markup, and the corrected numbers are less flattering to my argument than the first ones were.

    The evidence GUIDES in data/content-seo.ts. Eight URLs fetched 23 August 2026, headings extracted, chrome removed, leaf sections classified with the rule printed above the chart.

  2. 02

    Deserve

    Which URL, given one phrasing has twice the volume?

    Chose /content-seo/ over /seo-content/. "seo content" carries 3,600 US searches to "content seo"’s 1,800, and not one page ranking for the bigger phrasing has it as its own top keyword, while three have "content seo". The market’s own pages were built for the smaller one.

    The evidence Two SERP Overview pulls, 23 August 2026. Both queries return an identical nine-source AI Overview, so they are one need with two vocabularies.

    What it cost On paper this declines 1,800 monthly searches. In practice the bigger phrasing rolls up to the parent topic "seo" and is currently won by Google’s beginner guide, so the volume was never available.

  3. 03

    Know

    What will this page contain that the other eight do not?

    Three things, named before drafting: the guide audit, this site’s own corpus measured with a tool in the repository, and every load-bearing claim graded including four graded invented. Two of the four invented claims came from the brief for this page.

    The evidence CLAIMS, CORPUS and GUIDES. Level four, three and four on the gain ladder respectively.

  4. 04

    Know

    Publish the domain traffic number, or leave it out?

    Published it. Three organic keywords and eight estimated monthly organic visits on the day this went live, which puts this domain in the 1.94% band of the study the page opens with.

    The evidence Ahrefs Site Explorer, alstonantony.com, US, 23 August 2026.

    What it cost It is the worst number on the site and it is on a page about content quality. Leaving it out was the obvious call and would have made the opening argument dishonest.

  5. 05

    Satisfy

    What is the first screen?

    The answer to the core question, in the first sentence, with the two numbers that make it land: 96.55% and the eight visits. No definition of content SEO above the fold, because the definition is not in dispute and eleven other pages already have it.

    The evidence The TLDR block at the top of this page, written before the chapters.

  6. 06

    Show

    What is on the page besides prose?

    Sixteen interactive or graphic components, two of which carry the original data. The guide audit and the corpus audit would both still be worth something with every paragraph deleted, which is the test.

    The evidence The component list in src/components/content-seo/.

  7. 07

    Structure

    Chunk the content for AI, as the industry advises?

    No. Google published on 10 July 2026 that there is no requirement to break content into tiny pieces, that there is no ideal page length, and that Search ignores llms.txt. The retrieval chapter says so and quotes it, which cost this page the section every competing guide now has.

    The evidence Google’s AI optimization guide, read 23 August 2026.

    What it cost Refusing the popular advice on a page about content is a traffic decision as much as an editorial one. The sections still get written to stand alone, because that helps readers, which is a reason that does not need a ranking claim.

  8. 08

    Maintain

    What would make me delete this page?

    Two conditions, written before publication. If Google retires the commodity content framing, the spine loses its foundation. If the corpus audit stops being unflattering, the central receipt stops being evidence and becomes a boast.

    The evidence The second one is a real risk and it is the good kind: it means the site improved.

    What it cost The unresolved decision: "seo content strategy" carries 3,200 searches at difficulty 20 with its own parent topic, which by this site’s own rules in step five means a separate page. It is queued and not built, and saying so is cheaper than pretending this page covers it.

The reason this is the page rather than a client site: a worked example on somebody else’s property lets the author arrange the lesson so it comes out right. This one had to be published with its own traffic number attached, and that number is eight visits a month. If the method were being demonstrated on a site that already worked, you would have no way of telling whether the method or the site was doing the work.

Eight decisions, three with a price on them, one unresolved
Every figure here is reproducible from the sources at the bottom of this page

Now build the thing this page exists to produce

Not in this path.

Reading this page does not make you able to do any of it, and I would be selling you something if I implied otherwise. This next block is the part that does, and it produces one document: a brief with a field for the thing no brief template asks for.

The deliverable

Build a brief that can say no

Thirteen fields, in decision order, and two of them do not appear on any brief template I have seen. The gain field says what this page will contain that the current top three do not. The kill field says what would make you delete it. Leave the first one empty and this is not a brief, it is a topic, and the builder will keep telling you so.

Load an example
Deserve

One. From step three, confirmed against step four.

What they are trying to finish. Not what they are interested in.

What this page does for you. "Traffic" is not an answer.

Know

What this page will contain that the current top three do not. Empty means no brief.

The specific artefact, and where it comes from.

From the ladder. Below three, say why that is acceptable.

Satisfy

From the results page count, not from the calendar.

The actual sentence. Write it now, not during the draft.

Show

Table, chart, screenshot, tool, dataset. At least one.

Structure

Each one says what it answers, and each survives being lifted out.

Tools, versions, dates, people, documents. Not "a popular tool".

Maintain

A condition, not a date. "Quarterly" is a calendar, not a trigger.

The condition under which this page stops deserving to exist.

Everything you type stays in this browser and nowhere else. Nothing is sent anywhere, which matters on this form more than on the drills above it, because the gain field is the thing you know and nobody else does.

Three worked briefs, and one of them should have been killed

A page worth building Level three gain, named evidence, a first-screen sentence written before the draft, and a condition that would kill it.
Primary query
content decay
The reader’s job
Work out why one of their pages is falling, before spending a day on the wrong fix
The business job
Routes to the refresh service page, and proves I diagnose before I bill
The gain, in one sentence
Four named decay shapes, each with the Search Console pattern that identifies it, from twelve real diagnoses. Nobody has published the diagnostic; everybody has published the definition.
The named evidence
Twelve anonymised GSC charts with dates, plus my own corpus audit from 23 August 2026
Gain level
3, first-hand
Format
Explainer with a diagnostic table. 8 of 10 results are guides, so the format is not in dispute.
The first screen answer
Content decay is four different failures sharing one word, and three of them are not fixed by a refresh.
What is on the page besides prose
A four-row diagnostic table, four sparkline shapes, one annotated GSC screenshot
Section headings
What decay is not · The four shapes · Which one you have · What to do about each · The one that means stop
Things named specifically
Search Console, Ahrefs Site Explorer, Google’s "Is freshness an important signal" video, 23 August 2026
Review trigger
When GSC impressions fall 25% over 90 days, or when Google publishes anything new on freshness
What would make me delete this
If I lose permission to publish the client charts, the gain is gone and the page becomes a definition nobody needs. Delete it and fold the diagnostic into the refresh lesson.
A brief that should have been killed It looks complete. The gain field is a description of thoroughness, which is the tell.
Primary query
seo content
The reader’s job
Learn what SEO content is
The business job
Traffic and brand awareness
The gain, in one sentence
A more comprehensive and better structured guide than the current results, covering every subtopic in more depth.
The named evidence
Research from the top 10 results and industry best practices
Gain level
1, better arrangement
Format
Long-form guide, 3,000+ words
The first screen answer
SEO content is content designed to rank in search engines.
What is on the page besides prose
Stock images and a summary infographic
Section headings
What is SEO content · Types of SEO content · Why it matters · How to create it · Tools · FAQ
Things named specifically
Google, Ahrefs, Semrush
Review trigger
Quarterly
What would make me delete this
empty
A small page with real gain Three hundred words, one screenshot, and the only page on the internet that has it. Level three does not require a project.
Primary query
cloudflare sitemap 403 googlebot
The reader’s job
Stop their sitemap returning 403 to Googlebot today
The business job
Proves technical depth to exactly the reader who hires me
The gain, in one sentence
The specific Cloudflare rule that caused it, the request header dump from both a browser and Googlebot, and the two fixes that did not work.
The named evidence
My own header dump and the Search Console error, both dated
Gain level
3, first-hand
Format
Troubleshooting page. 6 of 10 results are forum threads, which means nobody has written the page.
The first screen answer
A Cloudflare managed rule was returning 403 to Googlebot on /sitemap.xml while every browser got a 200. Here is the header that proves it and the rule to change.
What is on the page besides prose
Two header dumps side by side and the Search Console error screenshot
Section headings
The symptom · Proving it is Cloudflare and not your CMS · The rule · What did not work
Things named specifically
Cloudflare WAF managed rules, Googlebot user agent, Search Console Page Indexing report
Review trigger
When Cloudflare changes its managed ruleset naming
What would make me delete this
If Cloudflare fixes the default behaviour, this page describes a problem that no longer exists. Delete it.

Read the second one carefully. Every field is filled in, it would pass any editorial review, and it is a commodity page: the gain field describes thoroughness rather than a finding, the evidence field says "research from the top 10 results", and the kill field is empty because nothing about that page could ever stop being true. It is the most common brief in the industry and it is the one that produces the 96.55%.

Then the seven assignments. One per decision, in the order the decisions come, and each one produces something real rather than a note in a document.

The seven assignments

One assignment per decision

Bigger than the exercises and meant to be done once, properly. Finish all seven and you have a content operation that can name what its next page will contain before it is written, and that has measured what its last twenty contained.

Ticks are stored in your browser and nowhere else. Nothing is sent anywhere, which matters because four of these ask you to write down where your own content is weak.

And if you would rather have a schedule than a list, the same work spread across a week. About half an hour a day, and day seven is the one people skip.

If you want a schedule

The seven-day content audit

Same material, paced. Days one and two measure, day three grades, day four kills, day five produces something that did not exist, day six fixes what you already have, and day seven writes one brief that could say no.

Put the site name and the date at the top of whatever you end up with. In six months it is the only thing that will tell you whether your content changed or your standards did.

Then the checkpoint. Eleven questions, and if you can answer all eleven about your next page you are ahead of every content plan I have been shown.

Before you call it finished

The content checkpoint

Answerable yes or no about a real page. Any no is a work item, and the last one is the one that catches people who have done everything else.

Question eleven is the honest test. If nothing could ever make you delete the page, you have not defined what it is for.

The six decisions, one more time

Not in this path.

If you keep one thing from this page, keep these. Every tactic, framework and tool in content SEO serves one of these six decisions, and any advice that does not answer one of them is decoration.

The spine

Deserve, Know, Satisfy, Show, Structure, Maintain

In order, and the order matters. The first two can cancel the other four, which is why they are first and why almost nothing published on this subject starts there. Each card ends with what you are actually holding when the decision is made.

  1. 1

    Deserve

    Should this page exist at all?

    The only decision on this list that can save you the whole cost of the page, and the one no content calendar has a column for. Google now has a word for the failure: commodity content, meaning something built from common knowledge that could have originated with anyone. If the honest answer to "what does this add" is "it covers the topic", the page is a retelling and the results page already has six.

    You end up holding A shorter list, with the rows you killed still visible and a reason next to each.

  2. 2

    Know

    What do you have that a model could not have guessed?

    Named before a word is written, not discovered during the draft. A number you measured, a screenshot of something you ran, a decision you made and regretted, an export, a named source with a date. Google filed a patent on scoring exactly this and calls it information gain. If you cannot name the thing, you are about to write a summary of the current top three.

    You end up holding One named piece of evidence per page, with where it comes from and who can check it.

  3. 3

    Satisfy

    Does the first screen answer the thing that was asked?

    The brief arrived from step four with a page type and a job. This decision is whether the job gets done immediately or gets withheld until paragraph nine. The failure is not length. It is sequence: the answer exists on the page and arrives after the reader has already gone back.

    You end up holding A first screen a stranger can read in twenty seconds and leave satisfied.

  4. 4

    Show

    Is the evidence visible, or only claimed?

    Tables, screenshots, diagrams, video, a dataset, a working tool. Not decoration and not "add images for SEO". Every one of these is a claim being made checkable, and the reason this decision sits fourth is that you cannot illustrate evidence you have not got. Google states that its generative AI features can bring in relevant images and video, which is opportunity rather than obligation.

    You end up holding At least one thing on the page that would still be worth something with the prose deleted.

  5. 5

    Structure

    Can a scanner navigate it and a machine extract it?

    Headings that say what a section answers, sections that survive being lifted out of the page, and things named specifically enough to be recognised. This is the decision most often oversold: Google published, in July 2026, that you can ignore chunking your content for AI, and the industry sold chunking anyway.

    You end up holding A heading outline a stranger could read alone and know what the page covers.

  6. 6

    Maintain

    Does it survive contact with next year?

    Freshness, decay, updating, pruning, and the audit that tells you which. This is the decision that makes the previous five compound instead of accumulate, and it is the one I am worst at: on the day this published, none of my eleven posts had ever been updated.

    You end up holding A dated audit of your own corpus, with a verb on every row.

Decisions one and two are the ones no content calendar has a field for
Everything published on this subject starts at three, because three is the teachable one

The test I would give a beginner for evaluating any advice about this subject, including mine: ask which of the six it answers, and ask what it would tell you not to publish. A method that only ever produces more pages is a content calendar wearing a method's clothes.

And the one sentence, if the whole page has to reduce to one. Before you write anything, name the one thing it will contain that a model could not have guessed, and if you cannot, do not write it. That sentence is free, it takes ten seconds, and applied honestly to your next ten briefs it will delete three of them. The three it deletes are the three that were going to end up in the 96.55%.

Next comes on-page SEO: how a page that deserves to exist communicates what it is, now that you know what is on it. That lesson is being written. In the meantime the roadmap has the rest of the sequence, the architecture lesson has the link graph audit that tells you whether anything points at the pages you are about to improve, and the technical SEO checklist covers the implementation half of chapters nineteen and twenty-five.

Glossary

Every word this lesson uses, defined once, in the plainest phrasing that is still correct. Where a term is normally taught with a ranking promise attached, the promise has been left out.

Not in this path.

Reference

Every word this lesson uses, in plain terms

Several are defined against Google’s own documentation rather than the industry’s retelling of it, which on this subject changes the definition and not just the wording. Where a term is usually taught with a ranking claim welded on, the claim has been left out on purpose, and the three terms Google itself defines are marked.

27 terms

Commodity content Google’s word
Google’s term for content built from common knowledge that could have originated from anyone, and typically adds little unique insight. Its published example is a generic first-time-homebuyer tips post.
Content brief
The document that decides a page before it is written. A brief that cannot say no is a topic with extra fields.
Content decay
A page losing traffic over time. Four distinct failures share the word: staleness, being outclassed, the intent moving, and the demand leaving.
Content pruning
Deliberately deleting or consolidating pages that no longer deserve to exist. The verb most content plans have no column for.
Content SEO
Deciding what information should exist and producing it. Distinct from on-page SEO, which makes one page communicate its relevance clearly.
Duplicate content
Several URLs answering one need, which makes Google pick one to represent the set. A consolidation problem, not a penalty.
E-E-A-T
Experience, expertise, authoritativeness and trustworthiness. A set of things quality raters look for, not a score in the algorithm. Google states trust is the most important of the four.
Evergreen content
Content whose usefulness does not depend on when it was published. A property of the subject, not a writing technique.
Extractability
Whether a section still makes sense when lifted out of the page. Useful for readers, quotes and retrieval systems alike.
Fan-out query Google’s term
One of the concurrent related queries a generative system issues behind a single prompt. Google defines it and separately warns against building a page for each one.
First-party evidence
Something you measured, ran, bought or decided. The category a language model cannot generate, which is what makes it the whole of the Know stage.
Freshness
A property of a query rather than of a page. Some queries deserve recent results and most do not, and no date change makes a query deserve freshness.
GEO
Generative engine optimization. Google’s own position is that optimizing for generative AI search is optimizing for search, and thus still SEO.
Grounding RAG
Retrieval-augmented generation. Google describes it as relying on core Search ranking systems to retrieve relevant pages, which the model then reviews to build a response.
Helpful content
Google’s framing for content created primarily for people rather than to manipulate rankings. The self-assessment questions are published and worth answering honestly once.
Information gain
How much a document adds beyond what the reader has already seen. Google filed a patent on scoring it. A patent is not a confirmed ranking system.
Intent
What would satisfy the searcher. Decided in step four of this roadmap and handed to this step as a page type with a count behind it.
Non-commodity content
Google’s counterpart to commodity content: unique expert or experienced takes that go beyond common knowledge, which a generative model could not easily produce.
On-page SEO
Making one page communicate its relevance: titles, headings, placement, links, markup. The step after this one, and the one this subject is most often confused with.
Original research
Producing a dataset nobody has. It does not require scale. Counting the section headings of eight guides is original research and takes an afternoon.
Pogo-sticking
A searcher clicking a result and immediately returning to the results page. A real behaviour, a badly documented ranking signal, and a good reason to write a better first screen anyway.
Retrieval
The step where a system selects passages to answer with. What makes a passage retrievable is that it stands alone, not that it is short.
Scaled content abuse Google’s policy
Google’s spam policy for many pages generated primarily to manipulate rankings rather than to help users. Defined by purpose, not by page count.
Site reputation abuse Google’s policy
Third-party content published on a host site mainly because of the host’s established ranking signals. A separate policy from scaled content, and often confused with it.
Thin content
Content with little or no added value, at any length. A four-thousand-word restatement of the top three is thin, which is why word count is the wrong diagnostic.
TL;DR
A short summary at the top of a page. A good idea for readers. The conversion numbers attached to it in circulation have no traceable source.
Topical authority
The idea that covering a subject broadly makes a site more likely to rank within it. Widely believed, not documented, and graded inference wherever this site uses it.

Watch the people who own the surface

Not in this path.

Where I would send you instead of a course. All free, filterable by channel, ordered by the chapter they belong to. Every channel here either owns a search surface or owns a database that measures one, and there are no independent SEO channels on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. That rule costs something here, because several of the best explanations of information gain and content pruning on YouTube are by individual practitioners and none of them is in this list.

Watch the people who own the surface

110 videos, 7 channels, ordered by chapter

Every channel here either owns a search surface or owns a database that measures one. No independent SEO channels, on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. Filter by channel, or read it top to bottom as a syllabus. Every id was checked against the YouTube oEmbed endpoint on 23 August 2026 and the channel shown is the channel oEmbed returned. All 118 candidates passed that check; the ones not in this list were cut for relevance rather than for failing it.

25 What Google published about AI features, and what it told you to ignore

Every id verified against the YouTube oEmbed endpoint on 23 August 2026
Two of the seven channels sell an SEO tool and two more own an AI answer surface. Watch them knowing that

What this lesson deliberately refuses to teach

Not in this path.

Ten subjects belong elsewhere, and cramming them in here would produce exactly the failure chapter two measured: a guide titled for content that is mostly about something else. Every item below is real, learnable, and worth nothing to somebody who cannot yet name what their next page will contain that nobody else has.

On-page SEO

Titles, headings, keyword placement, alt text, internal link anchors and structured data. It is the next step and it is a different job: this page decides what should exist, that one makes it legible.

Keyword research

Where demand comes from and how to size it. It is step three and it is finished before this page starts, which is exactly the boundary six of the eight audited guides fail to hold.

Step 3

Search intent and SERP analysis

Reading a results page to decide what kind of page deserves to exist. Step four hands this page a brief with a page type on it.

Step 4

Site architecture and internal linking

Where the page lives and what points at it. Step five, and the reason it comes first is that a page with no parent and no inbound link is a content problem you cannot write your way out of.

Step 5

Content marketing and distribution

Promotion, email, social, outreach. Publishing is not distribution and this lesson stops at publish. The surface map from step one is where distribution belongs.

Step 1

Copywriting and conversion

Persuasion, offers, calls to action, landing page structure. "seo copywriting" carries its own parent topic and a difficulty of zero, which by this site’s own rules means it is a separate page rather than a section here.

Editorial process and team workflow

Briefs at scale, freelancer management, editing pipelines, calendars. Real, learnable, and worth nothing to somebody who cannot yet say what their next page will contain that nobody else has.

Technical SEO for content

Rendering, indexing, canonical tags, pagination. Step two covers the mechanics and a content lesson that drifts into rendering has stopped being a content lesson.

Step 2

Programmatic SEO

Generating thousands of pages from a dataset. Chapter thirty-two states where the line is and refuses the implementation, because the implementation is an architecture problem and the line is the part people get wrong.

AI writing tool comparisons

Which tool to use. The tools change every quarter and the rule in chapter thirty-one does not, so the rule is here and the tools are in the directory.

The directory

The learner only needs six things from this page. A page earns its place by containing something a model could not have guessed. The evidence is named before the draft starts. The answer goes in the first screen. The format is decided by counting the results page. Every section has to survive being lifted out of it. And the audit of what you already have is more useful than any advice about what to write next.

Page history

What changed on this page, and when, because the central receipt on it is a measurement of my own content and it moves every time I publish or update anything.

Not in this path.

Page history

What changed, and when

The central receipt on this page is a measurement of my own content, and it will move the moment I publish or update anything here. So the changes get listed rather than absorbed silently into an "updated" stamp, and the next corpus run appears here as its own entry with the new numbers in it, including the ones that have not improved.

  1. 23 August 2026

    First publication

    • Published, as step six of the SEO roadmap.
    • Eight competing guides audited from their raw HTML: 147 leaf sections, 65 of them about content, and on-page SEO the largest non-content bucket at 36. Three of the eight are majority something other than content.
    • This site’s own blog measured with tools/corpus.mjs: 11 posts, 32,551 words, 3 tables, 14 external sources, 7 dated figures, 0 updates.
    • This domain’s own Ahrefs reading published: 3 organic keywords, 8 estimated monthly US organic visits.
    • Twenty keywords and two results pages pulled from Ahrefs, US database. Four of the twenty roll up to the parent topic "seo".
    • Eighteen claims graded. Seven documented, four on the record, three inference, four invented. Three of the four invented ones arrived in the brief for this page.
    • URL decision published: /content-seo/ over /seo-content/, against twice the volume, with the reason and the cost printed.

If you find something here that a current Google, Ahrefs or Semrush document contradicts, that is a bug in this page and I would rather hear about it than have it quietly rot. The contact page works.

Sources

Every claim above, traced to the document it came from, with the date it was read and the date the document says it was last updated. One of the entries is a tool in my own repository.

Not in this path.

Everything on this page, sourced

22 sources, each read on 22 August 2026. Where a document prints its own last-updated date, both dates are shown, because the document’s own date is the one that tells you whether the guidance has moved since I quoted it. Where an article has an author, the author is named rather than the brand.

Google Search Central 5

Google Search Central Blog 1

Google 1

Google Search Central, YouTube 2

Ahrefs 3

Semrush 2

Yoast 1

Backlinko 1

AIOSEO 1

  • SEO Content Read 23 August 2026. Two named case studies with before and after traffic figures.

Serpstat 1

Google LLC 1

Nielsen Norman Group 1

Liu et al., Transactions of the ACL 1

This repository 1

  • corpus.mjs, the content audit tool Run 23 August 2026 against src/content/blog. A hundred lines of regular expressions in content-machine/tools/. It counts presence, not quality.

Eight of these are Google’s own documentation and seven of them are quoted verbatim on this page, with the sentence intact rather than a paraphrase, because the paraphrases in circulation are how "chunk your content for AI" ended up being sold out of a document that says there is no requirement to do it. Two entries are documents I read and deliberately would not lean on as ranking evidence: a Google patent, and a paper about how language models handle long inputs. Both are graded on the record rather than documented in the claims table. If one of these sources now says something different from what I have written, the source wins and this page is out of date.

The last entry is mine and it is the fastest-rotting thing here. The corpus audit is whatever ’node tools/corpus.mjs’ said on 23 August 2026, and it changes the moment I publish or update anything, which is the point: the numbers on this page are bad and they are supposed to get better. The eight-guide audit will rot too, because six of those eight guides will be revised. The tool is in my own repository rather than in a screenshot, and it is the same one this page tells you to run against your own content. If your numbers disagree with mine, yours are about your site and they are the ones that matter.