SEO roadmap, step 6
Content SEO
What makes a page genuinely worth creating, and useful enough to compete. Six decisions, three datasets built for this page, and my own content measured against every rule in it before yours is.
Jump to a chapter 35
- 01 Commodity content is the whole problem
- 02 Content SEO is not on-page SEO
- 03 Why this comes after architecture
- 04 A page needs two jobs, not one
- 05 96.55%, and what those pages have in common
- 06 The pages you should not write
- 07 Information gain, and the patent
- 08 The four things a model cannot generate
- 09 Research is where the page is decided
- 10 A receipt, and what does not count as one
- 11 Examples, and the invented person rule
- 12 Original research at very small scale
- 13 What the brief decides and what content decides
- 14 The answer goes in the first screen
- 15 Pogo-sticking, and what Google has actually said
- 16 Word count is the wrong unit
- 17 Write for the scanner, then for the reader
- 18 Not every answer is an article
- 19 Visual content, and the four kinds that earn a slot
- 20 Tables and comparisons, the underproduced format
- 21 Video on a page, and what it is actually for
- 22 Every section has to survive being ripped out
- 23 A heading is a contract
- 24 Name things specifically
- 25 What Google published about AI features, and what it told you to ignore
- 26 Freshness is a property of the query
- 27 Content decay has four shapes
- 28 Updating a page, and the three ways it goes wrong
- 29 Duplicate content and commodity content are different problems
- 30 AI-assisted content, against what Google actually published
- 31 Human verification, and the one rule that makes drafting safe
- 32 Where scaled content stops working
- 33 Auditing a corpus, with mine as the example
- 34 The mistakes I would kill first
- 35 Test yourself
What the next hour buys you
What you will actually be able to do
Not "understand content quality". These are the specific things you should be able to do to your own content, by yourself, when you close the tab. If one of them still feels impossible afterwards, that chapter failed and I would rather know which.
- Name, before you write a word, the one thing a page will contain that a model could not have guessed
- Kill a planned page and give the reason out loud
- Tell commodity content from non-commodity content using Google’s own published example
- Say where content SEO stops and on-page SEO starts, and hold the line
- Grade a piece of evidence as strong, usable, weak or worthless before you cite it
- Rewrite a first screen so a stranger gets the answer in twenty seconds
- Pick a format other than "article" for at least a third of your briefs
- Build a comparison table from criteria you chose rather than from a feature list
- Test a section by ripping it out of the page and reading it alone
- Tell which of four decay shapes a falling page has, from the report rather than from a hunch
- State where the AI line is by quoting the policy rather than by guessing
- Audit your own corpus and name the four counts where the answer is nothing at all
Pick a path
Thirty-five chapters is a lot in one sitting. Choosing a path collapses the ones outside it, and you can open any of them anyway.
Your progress
0 / 0
I have been doing this since 2010. I have owned and run more than a hundred sites, lost two of them entirely to Panda and Penguin, bought and tested more than 500 SaaS products with my own money, and taught this to over 30,000 students. I have also published eleven blog posts on this domain since June, totalling 32,551 words, and on the day I wrote this Ahrefs valued the entire domain at 8 monthly organic visits. Both of those facts are relevant and the second one is more useful to you.
Here is the mistake this page exists to stop, and it survives because it looks like diligence. You research the topic, read the top three results, cover everything they cover plus a bit more, and publish something genuinely better organised than what was there. That page is a retelling. Google has a word for it now, published in its own guidance, and the word is not one the industry uses.
"Commodity content (for example, something like '7 Tips for First-Time Homebuyers') is often based on common knowledge, which could originate from anyone, and typically adds little unique insight for readers."
That is from Google's guide to optimizing for generative AI features, read on 23 August 2026 and last updated by Google on 10 July 2026. The same paragraph gives the counterexample, and it is worth reading twice because it is not what a content brief usually asks for: "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line". Not a better guide. A different kind of thing entirely, and one that requires somebody to have waived an inspection.
So the opponent for this whole page is coverage, and I can show you what it costs at scale. Ahrefs studied roughly 14 billion pages in its own index and published the distribution on 1 December 2023: 96.55% get zero organic traffic from Google. A further 1.94% get between one and ten visits a month, and that second band is where this domain currently is. I am not describing somebody else's failure rate.
This is step six of the roadmap and it assumes the five before it. It assumes you know which surfaces your buyers use, from search everywhere optimization, that you can read a crawl report, from SEO fundamentals, that you arrive with clusters rather than a blank page, from keyword research, that you know what kind of page each cluster deserves, from search intent, and that the page has a parent and something linking to it, from site architecture. None is a hard prerequisite. The last two make everything here easier, because they hand this step a brief instead of a blank page.
The spine
The whole lesson, in six decisions
In order, and the order matters. The first two can cancel the other four, which is why they are first and why almost nothing published on this subject starts there. Each card ends with what you are actually holding when the decision is made.
- 1
Deserve
Should this page exist at all?
The only decision on this list that can save you the whole cost of the page, and the one no content calendar has a column for. Google now has a word for the failure: commodity content, meaning something built from common knowledge that could have originated with anyone. If the honest answer to "what does this add" is "it covers the topic", the page is a retelling and the results page already has six.
You end up holding A shorter list, with the rows you killed still visible and a reason next to each.
- 2
Know
What do you have that a model could not have guessed?
Named before a word is written, not discovered during the draft. A number you measured, a screenshot of something you ran, a decision you made and regretted, an export, a named source with a date. Google filed a patent on scoring exactly this and calls it information gain. If you cannot name the thing, you are about to write a summary of the current top three.
You end up holding One named piece of evidence per page, with where it comes from and who can check it.
- 3
Satisfy
Does the first screen answer the thing that was asked?
The brief arrived from step four with a page type and a job. This decision is whether the job gets done immediately or gets withheld until paragraph nine. The failure is not length. It is sequence: the answer exists on the page and arrives after the reader has already gone back.
You end up holding A first screen a stranger can read in twenty seconds and leave satisfied.
- 4
Show
Is the evidence visible, or only claimed?
Tables, screenshots, diagrams, video, a dataset, a working tool. Not decoration and not "add images for SEO". Every one of these is a claim being made checkable, and the reason this decision sits fourth is that you cannot illustrate evidence you have not got. Google states that its generative AI features can bring in relevant images and video, which is opportunity rather than obligation.
You end up holding At least one thing on the page that would still be worth something with the prose deleted.
- 5
Structure
Can a scanner navigate it and a machine extract it?
Headings that say what a section answers, sections that survive being lifted out of the page, and things named specifically enough to be recognised. This is the decision most often oversold: Google published, in July 2026, that you can ignore chunking your content for AI, and the industry sold chunking anyway.
You end up holding A heading outline a stranger could read alone and know what the page covers.
- 6
Maintain
Does it survive contact with next year?
Freshness, decay, updating, pruning, and the audit that tells you which. This is the decision that makes the previous five compound instead of accumulate, and it is the one I am worst at: on the day this published, none of my eleven posts had ever been updated.
You end up holding A dated audit of your own corpus, with a verb on every row.
Those six are the model, and the order is load-bearing in a way most content frameworks are not. Decisions one and two can cancel the other four, and they are the two that no content calendar, brief template or editorial workflow has a field for. Every chapter below sits under one of the six, and the chips on each card jump to the chapters that serve it.
Notice what is not decision one. Not "pick a keyword", which was step three. Not "match the intent", which was step four. Not "structure the page", which is decision five here and is where every competing guide starts. The first question is whether the page should exist, and it is first because it is the only one that can still save you the whole cost.
Watch first, ten minutes
Google, asked whether more content is better, declining to say yes
The question this whole page answers, put to Google directly and answered at length. Notice what the answer is about: usefulness, audience, and whether anybody needed the page. Notice what it is never about, which is volume, length or coverage. Then read chapter one, which quotes the sentence that makes this answer make sense.
Published by Google Search Central. There are 110 videos from 7 channels in the library further down, filterable, and every one of those channels either owns a search surface or owns a database you might use to audit your own content.
Commodity content is the whole problem
If you remember one thing Google has a word for the failure and it is not "thin". Commodity content is built from common knowledge that could have originated with anyone, and it is what you get when the plan was to cover the topic.
Not in this path.
A page is worth creating when it contains something a language model could not have guessed. Everything else in this lesson is downstream of that sentence, and Google has published the vocabulary for the failure it describes.
The mechanism is worth understanding rather than accepting. A model trained on the web has read the top ten results for your query, and several thousand pages very like them. It can produce the average of that set on request, fluently, in seconds, at zero marginal cost. So the average of the set is now free, and anything a reader could have obtained for free is not a reason to visit a website. That is not a moral argument about AI. It is an observation about what remains scarce.
What remains scarce is anything that required somebody to do something. A number you measured. A tool you bought and broke. A decision you made and regretted. A conversation nobody else was in. Google's own phrasing for the standard is unusually blunt: do not recycle what others have already said, or what "could easily be produced by a generative AI model".
Here is the ten-item version of that judgement, and I would rather you got three of them wrong now than on a page you spend a week writing. Three have a surface signal pointing the wrong way, and in all three the tell is the same: effort spent writing is not the same thing as effort spent finding out.
Commodity, or not?
Ten page ideas, and the tidy answer is wrong three times
Google’s own framing: commodity content is built from common knowledge that could have originated from anyone. Commit to a verdict before you open the reasoning. 3 of these have a surface signal pointing the wrong way, and the pattern in all three is the same: effort is not the same thing as gain.
0
of 10
Judge the first one
-
Show the verdict
Commodity
Length is not the issue. Every fact in it exists in Google’s own documentation and in forty other guides, so a reader who has read one of them learns nothing. This is Google’s own description: common knowledge that could have originated from anyone.
What would rescue it Pick the one of those five where you have run something and been surprised, and write four hundred words about that instead.
-
Show the verdict
Not commodity
You hit it, you diagnosed it, and the header dump is evidence nobody else has. Short, specific, and the only page on the internet that has this exact thing.
-
Show the verdict
Not commodity
The comparison is everywhere. The same query run through all three on a named date, with the disagreement printed, is not, because it costs three subscriptions and an afternoon.
-
Show the verdict
Commodity
Assembled is the word doing the work. No tool was used, no criteria were chosen, and the page is a reformatting of fifteen marketing sites.
What would rescue it Buy three of them, run one identical brief through each, and publish the three outputs. That is a different page and it takes a morning.
-
Show the verdict
Commodity
Careful reading is level two on the gain ladder at best, and only if the synthesis says something no single source does. A faithful explainer of a public document is a better-written copy of a public document.
-
Show the verdict
Not commodity
Now there is an observation in it. Note that Google separately tells you not to build a page per fan-out query, which is a different claim from not looking at them.
-
Show the verdict
Commodity
Ahrefs puts the phrase at 150 US searches a month with a traffic potential of 80. Even winning it is worth nothing, and the definition is not in dispute.
What would rescue it Make it a chapter of something bigger. This lesson did exactly that and said so.
-
Show the verdict
Not commodity
It is a measurement, it is dated, the tool is public, and it is unflattering. The only reason nobody else has published it is that nobody else wants to.
-
Show the verdict
Commodity
Almost every statistic in these roundups traces back to another roundup. Try the three-hop test from chapter ten on any of them and see how many reach a method.
What would rescue it Pick the three that matter to your argument, trace them properly, and print the two that did not survive.
-
Show the verdict
Not commodity
The export is the gain, and naming the four confounders is what separates a case study from an advertisement. Most published case studies omit exactly that list.
Half of them are not commodity, and notice what those 5 have in common: every one required somebody to go and do something before the writing started, and not one required more words. That is the whole of decision two, and it is why the brief at the end of this lesson refuses to call itself finished without a gain sentence.
The pattern in the ten is that every non-commodity item required somebody to go and do something before the writing started, and not one of them required more words. The shortest item on that list, three hundred words about a Cloudflare rule, is the only page on the internet with that particular header dump on it. The longest, a four-thousand-word technical SEO guide, is one of several thousand.
One clarification before the rest of the page, because I am about to spend six chapters on gathering evidence and I do not want it read as an argument against writing well. Clarity, structure and sequence are real and they matter, and they are decisions three and five here. What they cannot do is rescue a page with nothing in it. A beautifully organised restatement of the current top three is still a restatement, and it will be indexed, crawled, technically flawless and worth nothing.
Content SEO is not on-page SEO, and I counted the confusion
If you remember one thing Content SEO decides what information should exist. On-page SEO makes one page legible. Six of the eight guides audited here spend more sections on keyword research and on-page than on content, in a guide titled for content.
Not in this path.
Content SEO decides what information should exist and produces it. On-page SEO makes one page communicate its relevance clearly. They are two different jobs and merging them deletes the first one, because only the second is checkable by a plugin.
The reason to separate them is diagnosis rather than tidiness. "Our content is fine, we just need better titles" and "our titles are fine, we just need better content" are two genuinely different diagnoses with two different budgets, and a team that cannot say which column a decision belongs in will spend the year on the cheap one.
The split
Two jobs that share a page and share nothing else
Not definitions, decisions. When somebody says content SEO and you are not sure what they mean, ask which of these two columns the thing they want belongs in. Every argument about this subject resolves in about four seconds once that question is on the table.
Content SEO
Deciding what information should exist, and producing it.
Is there anything here worth publishing, and do I have it?
It decides
- Whether the page should exist at all
- What it knows that nothing else on the results page does
- Which evidence gets gathered before the draft starts
- What format the answer takes
- What gets cut, merged, updated or deleted next year
When this side is wrong and the other side is right
You publish a competent, well-optimized page that restates the current top three. It gets indexed, it gets no traffic, and every audit tool reports it as healthy. This is the 96.55% failure and no amount of on-page work reaches it.
On-page SEO
Making one page communicate its relevance clearly.
Is this page as legible as it could be for the job it already does?
It decides
- Title, description and heading wording
- Where the target phrasing appears
- Internal links out of this page
- Image alt text and file naming
- Structured data and snippet control
When this side is wrong and the other side is right
You have something worth reading and the results page cannot tell. This is real and it is cheap to fix, which is why it gets fixed first and why it gets confused for the whole subject.
The reason these merge everywhere is structural rather than lazy. On-page work is checkable by a plugin and "does this page know anything" is not, so on any shared budget the checkable half wins. The audit in the next block is what that costs: across eight published guides, on-page SEO takes more sections than any other non-content subject, and the guide with the strongest title claim in the set gives content seven sections out of thirty-two.
Now the part I did not expect to be able to measure. Ahrefs' course splits SEO Content and On-Page SEO into two chapters, which is where the idea for this page came from, so I went and counted whether anybody else does. Eight guides, the seven the brief for this page named plus Google's own starter guide, fetched as raw HTML on the day this went live and sorted heading by heading.
Original data, 8 guides, 23 August 2026
Guides titled for content, audited by their own published headings
Every guide the brief for this page named, plus Google’s starter guide. Fetched as raw HTML on the day this published, every heading extracted, site chrome removed, and each remaining leaf section sorted by subject. Classification is by what a section is about rather than by how good it is, which flatters the guides with the most definition sections in them, and where a heading could sit in two buckets it went into the one that favours the guide. Every content share below is a ceiling.
-
80%
What it does, and what it does not
Best thing in it The highest content share in the set, and the only one of the eight whose parent course gives SEO Content and On-Page SEO separate chapters. That split is where the idea for this page came from.
The gap Ten sections. It is a course chapter rather than a lesson, so the depth goes to demand and page selection and the production half is seven short subsections.
-
63%
What it does, and what it does not
Best thing in it Three consecutive sections on extractability, and the clearest statement anywhere that you are now satisfying users, ranking systems and citation systems at once.
The gap Its worked example is a shoe guide with 15,000 monthly visitors and 629 AI Overview citations. A good outcome, presented without the ninety attempts behind it.
-
22%
What it does, and what it does not
Best thing in it Honest about its own scope in the first paragraph, and the copywriting third is genuinely about writing rather than about placement.
The gap Seven of thirty-two sections are about content. Two of its three named pillars are keyword research and site structure, which on this site are steps three and five, and it is the lowest content share in the set on the guide with the strongest title claim.
-
54%
What it does, and what it does not
Best thing in it Tip seven is information gain, named as such, which only one other guide in this set does at all.
The gap On 23 August 2026 this page had a Domain Rating of 90 and an estimated 13 US organic visits a month from 7 keywords. The advice is sound and the page is not winning its own subject.
-
63%
What it does, and what it does not
Best thing in it Two named case studies with before and after traffic figures attached, which is more first-party evidence than most of this set carries.
The gap The highest raw content count in the set, and rule three is doing the work: a large share of those twenty sections define what SEO content is and why it matters. Both case studies are customer outcomes on a vendor blog, so the variable being demonstrated is the product.
-
Serpstat
SEO Analysis of Content13%
What it does, and what it does not
Best thing in it It is about auditing content that already exists, which is the half of this subject the others barely touch, and five of its eight sections are on measurement.
The gap Its quality metrics are uniqueness, fluff, keyword distribution and readability, all of which a tool can score and none of which answer whether the page knows anything.
-
50%
What it does, and what it does not
Best thing in it An accurate picture of what tooling in this category automates: briefs, topic finding, optimization scoring and repurposing.
The gap It is a product page, and it is in this table because the brief named it and because a features page is a legitimate content format, which is chapter eighteen.
-
Google
SEO Starter Guide20%
What it does, and what it does not
Best thing in it It is the number one organic result for both phrasings of this subject and it is not a content SEO guide, which is the finding in chapter two rather than a criticism of the document.
The gap Four of twenty sections touch content, and everything load-bearing Google has published about quality lives in a different document that ranks nowhere for these queries.
Scroll the table sideways
147
sections across 8 guides
65
of them about content
44%
the content share, at its most generous
36
sections on on-page SEO, the largest non-content bucket
3
guides where content is a minority of the sections
Less than half of the average guide titled for content is about content, and the largest single competitor for that space is on-page SEO. The finding I expected and did not get is worth printing too: I assumed most of these would be dominated by keyword research and on-page, and only 3 of the eight are. Ahrefs and Semrush’s blog post are both comfortably majority-content. The case for this page rests on the average and on the worst of them, not on the best. Yoast’s is the sharpest single row: 7 of 32 sections, on the guide in this set with the strongest title claim. On this site keyword research is step three, intent is step four, architecture is step five, and all three finished before this page began. Run the same count on this page and hold me to it.
65 of 147 sections are actually about content, which is 44%. Yoast's is the sharpest single row and it is honest about itself: its three named pillars are keyword research, site structure and copywriting, and seven of its thirty-two sections are about content. Two thirds of the guide with the strongest title claim in the set is steps three and five of this roadmap.
The result I expected and did not get belongs here too, because a dataset that only ever agrees with the author is not a dataset. I assumed most of the eight would be dominated by keyword research and on-page, and only 3 of them are. Ahrefs and Semrush's blog post are both comfortably majority-content. The argument for this page rests on the average and on the worst three, not on a claim that everybody is doing it wrong.
I am not accusing anybody of padding. The drift is structural. Keyword placement can be scored, headings can be counted, and "does this page know anything" cannot, so on any shared word count the scoreable half wins. Which is also why this lesson can afford to be about content: the other buckets already have their own pages on this site and all of them shipped before this one.
One more finding from those two results pages, and it decides the URL of the page you are reading. Most guides settle their own slug in a commit message nobody sees. This one is settled here, with the numbers that argue against the choice printed first.
Two results pages, 2 pulls, 23 August 2026
The bigger query, and why this page is not on it
One subject, two phrasings, one with twice the volume of the other. Both pulled the day this published, both printed here, and the column that decides it is the last one. A page marked built for it has this query as its own top keyword. Almost none of them do.
AI Overview with 9 cited sources, then one organic result, then a four-question People Also Ask block.
| # | Page | DR | US traffic | Keywords | Its own top keyword | Built for it |
|---|---|---|---|---|---|---|
| 2 | | 99 | 426,040 | 4,867 | seo 498,000 | No |
| 4 | Siteimprove A creator’s guide to SEO content strategy | 81 | 12,549 | 207 | seo content strategy 3,200 | No |
| 5 | Ahrefs SEO Content: The Beginner’s Guide | 91 | 2,376 | 140 | ahrefs blog 2,500 | No |
| 6 | Semrush SEO Content: What It Is & How to Create It | 92 | 1,007 | 44 | content seo 1,800 | No |
| 7 | Bynder 12 tips for writing SEO-optimized content in 2026 | 82 | 6,811 | 355 | how to write seo content 1,500 | No |
| 9 | Yoast The ultimate guide to content SEO | 91 | 151 | 19 | content seo 1,800 | No |
| 10 | seoClarity A Complete Guide to Writing Content for SEO that Ranks | 77 | 330 | 40 | creating seo content 700 | No |
| 11 | Michigan State University Content Best Practices for SEO | 90 | 561 | 67 | seo for content 1,200 | No |
| 12 | Backlinko SEO Content: How to Create Content That Ranks | 90 | 13 | 7 | how to optimize content for seo 450 | No |
Scroll the table sideways
The AI Overview above all of it cited 9 sources
- Google also ranking
- Siteimprove also ranking
- YouTube (MyCaptain) not in the visible results
- Bynder also ranking
- Michigan State University not in the visible results
- Marketing Miner not in the visible results
- SimpleTiger not in the visible results
- Semrush also ranking
- SEOBoost not in the visible results
Nine organic results, nine different primary targets, and not one of them is this query. Position two is a beginner guide about SEO in general with 4,867 ranking keywords, which is why Ahrefs assigns this query a traffic potential of 443,000 and a parent topic of "seo". Underneath it sits a block of about thirty Reddit threads, LinkedIn posts, Instagram reels, TikToks and YouTube videos, and the Reddit titles in it are the most honest market research on the page: "Informational content is dying", "After years of ranking, competitor takes over most keywords using hundreds of AI written articles".
The same AI Overview with the same 9 cited sources, then three organic results, then a four-question People Also Ask block.
| # | Page | DR | US traffic | Keywords | Its own top keyword | Built for it |
|---|---|---|---|---|---|---|
| 2 | | 99 | 426,040 | 4,867 | seo 498,000 | No |
| 3 | Siteimprove A creator’s guide to SEO content strategy | 81 | 12,549 | 207 | seo content strategy 3,200 | No |
| 4 | Semrush SEO Content: What It Is & How to Create It | 92 | 1,007 | 44 | content seo 1,800 | Yes |
| 6 | Ahrefs SEO Content: The Beginner’s Guide | 91 | 2,376 | 140 | ahrefs blog 2,500 | No |
| 8 | DBS Interactive Technical SEO vs. Content SEO: How Each Works Differently | 69 | 155 | 12 | content seo 1,800 | Yes |
| 9 | Yoast The ultimate guide to content SEO | 91 | 151 | 19 | content seo 1,800 | Yes |
| 10 | Michigan State University Content Best Practices for SEO | 90 | 561 | 67 | seo for content 1,200 | No |
Scroll the table sideways
The AI Overview above all of it cited 9 sources
- Google also ranking
- Siteimprove also ranking
- YouTube (MyCaptain) not in the visible results
- Bynder not in the visible results
- Michigan State University also ranking
- Marketing Miner not in the visible results
- SimpleTiger not in the visible results
- Semrush not in the visible results
- SEOBoost not in the visible results
Half the volume, eleven points less difficulty, and three of the seven ranking pages have this exact phrasing as their own top keyword. One of them, DBS Interactive at position eight, is a comparison of technical SEO against content SEO on a DR 69 agency blog with twelve ranking keywords. The market’s own pages were built for this phrasing rather than the bigger one, and the AI Overview above them is identical to the other query’s, which means the two are one need with two vocabularies.
The decision, and what it cost
What was declined
"seo content", 3,600 US searches a month against 1,800. Twice the volume, and Ahrefs assigns it a traffic potential of 443,000 with the parent topic "seo", because the page winning it is Google’s beginner guide with 4,867 ranking keywords and a top keyword worth 498,000 a month. That traffic potential was never available to anybody.
Why /content-seo/ instead
Across both results pages, none of the 9 pages ranking on the bigger query were built for it, and 3 of 7 on the smaller one were. Semrush, Yoast and DBS Interactive all have "content seo" as their own top keyword. The market’s own pages were built for the phrasing with half the volume.
The uncomfortable part
Both queries return an identical 9-source AI Overview, and 3 of the guides audited in chapter two are absent from it: Ahrefs, Yoast, Backlinko. Two of the nine cited sources are a university web team and a YouTube video. Whatever is selecting those sources is not selecting on domain authority, and I cannot tell you what it is selecting on.
The one that should worry a publisher
Backlinko’s guide has a Domain Rating of 90 and Ahrefs put its US organic traffic at 13 visits a month from 7 keywords. It is a good guide by a strong domain on its own subject, and it is functionally invisible. Authority did not save it and neither will mine.
"seo content" carries 3,600 US searches a month and "content seo" carries 1,800. Twice the volume, and not one of the nine pages ranking for the bigger phrasing has it as its own top keyword, while three pages have the smaller one. The market's own pages were built for the phrasing with half the demand, which is the same finding step five produced on a completely different subject. Mature informational markets seem to get collected rather than targeted.
The uncomfortable half is the AI Overview. Both queries return an identical 9-source citation set, and Ahrefs, Yoast and Backlinko are absent from it while a university web team page and a YouTube video are in it. Whatever selects those nine is not selecting on domain authority, and I cannot tell you what it is selecting on. Anybody who can should be asked for their measurement.
Why this comes after architecture and before on-page
If you remember one thing Architecture decided where the page lives and what links to it. This step decides what is on it. Do it earlier and you are writing pages before you know whether they have a parent.
Not in this path.
Content belongs between deciding where a page lives and deciding how it reads, because it is the last step that can still say no at a reasonable price.
Run the chain out. Keyword research gives you what people want. Search intent gives you what would satisfy them, as a page type with a count behind it. Architecture gives the page a parent and a route in. Content asks what is actually on it, and on-page then makes that legible. Every step narrows, and the cost of reversing goes up at every stage: killing a row in a spreadsheet is free, killing a draft costs a morning, and killing a published page costs a redirect and a small amount of trust.
The usual order is keyword research, then write, then optimize. That order has one structural flaw and it is the one this whole page is about: nobody ever asks whether the page should exist, because by the time anybody is looking at it, it exists. A brief that arrives at a writer has already answered the question by arriving.
And a caveat I would rather give you than have you discover. If your pages have no parent and nothing links to them, content is not your bottleneck and this lesson will not help you. Step five has the audit for that, and it found ninety-seven pages on this site with no editorial link pointing at them. A brilliant page nothing points at is a brilliant page nobody reads.
A page needs two jobs, not one
If you remember one thing A page needs a job it does for the reader and a job it does for the business, and both have to be nameable in a sentence. If only one exists you have either a hobby or a brochure.
Not in this path.
Every page needs a job it does for the reader and a job it does for the business, and both have to be nameable in a sentence before the page is worth building.
With only the reader job, you get a hobby: useful, well made, and unconnected to anything that pays for it. With only the business job, you get a brochure: a page that exists to be found and does nothing for the person who found it. Both fail, and they fail differently enough that teams argue about which one they have.
The reader job has to be a verb. Not "learn about content decay", which is a topic wearing a job's clothes. "Work out why one of their pages is falling, before spending a day on the wrong fix." That version is checkable: you can read the page afterwards and ask whether somebody could now do it.
The business job has to be something other than traffic. Traffic is not a job, it is a measurement of one. Routes to a service page, proves a capability to somebody deciding whether to hire you, collects an email from exactly the person you want, gets cited by people who influence buyers, or supports a commercial page by being the thing that earns the link. If you cannot name which, the page has no business job and you will not miss it when it is gone.
reader job "work out why one of their pages is falling" ← a verb, checkable business job "routes to the refresh service, proves I diagnose" ← named, not "traffic" with only the first a hobby with only the second a brochure with neither the 96.55%
This is also the fastest filter available on an existing content plan. Take ten planned pages and write both jobs for each. The ones where you cannot is not usually two or three. In my experience it is closer to half, and every one of those is a page somebody was going to write.
96.55%, and what those pages have in common
If you remember one thing 96.55% of roughly 14 billion pages get zero organic traffic from Google. This domain sits in the 1.94% band that gets between one and ten visits a month, and that number is on this page.
Not in this path.
Almost all published content gets nothing. Ahrefs studied roughly 14 billion pages and found 96.55% receive zero organic traffic from Google, and the number is worth sitting with before reading any advice about how to write.
The band nobody quotes is the second one. 1.94% get between one and ten monthly visits, which is the band you land in when everything works: the page is indexed, it ranks for something, a person occasionally arrives, and it is worth nothing. That failure is far more common than the total failure and much harder to see, because every report on it looks fine.
I am in that band. Here is the distribution with my own reading marked on it.
The distribution
Almost all content gets nothing, and this site is nearly all content
Ahrefs, 1 December 2023, roughly 14 billion pages from its own index. The band everybody quotes is the first one. The band that matters is the second, because it is the one you land in when the page works, gets indexed, ranks for something, and is still worth nothing.
The three URLs Ahrefs sees on this domain
Ahrefs Site Explorer, US, subdomains mode, 23 August 2026. 3 organic keywords and 8 estimated monthly organic visits across roughly 140 published pages. Not a humble brag with a hidden good number underneath. This is the number.
| URL | Ranking for | Volume | Position | Visits | Where it resolves today |
|---|---|---|---|---|---|
/seo-tools/taja-ai-review/ | taja ai | 50 | 5 | 5 | 301 to zplatform.ai |
/tools/seo-google-penalty-checker/ | google penalty checker tool | 30 | 8 | 3 | 301 to /seo-tools/free/seo-google-penalty-checker/ |
/tools/column-to-comma/ | comma converter | 500 | 25 | 0 | 301 to /seo-tools/free/column-to-comma/ |
Scroll the table sideways
Two of the three are free tools and the third is a review. Not one of them is a blog post, on a site with eleven of those and thirty-two thousand words in them. That is the first finding, and it is the reason chapter one is about deserving to exist rather than about writing better.
The second finding is the last column, and I did not expect it. All three of those URLs are 301 redirects. Two moved when the tools directory was restructured and one now points at a different domain entirely. So the complete organic footprint of this site, as the largest backlink index sees it, is three addresses that no longer exist, and the things visibly working here are the two that do something rather than describe something.
Three URLs. Two free tools and a review, on a site with 11 blog posts and 32,551 words in them. Not one of those posts appears, and that is the finding rather than the embarrassment: the things ranking on this domain are the things that do something, and the things that describe things are invisible.
What do the 96.55% have in common? Ahrefs' own answer in that study is mostly about links and indexing, which is true and is steps two and five of this roadmap. The content answer is narrower and it is this: a page in that band could have been written by somebody who did not do anything. No measurement, no test, no purchase, no conversation, nothing that cost the author a morning before the writing started. It is not that those pages are badly written. Most of them are written perfectly well.
The useful thing about a number this large is that it reverses the default. The question is not "why would this page fail", it is "why would this page be one of the 3.45%". If you cannot answer that in a sentence, you have your answer.
The pages you should not write
If you remember one thing Six questions, and any one of them kills a page. The cheapest content decision available is the one that stops a page existing, and no content calendar has a column for it.
Not in this path.
The cheapest content decision available is the one that stops a page existing, and no content calendar has a column for it. Six questions, and any one of them can kill a row.
A page now costs almost nothing to produce and exactly as much as ever to maintain, link, update and eventually delete. That asymmetry is new, it is the defining fact of content work in 2026, and almost every process in the industry was designed before it.
- Can I name what this contains that the top three do not? One sentence, written before the draft. If the sentence is "it is more comprehensive", the answer is no, because comprehensive is the definition of commodity.
- Does the evidence exist yet? Not "can I find some". Does the specific artefact exist, and if not, will I actually go and produce it this week? A brief whose evidence is "research from the top 10 results" has already decided to be a retelling.
- Is the format one I can produce? If the results page is nine calculators and I write articles, the honest outcome is not a longer article. It is either building the calculator or declining the query in writing.
- Does the business need it? The question no results page can answer. You can satisfy an intent perfectly and gain nothing, and the pages where that happens are usually the ones with the best volume.
- Do I already have something that does this job? Search your own site for the job rather than for the words. Half the answer to this question is an improvement to an existing page wearing a new URL.
- Will I maintain it? Every page published is a permanent liability. If you would not update it when the interface changes, you are not publishing an asset, you are publishing a future redirect.
The sixth is the one this era needs most and the one my own audit fails hardest. Eleven posts, none of them ever updated, and I published this page anyway. The honest version of question six is not "will I maintain it in principle". It is "have I maintained anything yet", and if the answer is no then the answer to question six is no too.
A note on what killing a page actually looks like, because "just publish fewer things" is not a process. The rows you kill stay in the document, struck through, with the reason next to them. Otherwise the same row comes back in four months from somebody who does not know it was already considered, and the argument gets had twice.
Who wrote this, and how to check me on it
Placed here rather than at the end, because you have just read two chapters built entirely on measurements of my own content and you are entitled to know who took them and what they cost me to publish.
Not in this path.
Who is teaching this, and how to check me on it
Two receipts on this page, and both of them are about my own failures
This lesson argues that a page earns its place by containing something nobody else has, and that most published content contains nothing of the kind. It would be a poor lesson if I asked that of your content and not of mine. So here it is in Google’s four categories, then the method behind every dataset on the page, including the two that make me look worst.
Senior Digital Marketing Manager, Brainstorm Force · SEO since 2010 · Coimbatore, Tamil Nadu
01
Experience Has this person published content, at volume, and lived with the results?
Since 2010, across more than a hundred sites I have owned and run, two of which I lost entirely to Panda and Penguin. I have published 11 posts and 32,551 words on this domain since June 2026 and Ahrefs values the whole domain at 8 monthly organic visits. Every failure named in the commodity chapter is one I have personally shipped, and the audit that proves it is further up this page.
The version with the failures in it02
Expertise Do they know the mechanism, or only the vocabulary?
Senior Digital Marketing Manager at Brainstorm Force, MSc Computer Software Engineering with Distinction from University of Greenwich, and a dissertation that was an automated SEO management system. Professional Member of BCS, The Chartered Institute for IT since 2012. Which is why 7 claims on this page are quoted from a document with the sentence intact and the date attached, and why 4 of them are graded invented rather than argued with.
Credentials, dated and checkable03
Authoritativeness Does anyone else say so, or only them?
30,000+ students taught across six courses and free programs, 617 published videos, and a 15,000 member lifetime-deal community. The part I can prove is on the case studies with the raw exports attached. The part I will not claim: that better content caused any specific traffic number, because nobody publishing that claim has isolated the variable either.
Six case studies, with the exports04
Trust What happens when the evidence is embarrassing?
It goes on the page at full size. On the day this published, 9 of 11 of my posts contained no table, 3 cited no external source at all, and 11 of 11 had never been updated. Google states that of the four E-E-A-T categories, trust is the most important. This is the only one of the four that costs anything to demonstrate.
What this site earns from, and howThree datasets on this page. Here is how to go and get a different answer
Eight competing guides, audited
Raw HTML fetched on 23 August 2026, every heading extracted and sorted. 147 leaf sections across 8 guides and 65 of them about content, which is 44%. Eight fetches and an afternoon. Disagree with the classification row by row: the rule is printed above the chart, and the result that argues against my own expectation is printed under it.
The auditThis site’s own corpus
node tools/corpus.mjs, run against src/content/blog on 23 August 2026. A hundred lines of regular expressions in this repository, counting presence rather than quality. Point it at your own content folder and it will produce the same four counts about you.
Eleven posts, measuredThe URL decision, published
Two SERP Overview requests and one Keywords Explorer request, US, 23 August 2026. The bigger phrasing carries twice the volume and not one page ranking for it was built for it. Search both yourself, and count the publishers rather than reading the titles.
Why not the bigger queryThe experience part, as numbers you can go and check
74%
ZipWP organic click growth from a zero baseline
Brainstorm Force, Search Console · 2026
28
ZipWP keyword variations taken to #1 from nothing
Brainstorm Force, Search Console · 2026
302,037
Bing Copilot citations earned by owned properties
Bing AI Performance · 6-month windows
30,000+
Students taught across all platforms
Udemy plus direct and free courses · Aug 2026
500+
SaaS products personally bought and tested
Since 2019
Those are counts from properties I own, with the tool and the window named. They are evidence that I have done this at some scale. They are not evidence that better content caused any of them, and nobody presenting a content programme as the cause of a traffic number has isolated that variable either. That claim is graded inference on this page, in the table two chapters up, along with every other claim it leans on.
Information gain, and the patent nobody reads properly
If you remember one thing Google filed a patent on scoring how much a document adds beyond what the reader has already seen. Gain is comparative: your page is judged against the five things read before it, not on its own.
Not in this path.
Information gain is how much a document adds beyond what the reader has already seen, which makes it comparative rather than absolute. Your page is not judged on its merits. It is judged against the five things read before it.
Google filed a patent on this in October 2018, published in November 2020, titled "Contextual estimation of link information gain". Its own abstract is the clearest statement of the idea anywhere: an information gain score for a document "is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user."
Two things follow, and the second one is the one people miss. First, being thorough is not the same as being additive: a page containing everything the top three contain has an information gain of approximately zero by construction. Second, the comparison set is whatever the reader has already read, which means gain is contextual. Your beginner explainer has high gain for somebody who has read nothing and none at all for somebody who has read three.
The ladder below is how I use this in practice, because "add information gain" as advice produces nothing. Most content is not at level zero. It is at level one or two, which feels like work, reads well, and is still replaceable.
The ladder
Five levels of information gain, and the line is level three
Gain is comparative, not absolute. Google’s patent scores a document by what it adds beyond documents the reader has already seen, which means your page is judged against the five things read before it rather than on its own merits. The middle two levels are where most good content lives, and they are where it stays replaceable.
- 0
Restatement
A summary of what the current top three say, reorganized.
The test
Could a model with no access to your business have written it? If yes, this is level zero.
What it costs
An hour, and it is the hour Google now names in its own guidance as commodity content.
A real one Any of the eleven "what is SEO content" definitions on the two results pages behind this page. They agree with each other because they are copies of each other.
- 1
Better arrangement
The same information, sequenced better, written more clearly, easier to scan.
The test
Would a reader who had already read the top three learn a new fact? If no, you are at level one.
What it costs
A day. It is genuinely useful and it is the ceiling for most content operations.
A real one Google’s own SEO Starter Guide, which is the number one organic result for both phrasings of this subject and contains almost no fact the others do not.
- 2
Synthesis
Two or more bodies of knowledge joined, where the joining is the contribution.
The test
Does the page say something true that no single source it cites says on its own?
What it costs
Two or three days, most of it reading. This is the highest level available without doing anything.
A real one Reading Google’s spam policy against its AI optimization guide and noticing that one names scaled content abuse and the other tells you to stop targeting fan-out queries for exactly that reason.
- 3
Everything below this line could have been written by somebody who did nothing
First-hand
Something you did, ran, bought, broke or measured, reported with the date on it.
The test
Is there a sentence on the page that only you could have written, and could you be wrong about it in public?
What it costs
The cost of doing the thing. There is no shortcut and this is the point.
A real one The corpus audit in chapter thirty-three. Eleven posts, three tables, zero updates, produced by a tool in the repository and unflattering to me.
- 4
New data
A dataset nobody else has, produced deliberately, that other people will cite.
The test
Could somebody write a post whose central number comes from your page?
What it costs
Weeks, usually, and most of them are wasted. Do it once a year, not once a month.
A real one The eight-guide audit in chapter two. It exists because nobody had counted the sections, and counting them took an afternoon.
The honest status of the mechanism: information gain is a patent, not a confirmed ranking system. "Contextual estimation of link information gain, US20200349181A1", assigned to Google LLC, filed 18 October 2018 and published 5 November 2020. A patent proves the idea was worth protecting. It does not prove the system shipped, and this page grades it on the record rather than documented for exactly that reason. Use the ladder because it produces better pages, not because somebody told you it is a ranking factor.
Level three is the line, and the useful thing about it is how close most briefs are to it. A page sitting at level one usually needs one measurement, one screenshot or one honest decision to cross over, and the reason it does not cross is never difficulty. It is that nobody asked the question before the draft started, and by the time there is a draft, the draft is what gets edited.
Now the status of the mechanism, stated plainly because this is where the industry overreaches. A patent is not a confirmed ranking system. It proves somebody at Google thought the idea was worth protecting in 2018. It does not prove anything shipped, and this page grades information gain "on the record" rather than "documented" for exactly that reason. Use the ladder because it produces better pages. Do not cite it as a ranking factor, and be suspicious of anybody who does.
The four things a model cannot generate
If you remember one thing Four things a model cannot generate: something you measured, something you ran, something you decided and regretted, and something somebody told you privately. Everything else on your page is available to everyone.
Not in this path.
Four categories of thing are unavailable to a language model and to every competitor who has not done the work: something you measured, something you ran, something you decided and regretted, and something somebody told you privately. Everything else on your page is available to everyone.
Something you measured. A number that did not exist until you produced it. This is the highest-value category and the one people assume requires scale. It does not: the corpus audit on this page covers eleven posts and took an afternoon to build a tool for, and it is the most-quotable thing here precisely because nobody else has bothered.
Something you ran. A test, a migration, a tool you paid for, a setting you changed and broke. The evidence is a screenshot with a date in the frame and the version named. Google's own guidance names a first-hand review as its example of a unique point of view, and the operative word in that phrase is not "review".
Something you decided and regretted. The most underused category in the industry, because it costs status. A decision published with its reasoning and its outcome is unforgeable: nobody can copy it without having made it, and a model cannot invent one that is true. Every unflattering number on this page is in this category.
Something somebody told you privately. A client, a vendor rep, a support ticket, a conversation at a conference. Attributable or anonymised, but real. This is the category most likely to be legitimately unpublishable, and when it is, say the finding without the source rather than dressing it as your own analysis.
What is not on that list: opinion. An opinion is free to generate and free to hold, and a strong one is not evidence of anything. The distinction I use is whether being wrong would cost me something. "I think tables are underused" costs nothing. "Three tables across eleven posts on my own site" can be checked and would be embarrassing if I had made it up.
Research is where the page is decided
If you remember one thing The page is decided during research, not during writing. If you start the draft without a named piece of evidence, the draft will find one, and what it finds will be the median of the results page.
Not in this path.
The page is decided during research, not during writing. If you start a draft without a named piece of evidence, the draft will find one, and what it finds will be the median of the results page.
The mechanism is unglamorous. Writing is a search process: you write a sentence, you need support for it, you go and find support, and the fastest available support is whatever the top three already say. Do that for two thousand words and you have reconstructed the consensus, one honest sentence at a time. Nobody decided to write a retelling. The process wrote one.
So the order that works is: gather first, decide second, write third. Concretely, for this page, that meant reading eight competitor guides and counting their headings before writing a word, pulling two results pages and twenty keywords, and building a tool to measure my own corpus. Three artefacts existed before the first sentence, and the argument of the page came out of them rather than being illustrated by them.
Four research inputs, in the order I would spend time on them:
- The results page, read as evidence rather than as a checklist. Not to extract subheadings. To find what every result has in common, which tells you what the commodity floor is, and what none of them has, which is your opening.
- The forums and the reviews. Reddit and community threads contain the phrasing of the actual problem, and the complaints tell you which part of the standard answer does not work. Both results pages behind this lesson had large blocks of Reddit threads in them, and their titles were better market research than the ranking guides.
- Your own data. Search Console, your CRM, your support inbox, your sales calls. This is where level three gain comes from and it is the input almost nobody budgets time for.
- The primary document. If Google, or the vendor, has published on the subject, read the document rather than the coverage of it. Most of what makes this page different from the eight guides audited above is that I opened the documents they cite.
A note on using a model for this, since it is the obvious question. Research is one of the places Google's own guidance names generative AI as genuinely useful, and I use it that way: to find what has been said, build a first outline, and tell me what I have missed. It is very good at that and it is a research input, not a research output. Everything it hands you is by definition already in the commodity set.
A receipt, and what does not count as one
If you remember one thing A receipt is checkable by a stranger. A number nobody can trace is decoration with a decimal point in it, and four of the claims graded on this page are exactly that.
Not in this path.
A receipt is a claim a stranger can check. That is the whole definition, and it excludes most of what circulates in this industry as data.
The failure is not dishonesty. It is a chain: somebody measures something, somebody quotes it, somebody quotes the quote, and four hops later a number is being stated as fact by people who have never seen the method. Try it on any SEO statistic you like. Click through to the source. From there, click through again. A depressing share of them do not survive three hops, and when the chain breaks you cannot state the caveat because you never saw the method.
The receipt ladder
What counts as evidence, graded by whether a stranger can check it
Not by how impressive it sounds. A number you measured yourself on eleven items beats a survey of ten thousand people that you cannot trace to a method, because the first one can be argued with and the second can only be repeated.
- Strong A stranger can reproduce it or read it themselves.
Your own measurement, dated, reproducible
A number you produced, with the tool named, the date attached, and a way for the reader to run it themselves.
It cannot be copied by anybody who did not do the work, and being wrong about it is public. That is what makes it worth something.
"11 posts, 3 tables, 0 updates, produced by tools/corpus.mjs on 23 August 2026." Run it against your own folder and get a different answer.
A primary document, quoted with the sentence intact
The vendor’s own words, linked, with the date the document says it was last updated.
It removes you from the chain. The reader is arguing with Google rather than with your paraphrase of Google.
"Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)"
- Usable Secondhand, and honest about it.
A named third-party study, with its method and size
Somebody else’s data, cited with the sample size, the date and a link.
Usable and secondhand. You did not control the method, so quote the number and the caveat together or not at all.
Ahrefs, 1 December 2023: 96.55% of roughly 14 billion pages get zero organic traffic from Google.
A screenshot of something you actually ran
An interface, a report, an export, with the date visible in the image.
Hard to fake casually and immediately legible. Weaker than a number because a reader cannot recompute it.
A Search Console export with the date range in the frame, not a stock photo of a laptop.
- Weak You did not check, and it shows the moment somebody does.
A statistic with a source you did not read
A number quoted from a roundup that quoted a roundup.
Half of these do not survive being traced. If you cannot find the original method, you are repeating a rumour with a decimal point in it.
"Studies show 70% of buyers…" with a link to a listicle that links to a dead PDF.
- Worthless A number with nothing behind it.
A number with no source at all
A figure stated confidently, with nothing behind it.
It is the single most common form of evidence in this subject, and four claims on this page are graded invented for exactly this reason.
"A TL;DR at the top can lift conversions by 33%." I have been told this. I cannot find who measured it.
Two habits follow from that ladder and both are cheap. First, date every figure. A number without a date is an assertion, because the reader cannot tell whether it was true last week or in 2019. Nine of my eleven posts carry no dated figure at all, which is the count that surprised me most when I ran the tool. Second, quote the sentence rather than paraphrasing it. A paraphrase of a Google document puts you in the chain; the sentence takes you out of it, and the reader gets to argue with Google instead of with you.
The hardest habit is the third one: publish the numbers that argue against you. Every unflattering figure on this page was optional. Leaving them out would have produced a cleaner argument and a page nobody has any reason to believe, because a lesson about evidence whose every number flatters the author is itself a category of evidence, and the category is marketing.
Examples, and the invented person rule
If you remember one thing A real named thing beats an invented person every time, and inventing a person is worse than having no example at all because it tells the reader you had nothing.
Not in this path.
A real named thing beats an invented person every time, and inventing a person is worse than having no example, because it tells the reader you had nothing.
"Meet Sarah, a marketing manager at a mid-sized SaaS company" is a sentence that appears in an enormous amount of content and it carries exactly zero information. Sarah does what the author needs her to do. Her problem is the problem the article solves. Her objection is the objection the article answers. No reader has ever learned anything from Sarah, and every reader who has read three of these knows what she is for.
What works instead, in order of strength. A named real thing: a tool, a site, a document, a company, with the version and the date. A number from your own data. The reader addressed directly as "you", which is honest about being a generalisation. And a real person named with permission, which is the strongest and the most expensive.
The test I use is whether the example could be wrong. "Backlinko's guide to SEO content has a Domain Rating of 90 and 13 monthly US organic visits" can be checked and I would look foolish if it were false. "A well-known SEO blog struggles to rank its own content" cannot be checked and costs nothing to say. Both sentences convey the same idea and only one of them is evidence.
The exception worth naming: a hypothetical is fine when it is labelled as one and is doing a job an example cannot. The two mocked-up first screens in chapter fourteen are invented, and they are invented because the argument is about sequence and using a real page would turn it into a critique of somebody's writing. Say "here is a made-up example" and the reader can price it correctly.
Original research at very small scale
If you remember one thing You do not need fourteen billion pages. Counting the section headings of eight competing guides took an afternoon and produced the finding this whole lesson is built on.
Not in this path.
You do not need fourteen billion pages. Counting the section headings of eight competing guides took an afternoon and produced the finding this entire lesson is built on.
Original research has a reputation problem in this industry: it means a survey of a thousand marketers, or an analysis of a million SERPs, and both are out of reach for almost everybody. That reputation is doing real damage, because it stops people producing the small, specific, checkable datasets that are actually available to them and that nobody else has bothered to make.
Four kinds of very small original research, all of which I have done in the last month:
- Count something nobody has counted. Eight guides, seventy-nine section headings, four categories. The dataset is small, the classification rule is published so you can disagree with it, and it is the only count of its kind that exists.
- Measure your own thing with a tool you wrote. A hundred lines of regular expressions produced the corpus audit on this page. The tool is more reusable than the finding, which means the second run costs nothing.
- Run the same input through several tools and publish the disagreement. The disagreement is the finding. Nobody publishes it because it makes every vendor look imprecise, which is exactly why it is valuable.
- Do the thing and log it. A migration, a test, a recovery. The log is the dataset, and it only exists if you decided to keep it before you started.
The honest limit: small research is small. Eight guides is eight guides, and I would not generalise from it to the whole industry, which is why the component says the numbers are a ceiling and prints the classification rule. Stating the limit is not weakness. It is the difference between a finding and a claim, and it is what makes the finding citable.
What the brief decides and what content decides
If you remember one thing Step four handed you a page type and a job. Content decides what goes inside it. Confusing the two is how a brief that said "comparison" produces a two-thousand-word essay.
Not in this path.
Step four handed you a page type and a job. This step decides what goes inside it. Confusing the two is how a brief that said "comparison" produces a two-thousand-word essay with a table at the bottom.
The handover is worth being precise about, because the two steps get blurred in every workflow I have seen. Search intent decides the shape: what kind of page the results page rewards, in what format, at what depth, aimed at somebody at what stage. Content decides the substance: what specifically is in it that nothing else has. A brief that arrives with only the shape is half a brief, and the missing half is the one that decides whether the page is worth building.
from step four comparison page · 8 of 10 results are comparisons · commercial · buyer stage from step six the same query run through all three tools on 23 Aug, with the disagreement shape without substance a well-formatted comparison that says what the vendors say substance without shape a genuinely new finding in a format the market does not want
Both failures are real and the second one is rarer and more painful, because you did the expensive part and then put it in the wrong container. That is chapter eighteen, and it is why format is decided by counting the results page rather than by whichever shape your team produces fastest.
One rule for the handover: the brief carries the count. Not "commercial intent" but "eight of ten results are comparisons, two are product pages, none is a guide". A label is somebody's judgement and a count is an observation, and only one of them survives being disagreed with.
The answer goes in the first screen
If you remember one thing The answer goes in the first screen or the page has not started. The common failure is not length, it is sequence: the answer is present and arrives after the reader has already gone back.
Not in this path.
The answer goes in the first screen or the page has not started. The common failure is not length, it is sequence: the answer is on the page and it arrives after the reader has already gone back.
Here is the same page twice. Same facts, same author, same total word count. The only difference is which sentence is first, and the left-hand column is not a strawman: every individual sentence in it is defensible, which is exactly why the pattern survives.
The first screen
The same page, and the only difference is which sentence is first
Same facts, same author, same total length. The left column withholds the answer for about four hundred words, which is the standard opening of an informational page. Read the left column and notice that every individual sentence in it is defensible. The sequence is the failure.
Answer withheld A reader who leaves at twenty seconds learns nothing
Content Decay: The Complete Guide for 2026
In today’s fast-moving digital landscape, content marketing has become more competitive than ever before. Businesses of every size are investing heavily in blogs, guides and resources in the hope of capturing organic search traffic.
But what happens after you hit publish? Many marketers assume their work is done. In reality, the story is just beginning, and understanding what comes next is essential for anyone serious about long-term organic growth.
What is content decay?
Content decay refers to the gradual decline in organic traffic that a piece of content experiences over time after reaching its peak performance. In this comprehensive guide, we will explore everything you need to know.
Before we dive into the causes, it is worth understanding why this matters for your overall content strategy…
…the actual diagnostic arrives around here, at about the four hundredth word
Answer first A reader who leaves at twenty seconds has what they came for
Content Decay: The Four Shapes, and Which One You Have
Content decay is four different failures sharing one word, and three of them are not fixed by a refresh. Which one you have is readable off a single Search Console chart in about two minutes.
Impressions hold, clicks fall → staleness. Update the facts.
Both slide slowly together → outclassed. Go and read who passed you.
A cliff, and the page type above you changed → the intent moved. Change format or stop.
Position stable, impressions falling → the demand left. Do nothing.
Below the fold: what each one looks like in detail, the twelve charts, and the one that is most often misdiagnosed as the others.
What actually moved
One sentence. The diagnostic in the right column was already in the left-hand page, at about the four hundredth word. Nothing was written and nothing was cut. Almost every first-screen fix is this: a move, not an edit.
What the title did
"The Complete Guide for 2026" promises coverage. "The Four Shapes, and Which One You Have" promises a decision. The second one is also a contract, and chapter twenty-three is about what happens when you break it.
What this does not claim
Any ranking mechanism. The two numbers usually attached to this argument, about scroll depth and about a TL;DR lifting conversions, are both graded invented on this page. The reader came for an answer, which is sufficient.
Almost every first-screen fix is a move rather than an edit. The sentence you need is usually already written, sitting at the four-hundredth word where the author finally got to the point, and the whole intervention is cutting it and pasting it above the introduction that was warming up to it.
Three things belong in a first screen and nothing else does. The answer, in a sentence. The reason to believe it, which is usually a number with a date or a source. And the shape of what follows, so somebody deciding whether to keep reading can decide. Everything else, including who you are and why the subject matters, belongs further down or nowhere.
What I will not do is attach a ranking mechanism to this. The two numbers usually quoted alongside this advice, that sixty percent of visitors never scroll and that a summary lifts conversions by a third, are both graded invented on this page because I cannot find who measured either. The reader came for an answer. That is sufficient, and it does not need a statistic to be true.
Pogo-sticking, and what Google has actually said
If you remember one thing Returning to the results page is a real behaviour and a badly documented signal. Write the first screen for the reader, and be suspicious of anybody selling you a ranking mechanism for it.
Not in this path.
Clicking a result and immediately returning to the results page is a real behaviour and a badly documented ranking signal, and the gap between those two facts is where a lot of confident advice lives.
What is certainly true: a person who arrives, does not find the answer, and goes back has had a bad experience, and enough of those means your page is not doing its job. That statement needs no algorithm behind it and it is the reason to write a better first screen.
What is not established: that Google measures this per page and demotes you for it. Google has said repeatedly that click behaviour is not a straightforward ranking signal, that it is noisy, and that using it naively would be trivially manipulable. The industry's response has largely been to assume it happens anyway and to sell tactics for it.
My position, stated so you can discount it appropriately: I write first screens as though it matters and I do not claim it as a mechanism. That is not fence-sitting. It is the only position the evidence supports, and the practical advice is identical either way, which is usually the tell that a disputed mechanism is not worth arguing about.
Where the argument does have consequences is measurement. If you believe pogo-sticking is a ranking factor, you will chase a metric nobody can see and optimise for the appearance of engagement, which produces the intro that withholds the answer to increase time on page. That page is worse for everybody and it is a direct product of believing something undocumented with total confidence.
Word count is the wrong unit
If you remember one thing Google states it has no preferred word count, in a parenthesis, in the list of things that mark content as made for search engines. Depth is measured in questions answered, not in words.
Not in this path.
Google states it has no preferred word count, in a parenthesis, in a list of things that mark content as made for search engines rather than for people. The sentence is short enough to quote in full and almost nobody does.
"Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't.)"
The correlation that fuels the myth is real and it is backwards. Longer pages do tend to rank better on informational queries, because thorough answers to complicated questions take more words, and because more words gives more surface for links. Length is a symptom of having something to say, and target word counts treat the symptom.
The right unit is questions answered. A page is deep enough when somebody with the original problem can act, and the reasonable follow-up questions have been dealt with rather than deferred. That is measurable in a way word count is not: list the questions a reader will have after your first screen, and check each one is answered somewhere.
Two failure modes, and the second is much more common on good teams. Too thin is a page that answers the query and abandons the reader at the first obvious follow-up. Padded is a page that hits three thousand words by answering questions nobody asked, and it is worse than thin, because it buries the parts that were working underneath the parts that exist to hit a number.
The lesson you are reading is very long, and it is worth saying why that is not a contradiction. It is long because it is thirty-five chapters with three datasets and eight interactive drills in it, and because it has a reading-path control at the top that collapses it to seven chapters if you have twenty minutes. Length as an output of scope is fine. Length as an input to a brief is the thing Google put in a parenthesis.
Write for the scanner, then for the reader
If you remember one thing A reader scans before they read. Short paragraphs, headings that say what they answer, and bold used for the load-bearing sentence rather than for keywords.
Not in this path.
Everybody scans before they read, and the scan decides whether the read happens. So the page has to work twice: once as a set of headings and bolded sentences, and once as prose.
The Nielsen Norman Group has been documenting scanning patterns since 2006, and the F-shaped pattern is the famous one: a horizontal sweep across the top, a shorter sweep further down, and a vertical scan down the left. Worth reading their own caveat too, which almost nobody quotes: it is one of several scanning patterns rather than the only one, and a page that reads well when scanned is the goal rather than a specific shape.
What survives a scan, in order of what people actually see: headings, bolded phrases, the first few words of each paragraph, list items, table rows, and anything visually distinct. What does not survive: the middle of a long paragraph, a conclusion, and anything that depends on having read the paragraph before it.
- Headings say what the section answers. "Content decay" is a label. "Content decay is four different failures" is an answer, and somebody who reads only the headings has learned the page.
- Bold the load-bearing sentence, not the keyword. Bolding a phrase because it is a target term trains readers to ignore your bold text, which costs you the one formatting tool that survives a scan.
- Two or three lines per paragraph. Not because attention spans changed. Because a wall of text has no entry points and a scanner needs somewhere to land.
- Front-load every paragraph. The first clause carries the claim and the rest supports it. A paragraph that builds to its point has hidden the point from anybody who did not read all of it.
One thing to be careful of, because this advice taken to its limit produces a bad page. Formatting is not thinking. A page chopped into eight-word paragraphs with a bold phrase in every one reads like a slide deck and cannot sustain an argument, and arguments are what a non-commodity page is made of. Scannable is a floor, not a target.
Not every answer is an article
If you remember one thing A query ending in "checker" wants an input box and no length of article gets into it. Twenty-two formats exist, most content operations produce one, and format is decided before the first sentence.
Not in this path.
A query ending in "checker" wants an input box, and no length of article gets into that results page. Format is decided before the first sentence and it is decided by counting, not by whichever shape your team makes fastest.
This is the highest-cost mistake in the whole subject and it hides behind a process that feels rigorous. Research produces a cluster. The cluster goes into a calendar. The calendar produces a brief. The brief produces an article. At no point does anybody ask whether the market wanted an article, and step four of this roadmap has the receipt for what that costs: ten of ten results for "mortgage calculator" are calculators.
So here is the register. 22 formats, grouped by what the reader is trying to finish, with the query shape that asks for each one and what non-commodity looks like in that particular shape.
The register
Twenty-two formats, grouped by what the reader is trying to finish
Not by content type. "Blog post, landing page, video" is a production taxonomy and this is not a production decision. The middle column is the useful one: it says what non-commodity looks like in that particular shape, and it is different for every row.
Answer it
Somebody wants to know a thing and then leave.
-
Definition page
Wanted by A query of the shape "what is X"
What non-commodity looks like here A definition is level zero gain by default. It earns a page only when the definitions in circulation are wrong and you can show it.
Nothing on this site
-
Explainer
Wanted by "how does X work", "why does X happen"
What non-commodity looks like here The mechanism, one layer below the level everybody else stops at.
-
Question page
Wanted by A single question with a short true answer
What non-commodity looks like here Usually none. Most of these belong as a section of something bigger.
Nothing on this site
-
Glossary
Wanted by Vocabulary lookups across a subject
What non-commodity looks like here Coverage plus consistency. A glossary is a reference object, not an article.
Help them decide
Somebody is choosing between named options and has money involved.
-
Comparison
Wanted by "X vs Y"
What non-commodity looks like here Criteria you chose and defended, and a verdict you are willing to be wrong about in public.
-
Alternatives page
Wanted by "X alternatives", "X competitors"
What non-commodity looks like here Naming who each alternative is actually for, including the case where the answer is stay where you are.
Nothing on this site
-
Best-of list
Wanted by "best X for Y"
What non-commodity looks like here Stated criteria, disclosed testing, and at least one entry you removed and said why.
-
Pricing page
Wanted by "X pricing", "how much does X cost"
What non-commodity looks like here The number, current, with the date. Most pricing content is a paraphrase of a pricing table.
-
Review
Wanted by "X review", "is X any good"
What non-commodity looks like here First-hand use. Google names first-hand review as its own example of a unique point of view.
Help them do it
Somebody is mid-task and needs the next step to work.
-
Tutorial
Wanted by "how to X"
What non-commodity looks like here Screenshots of the actual interface, the version, and the step that goes wrong.
-
Troubleshooting page
Wanted by "X not working", "fix X"
What non-commodity looks like here The cause, ranked by how often it is the cause, from cases you have actually seen.
-
Checklist
Wanted by "X checklist", "X audit"
What non-commodity looks like here Order, and an item most checklists leave out with the reason.
-
Template
Wanted by "X template", "X spreadsheet"
What non-commodity looks like here The object itself. The page around it is documentation.
-
Course or lesson
Wanted by A subject somebody wants taught in order
What non-commodity looks like here Sequence, prerequisites, and a deliverable at the end.
Prove something
Somebody needs a reason to believe you rather than the other six.
-
Case study
Wanted by "X case study", "did X work"
What non-commodity looks like here The raw export, the window, and what else changed at the same time.
-
Original research
Wanted by A question nobody has counted the answer to
What non-commodity looks like here The dataset. This is the only format whose whole value is the gain.
-
Teardown
Wanted by A named thing examined in public
What non-commodity looks like here Specificity. A teardown of a category is a listicle; a teardown of one named page is evidence.
Nothing on this site
-
Experiment writeup
Wanted by "does X actually work"
What non-commodity looks like here A method, a control, and the willingness to publish it when it did not work.
Nothing on this site
Give them a thing
The answer is not prose. It is an object they take away.
-
Free tool
Wanted by A query ending in "checker", "calculator", "generator"
What non-commodity looks like here The tool works. Nothing you write instead of building it will rank for this.
-
Dataset or table page
Wanted by "list of X", "X statistics"
What non-commodity looks like here Completeness plus a last-updated date that is real.
Nothing on this site
-
Directory
Wanted by "X tools", "best X software"
What non-commodity looks like here Coverage, structure and maintenance. It is a product, not a post.
-
Interactive explainer
Wanted by A concept that is hard to describe and easy to demonstrate
What non-commodity looks like here The demonstration. Everything on this page you can click is one of these.
6 of the 22 have no example on this site and that gap is left visible rather than filled with something plausible. It tells you what this content operation can currently produce, which is the same finding the exercise below will produce about yours. Almost every content team can make one shape well and makes it for everything, and the queries that wanted a different shape are simply lost without anybody noticing.
The middle column is the one that makes this a decision table rather than a taxonomy. Non-commodity means something different in each shape: for a definition page it means the circulating definitions are wrong and you can show it, for a free tool it means the tool works, and for a comparison it means criteria you chose and a verdict you are willing to be wrong about in public.
6 of the 22 have no example on this site and the gaps are left visible rather than filled with something plausible. That tells you what this content operation can currently produce, which is the same finding the exercise below will produce about yours. Almost every team makes one shape well and makes it for everything, and the queries that wanted a different shape are lost silently.
Visual content, and the four kinds that earn a slot
If you remember one thing Four kinds of image earn a slot: evidence, structure, comparison and instruction. A stock photograph is none of them, and nine of eleven posts on this site have no table at all.
Not in this path.
An image is carrying a claim or it is filling space, and "add images to your content" is advice that produces the second one, because it never said what the image is for.
Google's position is genuinely encouraging and it is an opportunity rather than an obligation: its generative AI features can bring in relevant images and video, which means more ways for a site to appear than a blue link. The operative word in its guidance is "relevant", and a stock photograph of somebody at a laptop is not relevant to anything.
Four kinds that earn a slot
An image is carrying a claim, or it is filling space
"Add images" is advice that produces stock photography, because nobody said what the image is for. Each row below has a one-question test and they all reduce to the same thing: would deleting this weaken something on the page? A decorative image passes no test because it was never carrying anything.
- 01
Evidence
Shows the thing happened. A screenshot of the export, the interface, the error, with the date in the frame.
The test
Would deleting it weaken a claim on the page?
What fails it
A stock photograph of somebody at a laptop, which weakens nothing because it was never carrying anything.
- 02
Structure
Shows a relationship prose is bad at. A tree, a flow, a matrix, a before and after.
The test
Did you have to write two paragraphs explaining the same relationship anyway? Then the diagram is not doing the job.
What fails it
A diagram of five boxes with the section headings in them, which restates the outline and adds nothing.
- 03
Comparison
Puts two or more named things side by side on criteria you chose.
The test
Can a reader reach a decision from the image alone?
What fails it
A comparison chart where every row is a tick for every product.
- 04
Instruction
Shows exactly where to click, in the interface as it currently looks, with the version named.
The test
Does it match what the reader sees on their screen this month?
What fails it
A three-year-old screenshot of an interface that has been redesigned twice, which actively costs the reader time.
Mine, measured on 23 August 2026: 38 images across 11 posts, and three of those posts have none at all. Two of the three are the longest things I have published. The count is not the interesting part. Almost every one of those images is the instruction kind, because the posts are tutorials, and I have produced very close to nothing in the evidence, structure or comparison categories on the blog.
The fourth kind, instruction, is the one with a maintenance cost attached and the one most likely to go actively wrong. A screenshot of an interface that has been redesigned twice does not merely fail to help. It sends the reader looking for a button that no longer exists, which is worse than the page having no image at all, and it is the single most common form of staleness in tutorial content.
A note on generated images, since it is now one prompt away. The rule is the same as everywhere else on this page: an illustration generated to fill a slot is a stock photograph with extra steps. Where a generated image genuinely helps is the structure category, when you need a diagram of something abstract, and there the honest version is to draw it rather than to generate a picture of a concept. Google Merchant Center already requires AI-generated product images to be labelled in their metadata, which is worth knowing about where it applies.
Tables and comparisons, the most underproduced format on the web
If you remember one thing The densest evidence format on the web and the least produced. A table forces you to pick criteria, and picking criteria is the part of a comparison that is actually work.
Not in this path.
A table is the densest evidence format available and almost nobody builds a useful one, because the useful version requires deciding what matters and the useless version can be assembled from three marketing pages in twenty minutes.
The distinction is not presentational. A feature dump and a criteria table look identical and do opposite jobs, and the difference is entirely in what happened before the table was drawn.
The same three products, twice
A comparison table is criteria, or it is a feature dump
Left: assembled from three marketing pages in twenty minutes. Every row is a yes and the table decides nothing. Right: criteria chosen on paper before anybody opened a product, and four of the six rows disagree. The disagreement is the entire value, and it is why the second one takes an afternoon.
Feature dump Twenty minutes. Decides nothing
| Feature | Tool A | Tool B | Tool C |
|---|---|---|---|
| Keyword research | Yes | Yes | Yes |
| Rank tracking | Yes | Yes | Yes |
| Site audit | Yes | Yes | Yes |
| Backlink data | Yes | Yes | Yes |
| Content tools | Yes | Yes | Yes |
| API access | Yes | Yes | Yes |
Six rows, eighteen cells, one distinct value. A reader finishes this table knowing exactly what they knew before, and the page has spent a screen saying so.
Criteria first An afternoon. Two rows change the answer
| Criterion | Tool A | Tool B | Tool C |
|---|---|---|---|
| Cost of the cheapest plan that includes the API | $449/mo | $249/mo | No API at any tier |
| Rows per keyword export, cheapest plan | 1,000 | 10,000 | 500 |
| Seats included before the price changes | 1 | 1 | 3 |
| Index refresh, as the vendor states it | Every 15 min | Daily | Weekly |
| Can you cancel without contacting sales | Yes | Yes | Yes |
| Free trial without a card | No | No | No |
The last two rows are greyed because every column agrees, which makes them decoration. Deleting rows where nobody disagrees is the fastest way to make a comparison useful, and nobody does it because a longer table feels more thorough.
Four rules that produce the right-hand table
- Write the criteria before opening any product. Criteria derived from a feature list are a feature list with extra steps, and they will always favour whichever product you opened first.
- Every criterion is a question a buyer asked. If you cannot name the person who would care about a row, delete the row.
- Delete every row where all the answers match. It is true, it is accurate, and it costs the reader time to read a row that cannot change their answer.
- Put the number in the cell, not a tick. "Yes" hides the difference between a thousand rows and ten thousand, which is the difference the reader is actually buying.
Mine, from the corpus audit: 3 tables across 11 posts, and nine of the eleven have none at all. On a page arguing that this is the densest evidence format on the web. Notice also what the count implies about the two rules above that require work: I have not been picking criteria, because I have not been building tables.
The rule that does the most work is the one about deleting rows where every column agrees. A row that cannot change the reader's answer is costing them attention and buying nothing, and it survives because a longer table feels more thorough. The same instinct is what produces the twenty-row comparison where every cell is a tick.
On the claim that tables get you cited by AI systems, which you will have heard: it is graded inference on this page. It is plausible, it is repeated constantly, and I cannot find anybody who has isolated it, including the vendors selling the advice. Build the table because it lets a reader decide. That reason is sufficient, it is not disputed, and it does not stop being true if the citation claim turns out to be wrong.
My own count, from the audit: three tables across eleven posts, and nine of the eleven have none. I am arguing for the format I have most neglected, which is either the least credible position on this page or the most, depending on whether you think people should recommend things they have not done. I would rather print the number and let you decide.
Video on a page, and what it is actually for
If you remember one thing A video is not a ranking accessory. It is either the primary answer, a proof you did the thing, or it does not belong on the page.
Not in this path.
A video on a page is the primary answer, or it is proof you did the thing, or it does not belong there. There is no fourth reason, and "engagement" is not one.
The primary-answer case is straightforward: some things cannot be written. Anything with timing, spatial relationships, or an interface you have to watch somebody move through. For those queries the results page usually tells you directly, because video results appear in it, and step four has the counting method.
The proof case is the interesting one and it is badly underused. A thirty-second clip of a tool actually doing the thing you claim it does is evidence in a way a paragraph is not, and it is very hard to fake casually. This is the visual-evidence category from the last chapter with a play button on it.
What does not work is a video embedded because a checklist said pages with video perform better. That correlation exists, and it exists because people who make videos are usually investing more in the page overall. Embedding a loosely related clip does not import the investment. Google's own guidance on this is measured: relevant images and video are more ways for your site to appear, and both adjectives are load-bearing.
Eight of my eleven posts embed a video and that number flatters me. Most of them are the YouTube version of the same tutorial, which is legitimate, and none of them is the thirty-second proof clip that the proof case describes. Having a video and using video are different, and the audit only measures the first one.
Every section has to survive being ripped out
If you remember one thing Every section has to survive being ripped out of the page. A passage that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere, by anyone, ever.
Not in this path.
A section that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere, by anyone, ever. Not by a reader who skimmed to it, not by somebody sharing it, and not by a retrieval system.
The reason this happens to good writers is mechanical rather than careless. When you write a section, the previous section is on your screen, so a pronoun referring back to it is perfectly clear. It stops being clear the moment anybody arrives from a search result, a table of contents link, or a quotation, which is most of the ways sections get read.
The extraction test
Four sections, read in place and then ripped out
Three of these four broken versions are from drafts of this page. Nobody writes "as I mentioned above" on purpose. You write it because the paragraph you are referring to is on the screen while you type, and it is not on the screen when somebody arrives from a search result or quotes you in a newsletter.
Flip it and read the left column as though you found it with no idea what page it came from.
- 01
The four shapes
Cannot stand alone
It has four of them, and three are not fixed by the thing everybody does. The first is the one people mean when they say the word.
Stands alone
Content decay has four shapes, and three of them are not fixed by a refresh. The first is staleness, which is the one people mean when they say the word.
What breaks
Opens on "it" and "the thing everybody does". Lifted out, the passage does not say what it is about.
What the fix cost
Nine words longer. Both nouns were already known to the writer and neither was on the page.
- 02
Why this matters for updates
Cannot stand alone
As I mentioned above, this is why the second approach usually fails. Bear that in mind when you get to the next section.
Stands alone
Bumping a publication date without changing the content usually fails, because freshness is a property of the query rather than of the page.
What breaks
Two cross-references and a forward reference in twenty-two words. It cannot be quoted anywhere, by anyone, ever.
What the fix cost
One word shorter, and it now contains the claim instead of pointing at it.
- 03
How to check it
Cannot stand alone
Open the tool and run the same report as before. If the numbers differ from the earlier ones, you have found the problem described earlier.
Stands alone
Open Search Console, filter to the URL, and chart clicks against impressions for twelve months. If impressions hold while clicks fall, the page is stale rather than outranked.
What breaks
Three references to things outside the passage: the tool, the report, the earlier numbers. A reader arriving here cold cannot act on it.
What the fix cost
Longer, and it is now the only paragraph on the page somebody would screenshot.
- 04
The exception
Cannot stand alone
There is one case where this does not apply, and it is the case discussed in the previous chapter. Everything else follows the rule above.
Stands alone
One case does not follow the rule: a query that genuinely deserves freshness, such as a news event or a price. For those, a real update to the facts does move the page.
What breaks
The entire content of the passage is a pointer to another passage. Extracted, it says nothing at all.
What the fix cost
It now contains the exception instead of promising it, which is also better for a reader going in order.
This is not a retrieval hack and it is not chunking. Google published on 10 July 2026 that there is no requirement to break your content into tiny pieces for AI to understand it. A section that stands alone is better for a reader who skimmed straight to it, better for anybody quoting you, and better for retrieval as a side effect. The reader reason is sufficient, which is convenient, because it is the only one with a document behind it.
The fix costs one sentence per section and usually costs zero extra words, because you are replacing a pronoun with the noun it was standing in for. It is the highest-return editing pass in this lesson, and the only one I would run on a page I otherwise had no time for.
I want to be precise about what this is not, because the neighbouring advice is currently being oversold. This is not chunking. Google published on 10 July 2026 that there is no requirement to break your content into tiny pieces for AI to understand it, and that there is no ideal page length. Writing sections that stand alone is a readability practice that happens to help retrieval. Cutting your page into three-hundred word blocks because somebody said models like it is a different thing, and the document those people cite says not to.
A heading is a contract
If you remember one thing A heading is a promise about what the next section answers. Break it and the reader learns to skip your headings, which costs more than any keyword it might have carried.
Not in this path.
A heading promises what the next section answers. Break the promise and the reader learns to skip your headings, which costs more than any keyword the heading might have carried.
The mechanism is trust, and it decays fast. A reader who scans your headings, picks one, jumps to it, and does not find what it promised will not scan the next set. They will read linearly, get bored, and leave, which is the outcome the headings existed to prevent. Two broken headings is enough.
So the test for a heading is whether somebody who read only that heading would be able to predict the section. Three shapes that pass, and one that does not:
the question "Which decay shape do you have?" predicts a diagnostic the claim "Word count is the wrong unit" predicts an argument the object "Four shapes, and what each one needs" predicts a list of four the label "Content decay" predicts nothing at all
The label is not wrong, it is empty. It tells a scanner the subject of the section, which they could have guessed from the page title, and it tells them nothing about whether the section is the one they need. On a long page, labels are why people give up.
Two structural rules and then this chapter is finished, because heading hierarchy is on-page work and belongs to the next lesson. Use one H1, which is the title, and Google has a short video saying that more than one is not fatal. And do not skip levels for visual reasons, because the outline is the only structure a screen reader and a table of contents can see.
The check that catches everything: read your own headings alone, in order, with the body deleted. If they read as a coherent summary of the argument, the page is structured. If they read as a list of nouns, it is filed rather than structured.
Name things specifically
If you remember one thing Name the thing. "A popular SEO tool" is unrecognisable to a reader and to a machine; "Ahrefs, on the Lite plan, in August 2026" is checkable by both.
Not in this path.
"A popular SEO tool" is unrecognisable to a reader and to a machine. "Ahrefs, on the Lite plan, in August 2026" is checkable by both, and the difference costs nothing to produce.
Vagueness in content is almost never a stylistic choice. It is what you write when you do not know, when you are hedging a claim you cannot support, or when you are avoiding naming a competitor. All three are visible to a reader who does know, and the third one is visible to everybody.
What to name, specifically: tools, with the plan or version. Companies, rather than "a leading provider". Documents, with the publisher and the date it says it was updated. People, where they have said something publicly. Numbers, with the date and the source. Interfaces, with the current label of the button rather than a paraphrase of it.
There is a retrieval argument here and I will make the modest version of it. Systems that read your page work with the entities in it, and a page that says "a popular SEO tool" contains no entity to work with. That is a claim about what is legible, not about ranking, and I would make the same recommendation with no AI systems in the picture, because specificity is how a reader can tell whether you have used the thing.
What I will not do is recommend chasing mentions, and Google's July 2026 guidance is unusually direct about it: seeking inauthentic mentions across the web "isn't as helpful as it might seem". Name things because it makes your page checkable. Do not build a programme around getting named yourself.
What Google published about AI features, and what it told you to ignore
If you remember one thing Google published in July 2026 that you can ignore chunking, ignore llms.txt and ignore chasing mentions. Most of what is sold as generative engine optimization is contradicted by the vendor’s own document.
Not in this path.
Google published a guide to optimizing for its generative AI features on 10 July 2026, and the most useful part of it is the list of things it says you can ignore. Two of them are being sold hard by people citing that same document.
Start with what it says to do, because it is short. Create valuable, non-commodity content for your audience, which it says will "likely influence your website's presence in generative AI search in the long run more than any of the other suggestions in this guide". Keep the site technically crawlable. Add high-quality relevant images and video. Organise content so readers can follow it. That is the whole substance, and it is the first five decisions of this lesson.
Now the list of things to stop doing, quoted rather than paraphrased, because the paraphrases are the problem.
"There's no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users."
"You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them."
And the one that most directly contradicts current practice. Google defines a fan-out query in the same document, as one of the concurrent related queries a model issues behind a single prompt, and then says this about building for them: creating separate content for every possible variation of how people might search, including fan-out queries, primarily to manipulate rankings violates the scaled content abuse spam policy, and is an ineffective long-term strategy besides.
I want to be fair to the tactic, because there is a defensible version. Looking at fan-out queries to understand what a model is actually asking is research, and research is legitimate. Generating a page for each one is the thing being named. The distinction is the same one running through this whole lesson: reading the market is fine, and manufacturing pages to match it is what the policy is about.
So the practical answer to "how do I optimize for AI search" is unsatisfying and cheap: you already know. Have something worth citing, say it clearly, name things specifically, and make sure a crawler can reach it. Every one of those is a decision earlier in this lesson, which is why this page has no chapter on GEO tactics and a chapter on refusing them.
Freshness is a property of the query
If you remember one thing Freshness is a property of the query, not of your page. Some queries deserve it and most do not, and changing a date without changing the content is the one intervention that reliably does nothing.
Not in this path.
Some queries deserve fresh results and most do not. Freshness is a property of what was asked, not of your publication date, which is why changing the date on an unchanged page is the one intervention that reliably does nothing.
The mechanism is easy to reason about from the searcher's side. "Best iphone" needs a recent answer because the answer changes annually. "How does photosynthesis work" does not, and a results page full of pages published last month would be worse rather than better. Google's own channel has a video titled "Is freshness an important signal for all sites?" and the answer is no, which is a more useful sentence than most of what is written about this.
So the first question about any page is not "when did I last update this". It is does this query deserve freshness at all, and you can answer it by looking at the dates on the current results page. If the top ten are all from the last six months, the query wants recency and you are in a maintenance race. If they span five years, recency is not what is being rewarded and a refresh cycle is spending money on nothing.
Three things that are actually freshness work, as opposed to date changes. Updating facts that have changed: prices, versions, interfaces, statistics, the law. Adding what has happened since: a new competitor, a policy change, an outcome you now know. And removing what is no longer true, which is the one that gets skipped because deleting a paragraph feels like losing something.
Two anti-patterns worth naming because both are common and both are visible to readers. Putting the year in the title and bumping it annually, which is a maintenance commitment most people do not keep and looks exactly like what it is when they do not. And displaying an updated date that changed because somebody fixed a typo, which is a small lie that costs trust the first time a returning reader notices nothing is different.
Content decay has four shapes
If you remember one thing Four different failures share the word decay: staleness, being outclassed, the intent moving, and the demand leaving. Three of the four are not fixed by a refresh.
Not in this path.
Four different failures share the word decay, and three of them are not fixed by a refresh. Telling them apart takes two minutes on one Search Console chart and it is the difference between a useful afternoon and a wasted one.
The reason this matters more than it sounds: refreshing is the default response to a falling page in almost every content team, and it is the correct response to exactly one of the four situations below. Applied to the other three it produces a page with a new date and the same problem, and it feels like progress, which is the expensive part.
Four shapes
Content decay is four different failures sharing one word
Diagnosis happens by shape. Nobody reads a definition and then identifies their case; they look at twelve months of Search Console, recognise a pattern, and act. So here are the four patterns, with what each one actually is, and with the wrong move given the same space as the right one.
- 01
12 months
Staleness
The page is about something that changed. Prices, interfaces, versions, laws, the year in the title.
How you know
Impressions hold and clicks fall, or the page still ranks and the comments say it is out of date.
What to do
Update the facts, the screenshots and the date, and say what changed. This is the one case where a genuine update is the whole answer.
The wrong move
Rewriting the argument. The argument was fine; the facts moved.
- 02
12 months
Outclassed
Somebody published something better. Your page did not get worse; the results page got better.
How you know
A steady slide with no obvious event, and the new pages above you have something yours does not.
What to do
Go and read what beat you, then add the thing you do not have. If you cannot name that thing, you are about to change nothing.
The wrong move
A freshness edit. Bumping the date on a page that lost on substance is theatre.
- 03
12 months
The intent moved
The results page changed shape. Nine articles became nine tools, or a comparison replaced a guide.
How you know
A cliff rather than a slide, and the page type above you is no longer yours.
What to do
Change format or stop. The previous lesson has the count method for this and it takes ten minutes.
The wrong move
Adding two thousand words to an article on a results page that no longer wants articles.
- 04
12 months
The demand left
Fewer people search for it, or the answer now arrives without a click.
How you know
Your position is stable or improving and impressions fall anyway.
What to do
Nothing, usually. Consolidate it into something with a future, or leave it and stop measuring it monthly.
The wrong move
A quarterly refresh cycle on a page whose market is shrinking, which is how teams spend a year improving a page nobody wants.
Two of the four end in doing nothing, and that is the point of diagnosing rather than refreshing. A quarterly refresh cycle applied to all four produces one correct outcome, one expensive rewrite that would have happened anyway, and two pages that get worked on every quarter for the rest of their lives. The most expensive mistake is treating a competitive loss as a freshness problem, because it produces a page with a new date and the same gap, and it feels like progress.
Two of the four end in doing nothing, and that is the finding rather than a hedge. A page whose market has shrunk does not have a content problem, and a quarterly refresh cycle applied to it will consume hours every year for the rest of its life while the author wonders why the numbers never move.
The one most often misdiagnosed is the second. Being outclassed looks like staleness because both are gradual, and the tell is what is above you: if the new pages have something yours does not, you were beaten on substance and a date change is theatre. If you cannot name the thing they have in a single sentence, you are not ready to touch the page, because you do not yet know what you would be changing.
Updating a page, and the three ways it goes wrong
If you remember one thing Update the thing that lost, not the date. If you cannot name what a competitor has that you do not, you are about to spend a day changing nothing.
Not in this path.
Update the thing that lost, not the date. Everything else about refreshing follows from that sentence, and the diagnosis from the last chapter is what tells you what lost.
Here are eight real situations and five verbs. Pick one for each before you open the reasoning, and notice how often the reflex answer is the trap.
Update, rewrite, merge, prune or leave?
Eight pages, and twice the right answer is to do nothing
Each one gives you the signal from Search Console and nothing else, which is what you actually have. Pick a verb before you open the reasoning. The trap on every item is the move a quarterly refresh process would make automatically, because it decided the verb before it looked.
0
of 8
Diagnose the first one
-
Show the reasoning
Update
Textbook staleness. The argument is fine and the facts moved, so replace the screenshots, name the current version, and say what changed at the top.
What a refresh process would do instead Rewriting it. Nothing about the explanation was wrong.
-
Show the reasoning
Rewrite
You were outclassed on substance. The only useful move is to go and get the thing they have and you do not, and if you cannot name it in one sentence you are not ready to touch the page.
What a refresh process would do instead A freshness edit. Bumping the date on a page that lost on substance changes nothing and takes a day.
-
Show the reasoning
Merge
One need, two destinations, and the site cannot say which it means. Merge into the stronger URL, redirect the other, and keep the best paragraphs from both.
What a refresh process would do instead Improving both. You will spend twice the effort to keep competing with yourself.
-
Show the reasoning
Prune
Nothing points at it, nobody arrives, and the subject is gone. Delete it and redirect to the nearest genuinely relevant page, or return a 410 if there is no such page.
What a refresh process would do instead Redirecting it to the home page, which is treated as a soft 404 and discards whatever you were preserving.
-
Show the reasoning
Leave
It is working. A quarterly refresh process would touch this page four times a year for no reason, and every touch is a chance to make it worse.
What a refresh process would do instead Refreshing it because it is old. Age is not a diagnosis.
-
Show the reasoning
Rewrite
The intent moved and the format is wrong. This is a build decision rather than an editorial one: either produce the tool or stop competing here and say so in your map.
What a refresh process would do instead Adding two thousand words. No length of article gets into a results page that wants software.
-
Show the reasoning
Leave
You got better and the market got smaller. The demand left, and there is nothing on the page to fix. Note it, stop reporting on it monthly, and put the hours somewhere with a future.
What a refresh process would do instead A refresh cycle, which is how a team spends a year improving a page for a shrinking market.
-
Show the reasoning
Rewrite
The page is being collected rather than targeted. Pick the one query it should own, answer that in the first screen, and let the other thirty-nine go.
What a refresh process would do instead Adding sections for the forty queries, which is the fan-out farming Google names as a spam policy violation.
2 of the eight are leave. If you got both of those, you have the instinct this chapter is trying to build, and it is the hardest one to hold onto because it produces nothing to put in a report. The other transferable thing here: three items look like staleness and only one of them is, which is why the diagnosis comes before the verb and not after it.
The three ways an update goes wrong, in order of how often I see them.
- Cosmetic. A new date, a reworded introduction, three new sentences, and nothing that addresses why the page fell. This is the most common outcome of a scheduled refresh, because a schedule produces the activity without producing the diagnosis.
- Additive only. Two thousand words bolted onto a page that was already long enough, because adding feels safer than deleting. The page is now worse at the job it was doing, and the section that used to answer the query is four screens down.
- Identity loss. A rewrite so thorough that the page no longer answers the query it ranked for. This one is genuinely dangerous, and the check is cheap: before you publish, open Search Console, look at what the page currently ranks for, and confirm the new version still answers the top few.
And the honest measurement problem, which is why the claim "updating recovers traffic" is graded inference on this page rather than documented. Almost every published example of a successful refresh also changed the title, the internal links and the depth at the same time. Something worked. Nobody has isolated which thing, and I have not either.
What I would do about that, practically: change one thing at a time on pages you care about, write down what you changed and when, and give it enough weeks to mean anything. Google has a video on how long SEO takes for new pages, and the same patience applies here. A refresh judged after ten days has been judged on noise.
Duplicate content and commodity content are different problems
If you remember one thing Duplicate content is a consolidation problem with a documented mechanism. Commodity content is a value problem with no mechanism at all, and it is the expensive one.
Not in this path.
Duplicate content is a consolidation problem with a documented mechanism. Commodity content is a value problem with no mechanism at all. They get discussed together, they look nothing alike, and only one of them is expensive.
Duplicate means several URLs answering one need, on your site or across sites. Nothing is being punished. Google picks one URL to represent the set, using signals it can see, and your preference is one input among several. The fixes are structural and they belong to step five: merge and redirect, differentiate genuinely, or canonicalise. This is a solved problem with published documentation and a short Google video.
Commodity means your page is not a duplicate of anything and adds nothing anyway. Every sentence is original prose, no other URL is competing with it, and it is still the average of the results page. There is no mechanism to explain, no canonical tag to add, and no report that flags it. It is the expensive one and it is the subject of this lesson.
duplicate three URLs, one need → Google picks one. Consolidate, redirect, canonicalise. commodity one URL, nothing new → no penalty, no report, no traffic. Add something or delete it. thin any length, no added value → Google's own framing. Word count is not the diagnosis.
Thin is the third word in this cluster and it is the most misused. Google's quality material treats thin content as content with little or no added value, at any length, and a four-thousand-word restatement of the top three is thin by that definition. The industry reads "thin" as "short" and then writes longer pages, which is how a diagnosis produces exactly the wrong treatment.
The syndication case deserves a sentence, because it is where these two genuinely overlap. Republishing your own article on Medium or LinkedIn is a duplicate question, solvable with a canonical or by accepting that one version wins. Republishing somebody else's press release, as thousands of sites do, is both: duplicate by construction and commodity by definition, which is why those pages are invisible even when the technical setup is perfect.
AI-assisted content, against what Google actually published
If you remember one thing Google’s position has not moved since February 2023: appropriate use of AI is not against the guidelines, and generating pages primarily to manipulate rankings is. The tool is not the line, the purpose is.
Not in this path.
Google's position has not moved since February 2023: appropriate use of AI is not against the guidelines, and using automation to generate content primarily to manipulate rankings is a spam policy violation. The tool is not the line. The purpose is.
The most useful thing in that 2023 post is a historical analogy nobody quotes. Ten years earlier there were concerns about a rise in mass-produced human content, and Google's own observation is that nobody would have thought it reasonable to ban human writing in response. The answer then was to get better at rewarding quality, and it says that is the answer now.
Twelve uses, graded against the policy
The line is purpose, not the tool
Every row graded against one published sentence rather than against a feeling: appropriate use of AI is not against the guidelines, and using automation to generate content primarily to manipulate rankings is a spam policy violation. That has been Google’s position since February 2023 and it has not changed while everything written about it has.
"Our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years."
Google Search Central, February 2023. Guidance about AI-generated content . The same post points out that ten years earlier there were the same concerns about mass-produced human content, and that banning human writing was never the answer either.
- Fine 5 Named as useful, or not objected to anywhere in the documentation.
-
Researching a topic and finding what has already been said
Google names researching a topic as a case where generative AI is particularly useful. Google Search Central
-
Turning a transcript of yourself into a structured draft
The knowledge is yours and the machine did the typing. Google names adding structure to original content as a useful case. Google Search Central
-
Building an outline from the top ten results
It is competitor reading with fewer tabs. It also produces the median outline, which is why it cannot be the last step.
-
Drafting a section you then rewrite from your own knowledge
The draft is scaffolding. If your rewrite does not change the facts, you did not have any.
-
Generating alt text, then checking every one against the image
Google names alt text among the metadata to focus accuracy on when generating automatically. Google Search Central
-
- Depends 4 Legitimate when a human owns the output. Becomes spam at the point nobody is checking.
-
Translating your own content, then having a speaker check it
Fine when a human owns the result. Google has a specific video on whether translated content is a duplicate content issue and the answer is about quality rather than duplication.
-
Summarizing your own long content into a shorter page
Two pages for one need is a cannibalization decision before it is a content one. Chapter twenty-nine.
-
Generating a definition page for every term in your niche
Legitimate if each one is checked and each one is needed. It becomes scaled content abuse at exactly the point where nobody is checking.
-
Producing a hundred location pages from one template
The template is not the problem. A hundred pages that differ only in a place name and add nothing local is the problem.
-
- Policy violation 3 Matches the published spam policy, whether a person or a model produced it.
-
Publishing a first draft with the facts unchecked
Not because it is AI. Because nobody verified it, which is chapter thirty-one and the only rule in this section that matters.
-
Generating a page for every fan-out query you can extract
Google states directly that creating separate content for every variation, including fan-out queries, primarily to manipulate rankings violates the scaled content abuse policy. Google Search Central
-
Generating many pages to cover a topic space, without adding value
This is the scaled content abuse policy almost word for word: many pages generated for the primary purpose of manipulating rankings and not helping users. Google Search Central
-
Look at the middle band, which is nine of the twelve rows. Every one of them turns on the same hinge, and it is not a technical one: whether anybody was going to check the output. That is the whole of the next chapter, and it is why "is AI content allowed" is the wrong question. The right question is who is accountable for each fact on the page.
Look at the middle band, which is three quarters of the table. Every row in it turns on the same hinge, and it is not technical: whether anybody was going to check the output. That is the next chapter, and it is why "is AI content allowed" is the wrong question. The right question is who is accountable for each fact on the page.
Where I actually use it, since you should be able to price my advice. Research and finding what has been said. Turning a transcript of me talking into structured prose, which is typing rather than thinking. First-draft sections that I then rewrite, where the rewrite always changes the facts because the draft did not have any. Alt text I check against the image. And code, which is not content.
Where I do not, and this is a preference rather than a policy: anything that constitutes the gain. If a model could write the paragraph, the paragraph is not the reason the page exists, and the paragraph that is the reason cannot be written by anything that was not there. That is not a rule I can defend from a document. It is what falls out of taking the rest of this page seriously.
Human verification, and the one rule that makes drafting safe
If you remember one thing One rule makes AI drafting safe: nothing ships that a human has not checked against a source they opened. Every AI content disaster is a verification failure wearing a technology costume.
Not in this path.
One rule makes AI-assisted drafting safe: nothing ships that a human has not checked against a source they opened. Every AI content disaster is a verification failure wearing a technology costume.
The reason models are dangerous in content work is not that they lie. It is that they are fluent, and fluency reads as confidence. A wrong statistic delivered in the same register as a right one passes every quality check a human editor runs, because the quality checks were designed to catch bad writing and the writing is fine.
So the check has to be mechanical rather than editorial. Five things get verified, every time, by a person:
- Every number. Opened, at the source, this week. Not recognised as plausible. Not remembered from somewhere. Opened.
- Every quotation. Against the original document, word for word, including whether the person actually said it. Fabricated quotations are the highest-embarrassment failure available and they are common.
- Every link. Clicked, and confirmed to be the thing the sentence says it is. A link to a plausible URL that does not contain the claim is worse than no link, because it manufactures the appearance of sourcing.
- Every product claim. Against the vendor's current page, with the date noted. Pricing and features move constantly and a model's training data does not.
- Every instruction. By doing it. If the page says click Settings then Advanced, somebody opens the product and clicks Settings then Advanced.
The uncomfortable corollary is that verification does not scale, and that is the point rather than a problem to solve. If you cannot verify a hundred pages a week, you cannot responsibly publish a hundred pages a week, and the constraint is a feature: it puts a ceiling on volume that is set by the thing that actually makes content worth reading.
Google's own guidance adds a second habit worth adopting, which is about the reader rather than about accuracy: sharing how a piece of content was created gives people useful context. That does not mean a badge on every page. It means that where automation did something substantial, saying so is more comfortable than being found out.
Where scaled content stops working
If you remember one thing Scaled content abuse is defined by purpose, not by count. The line is not the number of pages. It is whether anybody was going to check them.
Not in this path.
Scaled content abuse is defined by purpose rather than by count. Google's policy is many pages generated for the primary purpose of manipulating rankings and not helping users, and there is no number in that sentence.
Which is genuinely useful, because it means programmatic content is not the problem. A hundred thousand product pages generated from a catalog are fine, and so are location pages that contain real local information, and so is a dataset published as a page per row. Every one of those is many pages from a template and every one of them can be excellent.
The line is whether each page is worth arriving on. Three tests, and any one of them failing puts you on the wrong side:
- Does each page contain something specific to it? Not a swapped city name. A different price, a different dataset, a different photograph, a different set of facts that somebody would want.
- Would you be comfortable if a human had written every one? If a hundred humans writing these pages would still have produced a hundred worthless pages, the generator is not the problem and stopping using it will not help.
- Is anybody going to check them? This is chapter thirty-one applied at scale, and it is the test that actually decides. The point at which nobody is verifying is the point at which the programme has become the thing the policy is about.
Two related policies worth knowing by name, because they catch things people do not think of as scaled content. Expired domain abuse, buying a domain for its history and repurposing it. And site reputation abuse, publishing third-party content on a host site mainly because of the host's established ranking signals, which is the policy behind the coupon sections and sponsored review directories that appeared on news sites.
The implementation of programmatic content is architecture work rather than content work, and it belongs to step five, which has the taxonomy and URL side of it. What belongs here is the decision, and the decision is the same one from chapter six asked once instead of ten thousand times: what does each of these pages contain that nobody could have guessed?
Auditing a corpus, with mine as the worked example
If you remember one thing A corpus audit is a count of what is present, not a judgement of what is good. Mine says three tables in eleven posts and zero updates, and the tool that found that is in the repository.
Not in this path.
A content audit is a count of what is present, not a judgement of what is good, and that limitation is what makes it useful. A count can be repeated in six months and compared; a judgement cannot.
Almost every content audit guide starts with traffic, which is the right place to start for deciding what to prune and the wrong place for understanding what you produce. Traffic tells you which pages worked. It does not tell you that you have written eleven posts and three tables, and that fact predicts the traffic better than any per-page report will.
So here is mine, produced by a tool in the repository, on the day this page published, including the four counts that came back at or near zero.
This site’s own blog, measured 23 August 2026
Every post I have published, counted against my own lesson
11 posts, 32,551 words, published between
11 June 2026 and 20 August 2026. Produced by node tools/corpus.mjs,
which is a hundred lines of regular expressions in this repository. It counts what is present,
not whether it was any good. Every zero in a column that should not be zero is marked.
| Post | Words | H2s | Tables | Images | Videos | Sources cited | Internal links | Dated figures | Updated |
|---|---|---|---|---|---|---|---|---|---|
| seo-keyword-research-chatgpt-prompts | 5,454 | 21 | 1 | 0 | 1 | 0 | 3 | 0 | never |
| common-chatgpt-words-to-avoid | 4,385 | 11 | 0 | 0 | 1 | 1 | 3 | 0 | never |
| seo-penalty-recovery-case-study | 4,339 | 14 | 0 | 7 | 2 | 0 | 15 | 5 | never |
| free-ai-seo-gpt-local-seo-assistant | 3,618 | 14 | 0 | 4 | 1 | 2 | 9 | 0 | never |
| windows-server-for-seo | 3,493 | 18 | 0 | 8 | 2 | 2 | 11 | 0 | never |
| how-to-create-custom-gpt-chatgpt | 3,338 | 11 | 0 | 5 | 1 | 1 | 10 | 0 | never |
| gmail-account-for-seo | 2,672 | 15 | 0 | 7 | 1 | 0 | 12 | 0 | never |
| blogger-not-indexing-google-fix | 2,642 | 11 | 2 | 0 | 0 | 2 | 2 | 0 | never |
| internal-linking-tool | 1,111 | 6 | 0 | 3 | 1 | 3 | 6 | 0 | never |
| cheapest-domain-registrars | 753 | 5 | 0 | 3 | 0 | 1 | 6 | 0 | never |
| query-fan-out | 746 | 4 | 0 | 1 | 0 | 2 | 3 | 2 | never |
| Totals | 32,551 | 130 | 3 | 38 | 10 | 14 | 80 | 7 | 0 of 11 |
Scroll the table sideways
The rows where the answer is nothing at all
9 of 11
posts contain no table
The format with the highest evidence density per pixel, and I produced three in eleven posts.
3 of 11
posts cite no external source
Not one link out to anybody. On three of them the reader has only my word for everything.
9 of 11
posts carry no dated figure
A number without a date is an assertion. Two posts have one and one of those two is the case study.
11 of 11
posts have never been updated
Zero updatedDate values in the whole collection, on a site about to spend three chapters on refreshing.
3 of 11
posts contain no image
Two of the three are the longest posts in the collection.
Publishing this was not a stylistic choice. Chapter one argues that most content fails because
nobody asked whether it should exist, and I have thirty-two thousand words that Ahrefs values at
eight monthly visits. If you want to run the same audit on your own folder, the tool is at
content-machine/tools/corpus.mjs and it takes one argument. Your numbers will be about your site
and yours will be the ones that matter.
The four counts at the bottom are the ones this lesson argues about and every one of them is bad. 9 of 11 posts contain no table, on a page with a chapter calling tables the most underproduced format on the web. 3 cite no external source at all. Nine carry no dated figure. And 11 of 11 have never been updated, on a page with three chapters on maintenance.
What a corpus audit is for, as distinct from a page audit: it tells you what your operation can produce. Mine says I produce long tutorials with screenshots, embed the matching YouTube video, cite almost nobody, and never come back. That is a coherent and quite limited machine, and no per-page report would ever have told me, because every one of those posts is individually fine.
The four counts to run on your own folder are in the exercise below, and none of them needs a tool. Posts with no table. Posts citing nobody. Posts with no dated figure. Posts never updated. Put today's date at the top and keep the file, because the second reading is the only one that means anything and it is worthless without the first.
The mistakes I would kill first
If you remember one thing Twelve claims about content, four of them stated most confidently by people who cannot source them, and three of those four arrived in the brief for this page.
Not in this path.
Twelve claims I have been told about content, several of them by people who know more than me, and the ones with no source anywhere are the ones stated with the most confidence.
The kill list
Twelve things people believe about content
Verdict first, then the reason, then where the reason comes from. 10 of the twelve cite a primary source and the rest are answered by measurements of this site. Most of the citations are Google’s own documentation, and on this subject the documentation says something noticeably less exciting, and noticeably more useful, than the tutorials built on top of it.
-
01
""Longer content ranks better.""
No, and Google put the answer in a parenthesis.
The helpful content guidance lists writing to a word count among the signs of search-engine-first content, and answers its own question: "(No, we don’t.)" Length correlates with rankings because thorough pages tend to be longer, and the correlation gets sold as a lever.
Creating helpful, reliable, people-first content Google Search Central, Read 23 August 2026. Google states last updated 10 December 2025.
-
02
""Publish consistently and the traffic follows.""
The distribution says otherwise.
96.55% of roughly 14 billion pages get zero organic traffic. Volume is the input that everybody can supply, which is precisely why it is not the input that decides anything.
96.55% of content gets no traffic from Google Ahrefs, Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.
-
03
""Google penalizes AI content.""
Not for being AI.
"Appropriate use of AI or automation is not against our guidelines." The policy is about generating many pages primarily to manipulate rankings. Plenty of AI content is spam. It is spam for reasons that would apply to a human who wrote it the same way.
Google Search’s guidance about AI-generated content Google Search Central Blog, Read 23 August 2026. Posted February 2023 by Danny Sullivan and Chris Nelson.
-
04
""Chunk your content so AI can extract it.""
Google published the opposite in July 2026.
"There’s no requirement to break your content into tiny pieces for AI to better understand it." Same paragraph: there is no ideal page length. This one is being sold hard by people who have not read the document they are citing.
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
05
""Reverse-engineer fan-out queries and build a page for each.""
Named in the guidance as a spam policy violation.
Google names fan-out queries directly and says creating separate content for every variation primarily to manipulate rankings violates the scaled content abuse policy, and is an ineffective long-term strategy besides. Understanding fan-out is useful. Farming it is not.
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
06
""Update the date and Google treats it as fresh.""
It treats it as a lie you told twice.
Freshness is a property of the query, not a property of your page. Changing a date without changing the content is the one intervention that reliably does nothing, and on a query that does not deserve freshness there was nothing to gain anyway.
Is freshness an important signal for all sites? Google Search Central, YouTube, Watched 23 August 2026. Paired with "Query deserves freshness. Fact or fiction?" on the same channel.
-
07
""Thin content means short content.""
Thin means no added value, at any length.
Google’s own quality material treats thin content as content with little or no added value, and a four-thousand-word restatement of the top three is thin. The word count is not the diagnosis.
Search Quality Rater Guidelines, sections 4.6.5 and 4.6.6 Google, Read 23 August 2026. Cited by Google’s own generative AI content guidance as the reference for scaled content abuse and for main content created with little effort, originality or added value.
-
08
""Duplicate content is a penalty.""
It is a consolidation problem.
Several URLs answering one need means Google picks one to represent the set, using signals you did not choose. Nothing is being punished. Commodity content is the different and more expensive problem, and it looks nothing like duplication.
How does Google handle duplicate content? Google Search Central, YouTube, Watched 23 August 2026. Paired with "How can I make the pages on my site unique?" on the same channel.
-
09
""Add statistics to make content authoritative.""
Add ones somebody can check.
A number with no traceable source is decoration with a decimal point in it. Four claims on this page are graded invented and three of them arrived in the brief from an experienced practitioner who believed them.
Answered from my own data: the corpus audit of this site run on 23 August 2026, plus the eight-guide audit, the twenty keywords and the two results pages pulled the same day. Every row of all of it is shown in full further up this page.
-
10
""Cover the topic comprehensively and you will win.""
Comprehensive is the definition of commodity.
If your page contains everything the top three contain, you have written the average of the results page. Google’s own phrasing for what to avoid is content that "could easily be produced by a generative AI model".
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
11
""Information gain is a confirmed Google ranking factor.""
It is a patent, and that is a different thing.
Filed 2018, published 2020, assigned to Google. A patent proves somebody thought it was worth protecting. It does not prove the system shipped. Treat it as a good way to think and a bad thing to cite as a ranking factor.
Contextual estimation of link information gain, US20200349181A1 Google LLC, Read 23 August 2026. Filed 18 October 2018, published 5 November 2020.
-
12
""AI will write your content for you.""
It will write the median page on the subject.
That is what it is for. A model trained on the results page produces the results page, which lands you at level one on the gain ladder with a page that reads well and adds nothing. The parts that cannot be generated are the parts you had to go and get.
Answered from my own data: the corpus audit of this site run on 23 August 2026, plus the eight-guide audit, the twenty keywords and the two results pages pulled the same day. Every row of all of it is shown in full further up this page.
And now the grading, which is the part that makes the rest of this page arguable. Every load-bearing claim on it, with a grade for how much weight the evidence bears, including the four with no traceable source anywhere. Three of those four arrived in the brief for this page.
18 claims, graded
Everything this page leans on, and how much weight each one holds
The grade is about the evidence, not about whether I believe it. I believe most of the inference rows and they are still graded inference. The bottom band is the one worth your time: four figures that circulate constantly with no traceable source, and three of them arrived in the brief for this page from somebody experienced who believed them.
- Documented
Google has no preferred word count
From the helpful content guidance, in a parenthesis: "Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)"
Google Search Central, Creating helpful, reliable, people-first content - Documented
Commodity content is a named failure, and Google gave the example
The AI optimization guide contrasts "7 Tips for First-Time Homebuyers" with "Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line", and says commodity content is often based on common knowledge that could originate from anyone.
Google Search Central, Optimizing your website for generative AI features on Google Search - Documented
Using AI is not itself against the guidelines
"Appropriate use of AI or automation is not against our guidelines." What is against them is using automation to generate content primarily to manipulate rankings.
Google Search Central Blog, Google Search’s guidance about AI-generated content - Documented
You can ignore chunking your content for Google’s AI features
"There’s no requirement to break your content into tiny pieces for AI to better understand it." Same document: ignore llms.txt files and inauthentic mentions too.
Google Search Central, Optimizing your website for generative AI features on Google Search - Documented
Targeting every fan-out query with its own page is a spam policy violation
Google names fan-out queries specifically and says doing this primarily to manipulate rankings violates the scaled content abuse policy, and adds that it is ineffective anyway.
Google Search Central, Optimizing your website for generative AI features on Google Search - Documented
Scaled content abuse is about purpose, not about volume
"Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users." The count is not in the policy.
Google Search Central, Spam policies for Google web search - Documented
Trust is the most important part of E-E-A-T
Stated in those words in the helpful content guidance: "Of these aspects, trust is most important."
Google Search Central, Creating helpful, reliable, people-first content - On the record
96.55% of pages get no organic traffic from Google
Ahrefs, 1 December 2023, roughly 14 billion pages in its Content Explorer index. Their index and their traffic estimates, published with the method.
Ahrefs, 96.55% of content gets no traffic from Google - On the record
Google scores documents on information gain
A granted patent, "Contextual estimation of link information gain", filed 18 October 2018 and published 5 November 2020. A patent is evidence that the idea was worth protecting. It is not evidence that the system is live in Search, and anybody telling you it is has gone past the document.
Google LLC, Contextual estimation of link information gain, US20200349181A1 - On the record
People scan pages in an F-shaped pattern
Nielsen Norman Group eyetracking, originally 2006, reviewed 19 August 2026. Their own article says the F-pattern is one of several scanning patterns rather than the only one, which is the half nobody quotes.
Nielsen Norman Group, F-Shaped Pattern of Reading on the Web - On the record
Long context degrades in the middle
Published research on language models retrieving information from long inputs. It is a finding about model behaviour on long prompts, not a documented statement about how Google ranks a web page, and the leap between those two is where the advice gets sold.
Liu et al., Transactions of the ACL, Lost in the Middle: How Language Models Use Long Contexts - Inference
Adding a comparison table increases the chance of being cited by an AI system
Plausible, widely repeated, and nobody has isolated it. Tables also help readers, which is the reason to build one. If somebody tells you they measured this, ask what the control was.
- Inference
Answering in the first screen reduces bounce and improves rankings
The first half is a reasonable claim about readers. The second half is a leap. Google has never documented a ranking use of return-to-results behaviour and has repeatedly said clicks are not a straightforward signal.
- Inference
Updating an old post reliably recovers traffic
It works often enough that everybody has a story. Almost every published example also changed the title, the internal links and the depth at the same time, so the variable is not isolated in a single public case I can find.
- Invented
At least 60% of visitors never scroll past the top of a page
Circulates constantly with no traceable original. Scroll-depth numbers exist, they vary enormously by page type and device, and none of the ones I can trace says this. Arrived in the brief for this page and it is not going on it as a fact.
- Invented
A TL;DR at the top lifts conversions by up to 33%
No source anywhere. A TL;DR is still a good idea for reasons chapter fourteen gives, and those reasons do not need a number that nobody measured.
- Invented
Expert quotes get you cited by AI Overviews within two hours of indexing
A specific, testable, unsourced claim of exactly the kind this lesson exists to teach you to refuse. Anybody who has this timing has a screenshot, and nobody has published one.
- Invented
Content decays on a predictable schedule and needs refreshing every six months
There is no schedule. There are four different failures wearing one word, and three of them are not fixed by a refresh at all. Chapter twenty-seven.
Filter to Invented and read the four. Every one of them is specific, testable and stated with total confidence in places you would trust, and none of them has an original anybody can find. That is the failure mode this whole lesson is about, and it does not happen because people are dishonest. It happens because a number gets repeated until it sounds measured, and because nobody wants to be the one who deletes it.
Test yourself
If you remember one thing Sixteen questions. Every answer is either quoted from a document or measured on this site.
Not in this path.
Sixteen questions. Six of them quote a sentence Google published, four are answered by measurements of this site or of the eight guides above, and none of them is a definition you could have looked up.
Test yourself
Sixteen questions, and every answer is quoted or measured
Every answer is sourced to Google’s own documentation, or to a measurement of this site you can reproduce with a free tool. Four of the sixteen have Google’s sentence quoted verbatim in the explanation. If an answer disagrees with something you were taught, check the source before you check me.
0
of 16
Answer the first question to start
-
Show the answer
Correct B. Commodity content
Commodity content, from the AI optimization guide. It contrasts "7 Tips for First-Time Homebuyers" with a first-hand piece about waiving a home inspection. Thin, duplicate and scaled are three different failures with their own definitions.
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
Show the answer
Correct C. About 96.55%
Ahrefs, 1 December 2023, roughly 14 billion pages. A further 1.94% get between one and ten visits a month, which is the band this site was in on the day this page published.
96.55% of content gets no traffic from Google Ahrefs, Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.
-
Show the answer
Correct C. It has no preferred word count and says so in a parenthesis
"Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)" It appears in the list of signs that content was made for search engines rather than people.
Creating helpful, reliable, people-first content Google Search Central, Read 23 August 2026. Google states last updated 10 December 2025.
-
Show the answer
Correct B. There is no requirement to break content into tiny pieces
The AI optimization guide says there is no requirement to chunk, that there is no ideal page length, and separately that Google Search ignores llms.txt files entirely.
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
Show the answer
Correct D. None
None. Nine results, nine different primary targets. Three of the pages on the two results pages behind this lesson have "content seo" as their top keyword instead, which is half the volume and is why this page lives at /content-seo/.
-
Show the answer
Correct B. Create separate content for every variation, primarily to manipulate rankings
The AI optimization guide names fan-out queries specifically and says doing this violates the scaled content abuse policy, and adds that a high quantity of pages does not make a site higher quality anyway.
Optimizing your website for generative AI features on Google Search Google Search Central, Read 23 August 2026. Google states last updated 10 July 2026.
-
Show the answer
Correct B. A granted Google patent, filed 2018 and published 2020
"Contextual estimation of link information gain", assigned to Google LLC, filed 18 October 2018, published 5 November 2020. A patent proves the idea was worth protecting. It does not prove the system is live in Search.
Contextual estimation of link information gain, US20200349181A1 Google LLC, Read 23 August 2026. Filed 18 October 2018, published 5 November 2020.
-
Show the answer
Correct D. The demand left
Stable position plus falling impressions means fewer people are searching, or the answer now arrives without a click. Refreshing it changes nothing, because your page is not what got worse.
-
Show the answer
Correct C. Generating many pages primarily to manipulate rankings without helping users
That is the scaled content abuse definition almost verbatim. The other three are cases Google either names as useful or does not object to. The policy is about purpose, not about the tool.
Spam policies for Google web search Google Search Central, Read 23 August 2026. Google states last updated 15 May 2026.
-
Show the answer
Correct C. A little under a half
65 of 147 sections, which is 44%. The largest single competitor for that space is on-page SEO at 36 sections. Three of the eight guides are majority something else, and Yoast’s, with the strongest title claim in the set, is 7 of 32.
-
Show the answer
Correct B. A number you measured, with the tool named and the date attached
The axis is checkability by a stranger. A measurement with a named tool and a date can be reproduced and can be wrong in public, which is exactly what makes it worth something.
-
Show the answer
Correct B. Its focus is on the quality of content rather than how content is produced
From the February 2023 post, restated since: "Our focus on the quality of content, rather than how content is produced, is a useful guide that has helped us deliver reliable, high quality results to users for years."
Google Search’s guidance about AI-generated content Google Search Central Blog, Read 23 August 2026. Posted February 2023 by Danny Sullivan and Chris Nelson.
-
Show the answer
Correct D. None
None. Zero updatedDate values in the whole collection, along with three tables across eleven posts and three posts citing no external source at all. The tool that produced those numbers is in the repository.
-
Show the answer
Correct B. It is the densest way to let a reader reach a decision
The citation claim is graded inference on this page and nobody has isolated it. The reader reason is sufficient on its own, which is lucky, because it is the only one with anything behind it.
-
Show the answer
Correct C. Change format or stop
Format is decided by the results page, and no length of article gets into a market that wants an input box. This is the intent-moved decay shape and the fix is not editorial.
-
Show the answer
Correct B. It survives being lifted out of the page and still makes sense
A passage that opens on "It" or leans on "as I mentioned above" cannot be quoted anywhere. The heading question and the word count are style choices; standing alone is the actual property.
The six decisions, run on the page you are reading
Not a hypothetical and not a client. Eight decisions taken while building this page, three of them with a price attached, and the last one ending with a question still open.
Not in this path.
The worked example
The six decisions, run on the page you are reading
Eight decisions, in the order they were taken, with the evidence behind each. 4 of them cost something and the cost is printed. The last one ends with a question still open, because it is still open and closing it in public would be the kind of tidiness this whole lesson argues against.
- 01
Deserve
Should this page exist, when eight good guides already do?
Audited all eight from their raw HTML before writing anything. 65 of 147 leaf sections are actually about content, so less than half of the average guide titled for content is about content. That gap is the reason to build. The first pass of this count was wrong, because I classified from summaries rather than from the markup, and the corrected numbers are less flattering to my argument than the first ones were.
The evidence GUIDES in data/content-seo.ts. Eight URLs fetched 23 August 2026, headings extracted, chrome removed, leaf sections classified with the rule printed above the chart.
- 02
Deserve
Which URL, given one phrasing has twice the volume?
Chose /content-seo/ over /seo-content/. "seo content" carries 3,600 US searches to "content seo"’s 1,800, and not one page ranking for the bigger phrasing has it as its own top keyword, while three have "content seo". The market’s own pages were built for the smaller one.
The evidence Two SERP Overview pulls, 23 August 2026. Both queries return an identical nine-source AI Overview, so they are one need with two vocabularies.
What it cost On paper this declines 1,800 monthly searches. In practice the bigger phrasing rolls up to the parent topic "seo" and is currently won by Google’s beginner guide, so the volume was never available.
- 03
Know
What will this page contain that the other eight do not?
Three things, named before drafting: the guide audit, this site’s own corpus measured with a tool in the repository, and every load-bearing claim graded including four graded invented. Two of the four invented claims came from the brief for this page.
The evidence CLAIMS, CORPUS and GUIDES. Level four, three and four on the gain ladder respectively.
- 04
Know
Publish the domain traffic number, or leave it out?
Published it. Three organic keywords and eight estimated monthly organic visits on the day this went live, which puts this domain in the 1.94% band of the study the page opens with.
The evidence Ahrefs Site Explorer, alstonantony.com, US, 23 August 2026.
What it cost It is the worst number on the site and it is on a page about content quality. Leaving it out was the obvious call and would have made the opening argument dishonest.
- 05
Satisfy
What is the first screen?
The answer to the core question, in the first sentence, with the two numbers that make it land: 96.55% and the eight visits. No definition of content SEO above the fold, because the definition is not in dispute and eleven other pages already have it.
The evidence The TLDR block at the top of this page, written before the chapters.
- 06
Show
What is on the page besides prose?
Sixteen interactive or graphic components, two of which carry the original data. The guide audit and the corpus audit would both still be worth something with every paragraph deleted, which is the test.
The evidence The component list in src/components/content-seo/.
- 07
Structure
Chunk the content for AI, as the industry advises?
No. Google published on 10 July 2026 that there is no requirement to break content into tiny pieces, that there is no ideal page length, and that Search ignores llms.txt. The retrieval chapter says so and quotes it, which cost this page the section every competing guide now has.
The evidence Google’s AI optimization guide, read 23 August 2026.
What it cost Refusing the popular advice on a page about content is a traffic decision as much as an editorial one. The sections still get written to stand alone, because that helps readers, which is a reason that does not need a ranking claim.
- 08
Maintain
What would make me delete this page?
Two conditions, written before publication. If Google retires the commodity content framing, the spine loses its foundation. If the corpus audit stops being unflattering, the central receipt stops being evidence and becomes a boast.
The evidence The second one is a real risk and it is the good kind: it means the site improved.
What it cost The unresolved decision: "seo content strategy" carries 3,200 searches at difficulty 20 with its own parent topic, which by this site’s own rules in step five means a separate page. It is queued and not built, and saying so is cheaper than pretending this page covers it.
The reason this is the page rather than a client site: a worked example on somebody else’s property lets the author arrange the lesson so it comes out right. This one had to be published with its own traffic number attached, and that number is eight visits a month. If the method were being demonstrated on a site that already worked, you would have no way of telling whether the method or the site was doing the work.
Now build the thing this page exists to produce
Not in this path.
Reading this page does not make you able to do any of it, and I would be selling you something if I implied otherwise. This next block is the part that does, and it produces one document: a brief with a field for the thing no brief template asks for.
The deliverable
Build a brief that can say no
Thirteen fields, in decision order, and two of them do not appear on any brief template I have seen. The gain field says what this page will contain that the current top three do not. The kill field says what would make you delete it. Leave the first one empty and this is not a brief, it is a topic, and the builder will keep telling you so.
0
of 11 required fields
Start with the query
Everything you type stays in this browser and nowhere else. Nothing is sent anywhere, which matters on this form more than on the drills above it, because the gain field is the thing you know and nobody else does.
Three worked briefs, and one of them should have been killed
A page worth building Level three gain, named evidence, a first-screen sentence written before the draft, and a condition that would kill it.
- Primary query
- content decay
- The reader’s job
- Work out why one of their pages is falling, before spending a day on the wrong fix
- The business job
- Routes to the refresh service page, and proves I diagnose before I bill
- The gain, in one sentence
- Four named decay shapes, each with the Search Console pattern that identifies it, from twelve real diagnoses. Nobody has published the diagnostic; everybody has published the definition.
- The named evidence
- Twelve anonymised GSC charts with dates, plus my own corpus audit from 23 August 2026
- Gain level
- 3, first-hand
- Format
- Explainer with a diagnostic table. 8 of 10 results are guides, so the format is not in dispute.
- The first screen answer
- Content decay is four different failures sharing one word, and three of them are not fixed by a refresh.
- What is on the page besides prose
- A four-row diagnostic table, four sparkline shapes, one annotated GSC screenshot
- Section headings
- What decay is not · The four shapes · Which one you have · What to do about each · The one that means stop
- Things named specifically
- Search Console, Ahrefs Site Explorer, Google’s "Is freshness an important signal" video, 23 August 2026
- Review trigger
- When GSC impressions fall 25% over 90 days, or when Google publishes anything new on freshness
- What would make me delete this
- If I lose permission to publish the client charts, the gain is gone and the page becomes a definition nobody needs. Delete it and fold the diagnostic into the refresh lesson.
A brief that should have been killed It looks complete. The gain field is a description of thoroughness, which is the tell.
- Primary query
- seo content
- The reader’s job
- Learn what SEO content is
- The business job
- Traffic and brand awareness
- The gain, in one sentence
- A more comprehensive and better structured guide than the current results, covering every subtopic in more depth.
- The named evidence
- Research from the top 10 results and industry best practices
- Gain level
- 1, better arrangement
- Format
- Long-form guide, 3,000+ words
- The first screen answer
- SEO content is content designed to rank in search engines.
- What is on the page besides prose
- Stock images and a summary infographic
- Section headings
- What is SEO content · Types of SEO content · Why it matters · How to create it · Tools · FAQ
- Things named specifically
- Google, Ahrefs, Semrush
- Review trigger
- Quarterly
- What would make me delete this
- empty
A small page with real gain Three hundred words, one screenshot, and the only page on the internet that has it. Level three does not require a project.
- Primary query
- cloudflare sitemap 403 googlebot
- The reader’s job
- Stop their sitemap returning 403 to Googlebot today
- The business job
- Proves technical depth to exactly the reader who hires me
- The gain, in one sentence
- The specific Cloudflare rule that caused it, the request header dump from both a browser and Googlebot, and the two fixes that did not work.
- The named evidence
- My own header dump and the Search Console error, both dated
- Gain level
- 3, first-hand
- Format
- Troubleshooting page. 6 of 10 results are forum threads, which means nobody has written the page.
- The first screen answer
- A Cloudflare managed rule was returning 403 to Googlebot on /sitemap.xml while every browser got a 200. Here is the header that proves it and the rule to change.
- What is on the page besides prose
- Two header dumps side by side and the Search Console error screenshot
- Section headings
- The symptom · Proving it is Cloudflare and not your CMS · The rule · What did not work
- Things named specifically
- Cloudflare WAF managed rules, Googlebot user agent, Search Console Page Indexing report
- Review trigger
- When Cloudflare changes its managed ruleset naming
- What would make me delete this
- If Cloudflare fixes the default behaviour, this page describes a problem that no longer exists. Delete it.
Read the second one carefully. Every field is filled in, it would pass any editorial review, and it is a commodity page: the gain field describes thoroughness rather than a finding, the evidence field says "research from the top 10 results", and the kill field is empty because nothing about that page could ever stop being true. It is the most common brief in the industry and it is the one that produces the 96.55%.
Then the seven assignments. One per decision, in the order the decisions come, and each one produces something real rather than a note in a document.
The seven assignments
One assignment per decision
Bigger than the exercises and meant to be done once, properly. Finish all seven and you have a content operation that can name what its next page will contain before it is written, and that has measured what its last twenty contained.
0/7
Ticks are stored in your browser and nowhere else. Nothing is sent anywhere, which matters because four of these ask you to write down where your own content is weak.
And if you would rather have a schedule than a list, the same work spread across a week. About half an hour a day, and day seven is the one people skip.
If you want a schedule
The seven-day content audit
Same material, paced. Days one and two measure, day three grades, day four kills, day five produces something that did not exist, day six fixes what you already have, and day seven writes one brief that could say no.
0/7
Put the site name and the date at the top of whatever you end up with. In six months it is the only thing that will tell you whether your content changed or your standards did.
Then the checkpoint. Eleven questions, and if you can answer all eleven about your next page you are ahead of every content plan I have been shown.
Before you call it finished
The content checkpoint
Answerable yes or no about a real page. Any no is a work item, and the last one is the one that catches people who have done everything else.
0/11
Question eleven is the honest test. If nothing could ever make you delete the page, you have not defined what it is for.
The six decisions, one more time
Not in this path.
If you keep one thing from this page, keep these. Every tactic, framework and tool in content SEO serves one of these six decisions, and any advice that does not answer one of them is decoration.
The spine
Deserve, Know, Satisfy, Show, Structure, Maintain
In order, and the order matters. The first two can cancel the other four, which is why they are first and why almost nothing published on this subject starts there. Each card ends with what you are actually holding when the decision is made.
- 1
Deserve
Should this page exist at all?
The only decision on this list that can save you the whole cost of the page, and the one no content calendar has a column for. Google now has a word for the failure: commodity content, meaning something built from common knowledge that could have originated with anyone. If the honest answer to "what does this add" is "it covers the topic", the page is a retelling and the results page already has six.
You end up holding A shorter list, with the rows you killed still visible and a reason next to each.
- 2
Know
What do you have that a model could not have guessed?
Named before a word is written, not discovered during the draft. A number you measured, a screenshot of something you ran, a decision you made and regretted, an export, a named source with a date. Google filed a patent on scoring exactly this and calls it information gain. If you cannot name the thing, you are about to write a summary of the current top three.
You end up holding One named piece of evidence per page, with where it comes from and who can check it.
- 3
Satisfy
Does the first screen answer the thing that was asked?
The brief arrived from step four with a page type and a job. This decision is whether the job gets done immediately or gets withheld until paragraph nine. The failure is not length. It is sequence: the answer exists on the page and arrives after the reader has already gone back.
You end up holding A first screen a stranger can read in twenty seconds and leave satisfied.
- 4
Show
Is the evidence visible, or only claimed?
Tables, screenshots, diagrams, video, a dataset, a working tool. Not decoration and not "add images for SEO". Every one of these is a claim being made checkable, and the reason this decision sits fourth is that you cannot illustrate evidence you have not got. Google states that its generative AI features can bring in relevant images and video, which is opportunity rather than obligation.
You end up holding At least one thing on the page that would still be worth something with the prose deleted.
- 5
Structure
Can a scanner navigate it and a machine extract it?
Headings that say what a section answers, sections that survive being lifted out of the page, and things named specifically enough to be recognised. This is the decision most often oversold: Google published, in July 2026, that you can ignore chunking your content for AI, and the industry sold chunking anyway.
You end up holding A heading outline a stranger could read alone and know what the page covers.
- 6
Maintain
Does it survive contact with next year?
Freshness, decay, updating, pruning, and the audit that tells you which. This is the decision that makes the previous five compound instead of accumulate, and it is the one I am worst at: on the day this published, none of my eleven posts had ever been updated.
You end up holding A dated audit of your own corpus, with a verb on every row.
The test I would give a beginner for evaluating any advice about this subject, including mine: ask which of the six it answers, and ask what it would tell you not to publish. A method that only ever produces more pages is a content calendar wearing a method's clothes.
And the one sentence, if the whole page has to reduce to one. Before you write anything, name the one thing it will contain that a model could not have guessed, and if you cannot, do not write it. That sentence is free, it takes ten seconds, and applied honestly to your next ten briefs it will delete three of them. The three it deletes are the three that were going to end up in the 96.55%.
Next comes on-page SEO: how a page that deserves to exist communicates what it is, now that you know what is on it. That lesson is being written. In the meantime the roadmap has the rest of the sequence, the architecture lesson has the link graph audit that tells you whether anything points at the pages you are about to improve, and the technical SEO checklist covers the implementation half of chapters nineteen and twenty-five.
Glossary
Every word this lesson uses, defined once, in the plainest phrasing that is still correct. Where a term is normally taught with a ranking promise attached, the promise has been left out.
Not in this path.
Reference
Every word this lesson uses, in plain terms
Several are defined against Google’s own documentation rather than the industry’s retelling of it, which on this subject changes the definition and not just the wording. Where a term is usually taught with a ranking claim welded on, the claim has been left out on purpose, and the three terms Google itself defines are marked.
27 terms
- Commodity content Google’s word
- Google’s term for content built from common knowledge that could have originated from anyone, and typically adds little unique insight. Its published example is a generic first-time-homebuyer tips post.
- Content brief
- The document that decides a page before it is written. A brief that cannot say no is a topic with extra fields.
- Content decay
- A page losing traffic over time. Four distinct failures share the word: staleness, being outclassed, the intent moving, and the demand leaving.
- Content pruning
- Deliberately deleting or consolidating pages that no longer deserve to exist. The verb most content plans have no column for.
- Content SEO
- Deciding what information should exist and producing it. Distinct from on-page SEO, which makes one page communicate its relevance clearly.
- Duplicate content
- Several URLs answering one need, which makes Google pick one to represent the set. A consolidation problem, not a penalty.
- E-E-A-T
- Experience, expertise, authoritativeness and trustworthiness. A set of things quality raters look for, not a score in the algorithm. Google states trust is the most important of the four.
- Evergreen content
- Content whose usefulness does not depend on when it was published. A property of the subject, not a writing technique.
- Extractability
- Whether a section still makes sense when lifted out of the page. Useful for readers, quotes and retrieval systems alike.
- Fan-out query Google’s term
- One of the concurrent related queries a generative system issues behind a single prompt. Google defines it and separately warns against building a page for each one.
- First-party evidence
- Something you measured, ran, bought or decided. The category a language model cannot generate, which is what makes it the whole of the Know stage.
- Freshness
- A property of a query rather than of a page. Some queries deserve recent results and most do not, and no date change makes a query deserve freshness.
- GEO
- Generative engine optimization. Google’s own position is that optimizing for generative AI search is optimizing for search, and thus still SEO.
- Grounding RAG
- Retrieval-augmented generation. Google describes it as relying on core Search ranking systems to retrieve relevant pages, which the model then reviews to build a response.
- Helpful content
- Google’s framing for content created primarily for people rather than to manipulate rankings. The self-assessment questions are published and worth answering honestly once.
- Information gain
- How much a document adds beyond what the reader has already seen. Google filed a patent on scoring it. A patent is not a confirmed ranking system.
- Intent
- What would satisfy the searcher. Decided in step four of this roadmap and handed to this step as a page type with a count behind it.
- Non-commodity content
- Google’s counterpart to commodity content: unique expert or experienced takes that go beyond common knowledge, which a generative model could not easily produce.
- On-page SEO
- Making one page communicate its relevance: titles, headings, placement, links, markup. The step after this one, and the one this subject is most often confused with.
- Original research
- Producing a dataset nobody has. It does not require scale. Counting the section headings of eight guides is original research and takes an afternoon.
- Pogo-sticking
- A searcher clicking a result and immediately returning to the results page. A real behaviour, a badly documented ranking signal, and a good reason to write a better first screen anyway.
- Retrieval
- The step where a system selects passages to answer with. What makes a passage retrievable is that it stands alone, not that it is short.
- Scaled content abuse Google’s policy
- Google’s spam policy for many pages generated primarily to manipulate rankings rather than to help users. Defined by purpose, not by page count.
- Site reputation abuse Google’s policy
- Third-party content published on a host site mainly because of the host’s established ranking signals. A separate policy from scaled content, and often confused with it.
- Thin content
- Content with little or no added value, at any length. A four-thousand-word restatement of the top three is thin, which is why word count is the wrong diagnostic.
- TL;DR
- A short summary at the top of a page. A good idea for readers. The conversion numbers attached to it in circulation have no traceable source.
- Topical authority
- The idea that covering a subject broadly makes a site more likely to rank within it. Widely believed, not documented, and graded inference wherever this site uses it.
Nothing matches that. If it is a real term and it is not here, it is either something this lesson deliberately leaves to technical SEO, or one of the acronyms somebody coined last quarter for work that already had a name.
Watch the people who own the surface
Not in this path.
Where I would send you instead of a course. All free, filterable by channel, ordered by the chapter they belong to. Every channel here either owns a search surface or owns a database that measures one, and there are no independent SEO channels on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. That rule costs something here, because several of the best explanations of information gain and content pruning on YouTube are by individual practitioners and none of them is in this list.
Watch the people who own the surface
110 videos, 7 channels, ordered by chapter
Every channel here either owns a search surface or owns a database that measures one. No independent SEO channels, on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. Filter by channel, or read it top to bottom as a syllabus. Every id was checked against the YouTube oEmbed endpoint on 23 August 2026 and the channel shown is the channel oEmbed returned. All 118 candidates passed that check; the ones not in this list were cut for relevance rather than for failing it.
01 Commodity content is the whole problem
- Is more content better? SEO Mythbusting embedded above The question this whole lesson answers, put to Google directly. Watch it before chapter one and notice that the answer is never about volume.
- 3 ways to improve your content Three minutes, and none of the three is "make it longer". Useful as a check on any content plan whose only lever is production.
- Creating great content that performs well in Google search results The long version from Google’s own channel. Watch what it spends its time on and compare that against the eight guides audited in chapter two.
- Which ranking signals do SEOs worry about too much? A useful corrective before you spend a week on the wrong half of this subject.
02 Content SEO is not on-page SEO
- How to Create Content for SEO. Lesson 8/8, Semrush Academy Where content sits in a vendor curriculum, and which decisions the course hands to a different lesson. Compare the split against the one in chapter two.
- SEO Copywriting Tutorial: From Start to Finish Deliberately out of scope for this lesson and here so you can see the boundary. Copywriting is a separate page on this site for a reason printed in chapter two.
- The importance of content in the body of a page Old, short, and the clearest statement that the words in the body are the thing being evaluated rather than the words in the tags.
03 Why this comes after architecture
- How Google Search Works (in 5 minutes) Five minutes on the machine your content is being read by. Step two of the roadmap in one clip.
- The most common SEO mistake Worth a minute before you decide content is the bottleneck on your site. Sometimes it is not.
- Which aspect of my site should I focus on? Search Off the Record A long conversation about prioritizing from people who see a lot of sites. The relevant part is how rarely the answer is "publish more".
05 96.55%, and what those pages have in common
- SEO Mistakes: Why 91% of Content Gets No Organic Traffic embedded above Ahrefs’ own video on their own distribution study. Their published figure has since moved to 96.55% on a larger sample, which is worth noticing about round numbers in this industry.
- Ranking #1 on Google is Overrated The other half of the distribution argument: position is not the same as value, and a first place on the wrong query is still nothing.
- Why does a new page’s ranking change over time? Before you conclude a page failed, watch this. Some of the zero band is patience rather than quality.
06 The pages you should not write
- 5 Things in SEO that Aren’t Important A vendor arguing against work, which is rarer than it should be and worth watching for that alone.
- How to Find Great Blog Topics to Write About [4.1] The standard version of topic selection. Chapter six is the filter that should run after it and usually does not.
- How much content should be on a homepage? A narrow question with a broadly useful answer about publishing content nobody asked for.
07 Information gain, and the patent
- How to Create Content that’s "Better" than Your Competitor’s embedded above The scare quotes are the lesson. Watch for how much of "better" turns out to mean "different", which is the gain ladder in other words.
- How can I make the pages on my site unique? Google answering the question underneath this chapter: if your page is not meaningfully different, why does it exist.
- How can I make sure that Google knows my content is original? Short, and the answer is less about markup than the question expects.
- Is it a good practice to combine small portions of content from other sites? The stitching question, answered years before anybody was doing it at scale with a model.
08 The four things a model cannot generate
- How does Google think about search quality when it relies on subjective signals? Where the E-E-A-T vocabulary comes from and how loosely it maps onto anything mechanical.
- Why doesn’t Google release an SEO quality check up calculator? The answer explains why no tool scores the thing this chapter is about, and why every tool that claims to is scoring something else.
- How to Write High-Quality Content Optimized to Rank The vendor version. Watch how much of "high quality" is process and how little is knowledge, then read chapter eight.
09 Research is where the page is decided
- How to Do an Effective Content Gap Analysis for SEO Gap analysis done properly, which is a research input rather than a content plan. The difference is chapter nine.
- How to Write a Blog Post That Actually Gets Traffic The full workflow from one of the two companies whose data this lesson leans on. Notice where research stops and drafting begins.
- What Is a Content Brief The standard brief, which is a good starting point and has no field for the two things chapter nine says decide the page.
- Content Writing for SEO: How to Create Content that Ranks in Google End to end, and useful for seeing how much of the work happens before a word is written.
10 A receipt, and what does not count as one
- Do spelling and grammar matter when evaluating content and site quality? A good calibration for what Google does and does not treat as a quality signal, from the source rather than from a checklist.
- Should I focus on clarity or jargon when writing content? Two minutes, and it settles an argument that costs teams a lot of time.
- How Google Search continues to improve results How the evaluation actually happens, including the human raters, which is the context most E-E-A-T advice is missing.
- How real people make Google Search better The quality rater programme from Google’s consumer channel. Watch it before you treat rater guidelines as an algorithm specification.
11 Examples, and the invented person rule
- Which is more important: content or links? Old and still the question everybody asks. The answer is less satisfying and more useful than either camp wants.
- How can content be ranked if there aren’t many links to it? The corollary, and a reasonable answer to anybody who says content alone cannot work.
12 Original research at very small scale
- How to Write a Blog Post That Attracts Backlinks (Case Study) What original data does after publication. The link acquisition is a side effect of having something nobody else has.
- How to Increase Organic Traffic with a Content Audit An audit is also a dataset. Yours, about your own site, which is the cheapest original research available to anybody.
13 What the brief decides and what content decides
- Content Hubs: Where SEO and Content Marketing Meet Where the architecture lesson hands off to this one. Useful for seeing the seam between step five and step six.
- What Is a Content Pillar (And How It Builds Authority) The standard pillar framing, which step five argues with. Watch it as the thing being argued with rather than as instructions.
- Topic Clustering Made Simple, Beginner’s Business Guide 2025 The clearest short explanation of the model that decides how many pages a subject gets.
14 The answer goes in the first screen
- How to Create SEO Content (That Actually Ranks in 2025) Watch the first two minutes for the framing and the rest for how a vendor sequences a page. The above-the-fold advice is the durable part.
- How to Make Your Content SEO-Friendly in 5 MINUTES Five minutes, mostly on-page, and useful for spotting where the boundary in chapter two sits in practice.
15 Pogo-sticking, and what Google has actually said
- What are some myths about SEO? Watch this before quoting anybody on user signals as a ranking mechanism.
- What are some misconceptions in the SEO industry? The companion, and the reason four claims on this page are graded invented rather than argued with.
- What’s the latest SEO misconception that you would like to put to rest? Short. The value is in noticing how many of these misconceptions have a confident number attached to them.
16 Word count is the wrong unit
- How Long Should a Blog Post Be (It’s NOT 2,000+ Words) embedded above A vendor with a content database arguing against the number their own industry sells. Pair it with the parenthesis Google put in its own guidance.
- What is the ideal keyword density of a page? The ancestor of the word-count question, answered the same way and still being ignored.
- Hidden text and/or keyword stuffing What happens at the far end of writing for a target rather than for a reader.
17 Write for the scanner, then for the reader
- How to Optimize Content for SEO in 2025 The formatting half of this subject, done competently. Take the scanning advice and leave the density advice.
- More than one H1 on a page: good or bad? Settles a recurring argument in thirty seconds and frees up the hour it usually costs.
- Is there a way to indicate boilerplate content on a page? Useful for understanding what a machine treats as the main content, which is the same thing a scanner does.
18 Not every answer is an article
- The Content Format AI Overviews Haven’t Taken Over Yet Format as a strategic choice rather than a template decision, with data behind which formats are still returning clicks.
- How do I optimize an e-commerce site without rich content? A useful corrective for anybody whose only format is "article". The answer is not "add a blog to your category page".
- How to Optimize Your Content to Rank for More Keywords [5.3] Where one page can legitimately serve several queries, which is the boundary between this and building more pages.
19 Visual content, and the four kinds that earn a slot
- Does using stock photos on your pages have a negative effect on rankings? The ranking answer is reassuring and beside the point. Chapter nineteen argues the reader reason, which is the one that matters.
- SEO for Google Images The implementation half, which is on-page work. Here so you can see the boundary rather than because this lesson covers it.
- How to know if an image is AI generated or not Relevant now that "add images" and "generate images" have become the same sentence in a lot of content plans.
20 Tables and comparisons, the underproduced format
- How I Structure Content so AI Actually Cites It embedded above The best available version of the structure-for-citation argument. Note that the tables advice is graded inference on this page, and that it is a good idea for readers regardless.
- This Is NOT What We Expected To Find Analyzing 150,000 AI Citations A vendor publishing a result that undercuts the simple version of their own advice, which is rare enough to be worth an hour.
21 Video on a page, and what it is actually for
- How can a site that focuses on video or images improve its rankings? The honest answer about media-first sites, and why a video on a page is not the same as a video site.
- How To Make Video Content More Engaging: 7 Ways Production rather than SEO, which is the correct emphasis. A video nobody watches is not helping the page.
22 Every section has to survive being ripped out
- How to Optimize Content for AI Search Engines. 3.1. AEO Course by Ahrefs The most careful vendor treatment of this. Read Google’s July 2026 guidance alongside it and notice the two places they disagree.
- How Ranking in Google AI Overviews, ChatGPT, and Perplexity are Different Three surfaces, three retrieval behaviours. Useful before you optimize for "AI" as though it were one thing.
- How to Optimize Your Content For Your Target Keyword [5.2] On-page work, included at the seam. Headings are the one element both jobs have a claim on.
- How to Create SEO-Friendly Content With the SEO Content Template What a tool-generated brief actually contains. Compare its fields against the builder at the end of this lesson.
- 12 Brand Authority Signals That Make AI Recommend You Naming things specifically, framed as brand work. Take the specificity argument and treat the twelve signals as inference.
- The Real Reason Your Brand is Invisible to AI Watch it against Google’s line about not chasing inauthentic mentions. Both can be right and the distinction is the useful part.
25 What Google published about AI features, and what it told you to ignore
- Thoughts on SEO & SEO for AI, part 1 embedded above Google’s own people on what changes and what does not. The most quotable ten minutes on this subject and almost nobody cites it.
- How AI Is Changing Google Search and SEO The companion. Watch both before buying anything with "GEO" in the product name.
- Integrating generative AI in Google Search How the layer was built, from the people who built it.
- Introducing Grounding with Google Search Grounding from the developer side, which is the clearest explanation of what retrieval actually does with your page.
- Level up AI with Grounding with Google Search The applied version. Worth watching to understand why a passage that stands alone is easier to use than one that does not.
- Is searching my content with AI actually better? A developer asking the question a publisher should ask, and getting a more honest answer than the marketing gives.
- Search, 12 Days of OpenAI: Day 8 The surface where your content is read and never clicked, introduced by the company that owns it.
- Improving Web Search Results in GPT-5.3 Instant What changes when the retrieval layer improves, from the vendor rather than from a case study.
- Updates to deep research in ChatGPT Deep research reads dozens of pages per question, which changes what a page has to contain to be the one that gets used.
- ChatGPT Shopping: research and compare products without all the tabs The comparison behaviour that chapter twenty is really about, demonstrated by the surface doing it.
- Perplexity: A New Search Perspective The other citation surface, introduced by its own team. Watch what it does with sources.
- Incognito Mode & Source Preview Source preview is the clearest demonstration anywhere of what a retrieval system actually extracts from your page.
- Introducing: Deep Research on Perplexity Same behaviour, second surface. Two vendors converging on reading many pages and citing few.
26 Freshness is a property of the query
- Is freshness an important signal for all sites? embedded above The answer is no, and the reasoning is the whole chapter. Watch it before scheduling a quarterly refresh cycle.
- "Query deserves freshness." Fact or fiction? Freshness as a property of the query rather than of your page, stated plainly.
- Do dates in URLs determine freshness? Settles a decision people make once and then live with for years.
- How important is the frequency of updates on a blog? The publishing-cadence question, answered by the only party whose opinion on it is evidence.
27 Content decay has four shapes
- Let’s talk Content Decay, Search Off the Record embedded above Google’s own podcast on the subject of chapter twenty-seven, and the single best listen in this library.
- How can an older site maintain its ranking over time? The maintenance question from the other end, and a useful reality check on how much of decay is you.
- Let’s talk ranking updates, Search Off the Record What a core update actually is, from the people who ship them. Half of what gets diagnosed as decay is this.
28 Updating a page, and the three ways it goes wrong
- Republishing Content: How to Update Blog Posts For More Organic Traffic embedded above The best available walkthrough of the update itself. Run the diagnosis from chapter twenty-seven first so you are updating the right thing.
- How to do a Content Audit for Your Blog [5.4] The audit that decides which pages get updated. Pair it with the four counts in chapter thirty-three.
- How to Analyze and Optimize Your Content With the Content Audit Tool The tooled version of the same job, useful for seeing which parts a tool can decide and which it cannot.
- How long does SEO take for new pages? Watch this before concluding that an update failed. A lot of "it did not work" is measured too early.
29 Duplicate content and commodity content are different problems
- How does Google handle duplicate content? embedded above Consolidation rather than punishment, from the source. The distinction chapter twenty-nine is built on.
- How to avoid duplicate content Older and still the clearest short statement of the mechanism.
- If I quote another source, will I be penalized for duplicate content? A worry that stops people citing anybody, answered in ninety seconds.
- Thin content with little or no added value Thin defined by added value rather than by length, which is the half of the definition that gets dropped.
- Thin content (and why quality content matters) The longer treatment from Google’s monetized-sites series, which is franker than the documentation.
- Does translated content cause a duplicate content issue? Directly relevant now that translation is one prompt away and nobody is checking the output.
- Duplicate content (and what to do with it) The practical version, with the fixes ranked by how often they are the right one.
30 AI-assisted content, against what Google actually published
- Does Google see automatically generated content as a bad thing? embedded above Recorded long before this argument started, and the position has not moved since. Watch it before quoting anybody on Google penalising AI.
- AI Websites, Crawling and Search Console updates (Q1 ’26) The current state, from the channel that ships it, rather than from a conference talk about it.
- Google Search Gen AI Reports, Search Profiles & more (Q2 ’26) Where the measurement now lives, which matters more than most of the tactics being sold alongside it.
- Why You Should Never Post AI-Generated Content Without Editing A minute, and it is the whole of chapter thirty-one compressed into a warning.
- SEO in 2026: How I’d Rank in Google in the AI Era A current, opinionated take from a vendor with the data to back parts of it. Notice which parts are backed.
31 Human verification, and the one rule that makes drafting safe
- Should I spend more time on improving my content or on fighting scrapers? A question about priorities that answers a question about effort, and the answer applies directly to verification time.
- User-generated content (and building your community) Content you did not write and are still responsible for, which is the same problem as content a model wrote.
32 Where scaled content stops working
- The New SEO Playbook for AI Search (Top GEO Ranking Factors) Watch it against Google’s July 2026 guidance and mark every claim that the document contradicts. That exercise is worth more than the video.
- Learn 80% of AEO in 19 Minutes Nineteen minutes for the vocabulary, so you can tell what somebody means when they use it at you.
- How AI Search Engines Work. 1.1. AEO Course by Ahrefs The mechanism, explained well, by a company selling tooling for it. Both halves of that sentence are worth holding in mind.
What this lesson deliberately refuses to teach
Not in this path.
Ten subjects belong elsewhere, and cramming them in here would produce exactly the failure chapter two measured: a guide titled for content that is mostly about something else. Every item below is real, learnable, and worth nothing to somebody who cannot yet name what their next page will contain that nobody else has.
On-page SEO
Titles, headings, keyword placement, alt text, internal link anchors and structured data. It is the next step and it is a different job: this page decides what should exist, that one makes it legible.
Keyword research
Where demand comes from and how to size it. It is step three and it is finished before this page starts, which is exactly the boundary six of the eight audited guides fail to hold.
Step 3Search intent and SERP analysis
Reading a results page to decide what kind of page deserves to exist. Step four hands this page a brief with a page type on it.
Step 4Site architecture and internal linking
Where the page lives and what points at it. Step five, and the reason it comes first is that a page with no parent and no inbound link is a content problem you cannot write your way out of.
Step 5Content marketing and distribution
Promotion, email, social, outreach. Publishing is not distribution and this lesson stops at publish. The surface map from step one is where distribution belongs.
Step 1Copywriting and conversion
Persuasion, offers, calls to action, landing page structure. "seo copywriting" carries its own parent topic and a difficulty of zero, which by this site’s own rules means it is a separate page rather than a section here.
Editorial process and team workflow
Briefs at scale, freelancer management, editing pipelines, calendars. Real, learnable, and worth nothing to somebody who cannot yet say what their next page will contain that nobody else has.
Technical SEO for content
Rendering, indexing, canonical tags, pagination. Step two covers the mechanics and a content lesson that drifts into rendering has stopped being a content lesson.
Step 2Programmatic SEO
Generating thousands of pages from a dataset. Chapter thirty-two states where the line is and refuses the implementation, because the implementation is an architecture problem and the line is the part people get wrong.
AI writing tool comparisons
Which tool to use. The tools change every quarter and the rule in chapter thirty-one does not, so the rule is here and the tools are in the directory.
The directoryThe learner only needs six things from this page. A page earns its place by containing something a model could not have guessed. The evidence is named before the draft starts. The answer goes in the first screen. The format is decided by counting the results page. Every section has to survive being lifted out of it. And the audit of what you already have is more useful than any advice about what to write next.
Page history
What changed on this page, and when, because the central receipt on it is a measurement of my own content and it moves every time I publish or update anything.
Not in this path.
Page history
What changed, and when
The central receipt on this page is a measurement of my own content, and it will move the moment I publish or update anything here. So the changes get listed rather than absorbed silently into an "updated" stamp, and the next corpus run appears here as its own entry with the new numbers in it, including the ones that have not improved.
-
23 August 2026
First publication
- Published, as step six of the SEO roadmap.
- Eight competing guides audited from their raw HTML: 147 leaf sections, 65 of them about content, and on-page SEO the largest non-content bucket at 36. Three of the eight are majority something other than content.
- This site’s own blog measured with tools/corpus.mjs: 11 posts, 32,551 words, 3 tables, 14 external sources, 7 dated figures, 0 updates.
- This domain’s own Ahrefs reading published: 3 organic keywords, 8 estimated monthly US organic visits.
- Twenty keywords and two results pages pulled from Ahrefs, US database. Four of the twenty roll up to the parent topic "seo".
- Eighteen claims graded. Seven documented, four on the record, three inference, four invented. Three of the four invented ones arrived in the brief for this page.
- URL decision published: /content-seo/ over /seo-content/, against twice the volume, with the reason and the cost printed.
If you find something here that a current Google, Ahrefs or Semrush document contradicts, that is a bug in this page and I would rather hear about it than have it quietly rot. The contact page works.
Sources
Every claim above, traced to the document it came from, with the date it was read and the date the document says it was last updated. One of the entries is a tool in my own repository.
Not in this path.
Everything on this page, sourced
22 sources, each read on 22 August 2026. Where a document prints its own last-updated date, both dates are shown, because the document’s own date is the one that tells you whether the guidance has moved since I quoted it. Where an article has an author, the author is named rather than the brand.
Google Search Central 5
- Creating helpful, reliable, people-first content Read 23 August 2026. Google states last updated 10 December 2025.
- Optimizing your website for generative AI features on Google Search Read 23 August 2026. Google states last updated 10 July 2026.
- Google Search’s guidance on using generative AI content on your website Read 23 August 2026. Google states last updated 10 July 2026.
- Spam policies for Google web search Read 23 August 2026. Google states last updated 15 May 2026.
- SEO Starter Guide Read 23 August 2026. Google states last updated 10 December 2025.
Google Search Central Blog 1
- Google Search’s guidance about AI-generated content Read 23 August 2026. Posted February 2023 by Danny Sullivan and Chris Nelson.
Google 1
- Search Quality Rater Guidelines, sections 4.6.5 and 4.6.6 Read 23 August 2026. Cited by Google’s own generative AI content guidance as the reference for scaled content abuse and for main content created with little effort, originality or added value.
Google Search Central, YouTube 2
- Is freshness an important signal for all sites? Watched 23 August 2026. Paired with "Query deserves freshness. Fact or fiction?" on the same channel.
- How does Google handle duplicate content? Watched 23 August 2026. Paired with "How can I make the pages on my site unique?" on the same channel.
Ahrefs 3
- 96.55% of content gets no traffic from Google Read 23 August 2026. Published 1 December 2023, roughly 14 billion pages from Ahrefs’ Content Explorer index.
- SEO Content: The Beginner’s Guide Read 23 August 2026. Chapter four of Ahrefs’ SEO course, and the only guide in the audit that separates SEO Content from On-Page SEO.
- SERP Overview and Keywords Explorer, US database Pulled 23 August 2026. Two SERP Overview requests, one Keywords Explorer request of twenty keywords, one Site Explorer request against this domain.
Semrush 2
- SEO Content: What It Is & How to Create It Read 23 August 2026. Seventeen sections, three of them on extractability.
- Content Marketing Platform Read 23 August 2026. A product page, included because the brief named it and because a features page is a content format.
Yoast 1
- The ultimate guide to content SEO Read 23 August 2026. Three named pillars: keyword research, site structure, copywriting.
Backlinko 1
- SEO Content: How to Create Content That Ranks Read 23 August 2026. Names information gain as tip seven. Ahrefs put its US organic traffic at 13 visits a month on the same day.
AIOSEO 1
- SEO Content Read 23 August 2026. Two named case studies with before and after traffic figures.
Serpstat 1
- SEO Analysis of Content Read 23 August 2026. Published 2 April 2024. About auditing content that already exists.
Google LLC 1
- Contextual estimation of link information gain, US20200349181A1 Read 23 August 2026. Filed 18 October 2018, published 5 November 2020.
Nielsen Norman Group 1
- F-Shaped Pattern of Reading on the Web Read 23 August 2026. Original eyetracking 2006, article reviewed 19 August 2026 by Kara Pernice.
Liu et al., Transactions of the ACL 1
- Lost in the Middle: How Language Models Use Long Contexts Read 23 August 2026. A finding about model behaviour on long inputs, not about web ranking.
This repository 1
- corpus.mjs, the content audit tool Run 23 August 2026 against src/content/blog. A hundred lines of regular expressions in content-machine/tools/. It counts presence, not quality.
Eight of these are Google’s own documentation and seven of them are quoted verbatim on this page, with the sentence intact rather than a paraphrase, because the paraphrases in circulation are how "chunk your content for AI" ended up being sold out of a document that says there is no requirement to do it. Two entries are documents I read and deliberately would not lean on as ranking evidence: a Google patent, and a paper about how language models handle long inputs. Both are graded on the record rather than documented in the claims table. If one of these sources now says something different from what I have written, the source wins and this page is out of date.
The last entry is mine and it is the fastest-rotting thing here. The corpus audit is whatever ’node tools/corpus.mjs’ said on 23 August 2026, and it changes the moment I publish or update anything, which is the point: the numbers on this page are bad and they are supposed to get better. The eight-guide audit will rot too, because six of those eight guides will be revised. The tool is in my own repository rather than in a screenshot, and it is the same one this page tells you to run against your own content. If your numbers disagree with mine, yours are about your site and they are the ones that matter.