SEO roadmap, step 5
SEO Site Architecture
How to turn validated search demand into a logical system of pages, before you create anything. Six decisions, the link graph of my own site published unflattered, and one deliverable at the end that is not an XML sitemap.
Jump to a chapter 33
- 01 Architecture is the link graph
- 02 Five layers people call one thing
- 03 Why this step is fifth
- 04 Cluster, intent, destination
- 05 One page or several
- 06 Cannibalization is an architecture decision
- 07 Page types, and the queries that want them
- 08 The keyword-to-URL map
- 09 The pages you should not build
- 10 Four architectures, four business models
- 11 The test a parent page has to pass
- 12 Categories, taxonomy and the debt
- 13 Hubs without the mythology
- 14 Topic clusters are a model, not a formula
- 15 Silos leak, and that is the point
- 16 What earns a permanent nav slot
- 17 The mobile menu that hides half your links
- 18 A link Google cannot follow is not a link
- 19 URL architecture
- 20 URL hierarchy is not conceptual hierarchy
- 21 Subfolders, subdomains and root clutter
- 22 Breadcrumbs, and what they are not
- 23 Faceted navigation and pagination
- 24 Click depth, and the three-click religion
- 25 Does the structure treat this page as important?
- 26 Orphans, dead ends and the sitemap that hides them
- 27 Internal links do three jobs
- 28 Building for 200 pages, not 20
- 29 Competitor structure is evidence, not instructions
- 30 One structure for humans, crawlers and retrieval
- 31 SEO architecture debt
- 32 The mistakes I would kill first
- 33 Test yourself
What the next hour buys you
What you will actually be able to do
Not "understand site architecture". These are the specific things you should be able to do to your own site, by yourself, when you close the tab. If one of them still feels impossible afterwards, that chapter failed and I would rather know which.
- Export your own internal links and name every page nothing on the site points at
- Explain why a perfect folder tree is not an architecture, quoting the sentence Google published
- Take a cluster list and decide how many URLs it actually needs, with a reason per row
- Name a page type other than "blog post" for at least a third of your clusters
- Apply one test to a parent page and be willing to delete it when it fails
- Turn a keyword list into a map with one of seven verbs on every row
- Say which four things belong in your primary navigation, and what you demoted to make room
- Spot a link Google cannot follow by reading the markup rather than by clicking it
- Decide against a URL migration and give the reason out loud
- Diagnose a cannibalization case as an architecture decision and pick between merge, differentiate and redirect
- Grade an architecture claim as documented, on the record, inference or invented
- Draw an SEO site map on one page and hand it to somebody who could build from it
Pick a path
Thirty-three chapters is a lot in one sitting. Choosing a path collapses the ones outside it, and you can open any of them anyway.
Your progress
0 / 0
I have been doing this since 2010. I have owned and run more than a hundred sites, lost two of them entirely to Panda and Penguin, bought and tested more than 500 SaaS products with my own money, and taught this to over 30,000 students. I have also, on this domain, in the week I wrote this page, deleted a lesson four days after publishing it, merged a post into its own category page, and replaced a redirect file with forty-six rules in it. All three of those are architecture decisions arriving late, and the bill for late is always a redirect.
Here is the mistake this page exists to stop, and it survives because it looks like competence. You draw a folder tree, keep it three clicks deep, write readable slugs, and believe you have designed an architecture. You have designed an address scheme. Whether it is an architecture depends entirely on whether anything links along it, and Google says so in a sentence almost nobody quotes.
"Google generally doesn’t look at the structure of URLs to work out the structure of a site. Instead, it analyzes the linkages between pages to gain insights about the relative importance of different pages on a site."
That is from Google’s own ecommerce site structure documentation, read on 22 August 2026 and last updated by Google on 10 December 2025. Read it twice, because it reorders the entire subject. It means the folder tree is a description, the link graph is the architecture, and a site can have a flawless one and none of the other.
So the opponent for this whole page is the folder diagram, and I can show you exactly what
it costs. I have 103 commercial tool pages sitting in
/seo-tools/{tool}/. Consistent slugs. Two clicks from the home
page. Canonical tags, sitemap entries, breadcrumbs, the lot. On 22 August 2026 I ran the link
graph and 97 of them had no editorial link pointing at them. Not a broken link.
No link. Nothing anywhere on my site argues that those pages should be read, and every
audit tool I own reports that section as healthy.
This is step five of the roadmap and it assumes the four before it. It assumes you know which surfaces your buyers use, from search everywhere optimization, that you can read a crawl report, from SEO fundamentals, that you arrive with clusters rather than a blank page, from keyword research, and that you know what kind of page each cluster deserves, from search intent and SERP analysis. None is a hard prerequisite. The fourth one in particular makes everything here easier, because this page decides where a page lives and that one decides what kind of page it is.
The framework
The whole lesson, in six decisions
Six questions, in order, and the order is not decorative. You cannot pick a parent before you know what kind of page it is, you cannot design a route to a page you have not decided to build, and you cannot rank six things by importance until all six exist on paper. Steps one and two arrive from the two lessons before this one. Publishing comes after all six, which is the only genuinely controversial thing on this card.
- 01
Demand
Is there something here worth satisfying?
It arrives from step three as a cluster with a number on it. Architecture does not create demand and it cannot rescue a cluster that has none. The only work at this step is throwing rows away, which is the step people skip because it produces nothing to show.
You end up holding A shortened list, with the rows you deleted still visible and a reason next to each.
01 Architecture is the link graph 02 Five layers people call one thing 03 Why this step is fifth
- 02
Intent
What would actually satisfy it?
It arrives from step four as a brief: a dominant page type with a count, a format, an angle. If you skipped that step you are about to decide where a page lives without knowing what kind of page it is, which is the most expensive order to do this in.
You end up holding One brief per cluster, with a page type and a verb on it.
04 Cluster, intent, destination 05 One page or several 06 Cannibalization is an architecture decision
- 03
Page
Which destination satisfies it, and does it already exist?
The first genuinely architectural question. Not "what do I write" but "how many URLs does this need, of what kind, and do I have one already". Two clusters can share a destination. Twelve clusters can share a destination. Occasionally one cluster needs three.
You end up holding A keyword-to-URL map: cluster, intent, page type, existing URL, verb.
07 Page types, and the queries that want them 08 The keyword-to-URL map 09 The pages you should not build
- 04
Parent
Where does that page belong, and what does it belong under?
Grouping, and the test every group has to pass: if somebody lands on the parent, does having these children collected there help them finish something? A parent that exists because a framework said to have a pillar page is a filing cabinet with a URL.
You end up holding A tree. Named parents, assigned children, and the overlaps marked rather than hidden.
10 Four architectures, four business models 11 The test a parent page has to pass 12 Categories, taxonomy and the debt 13 Hubs without the mythology 14 Topic clusters are a model, not a formula 15 Silos leak, and that is the point 29 Competitor structure is evidence, not instructions
- 05
Path
How does anybody, or anything, get there?
Navigation, breadcrumbs, contextual links, URLs. This is the step where the drawing becomes a website, and the step most architecture advice skips by assuming that a folder implies a route. It does not. Google reads the links.
You end up holding At least one crawlable route to every page you care about, written down.
16 What earns a permanent nav slot 17 The mobile menu that hides half your links 18 A link Google cannot follow is not a link 19 URL architecture 20 URL hierarchy is not conceptual hierarchy 21 Subfolders, subdomains and root clutter 22 Breadcrumbs, and what they are not 23 Faceted navigation and pagination
- 06
Priority
How prominently should the structure treat it?
Everything cannot be important. Nav slots, link counts and click depth are the site saying out loud what it thinks matters, and on most sites what it says contradicts what the owner would say. Fixing that contradiction is the cheapest hour in SEO and it produces no content.
You end up holding A ranked list, and the pages you demoted named alongside the ones you promoted.
24 Click depth, and the three-click religion 25 Does the structure treat this page as important? 26 Orphans, dead ends and the sitemap that hides them 27 Internal links do three jobs 28 Building for 200 pages, not 20 30 One structure for humans, crawlers and retrieval 31 SEO architecture debt 32 The mistakes I would kill first 33 Test yourself
Then publish. Not before. The whole argument for putting this lesson fifth is that a page whose parent, path and priority were decided after it shipped is a page whose architecture is an accident.
Those six are the model, and the order is load-bearing. You cannot pick a parent before you know what kind of page it is. You cannot design a route to a page you have not decided to build. And you cannot rank six things by importance until all six exist on paper. Every chapter below sits under one of the six, and the chips on each card jump to the chapters that serve it.
Steps one and two are the outputs of the two previous lessons, which is worth saying out loud: architecture is not a fresh start. If you skipped keyword research and search intent, this page will still teach you the mechanics and you will be applying them to guesses.
Watch first, two minutes
Google, asked for the preferred site structure, declining to name one
The question this whole page answers, put to Google directly. Notice what the answer is about: usefulness, and links. Notice what it is not about, which is the shape of a tree or the depth of a folder. Then read chapter one, which quotes the sentence that makes this answer make sense.
Published by Google Search Central. There are 68 videos from 8 channels in the library further down, filterable, and every one of those channels either owns a search surface or owns a crawler you might use to find your own orphan pages.
Site architecture is your link graph, not your folder tree
If you remember one thing Google says it works out your site’s hierarchy from the links between your pages, not from your URL folders. So the tree you drew is a drawing until something links along it.
Not in this path.
Your architecture is the set of links between your pages. Not the folders those pages sit in. Google says so in one sentence in its ecommerce documentation, and that sentence is the single most useful thing published about this subject: it generally does not look at URL structure to work out the structure of a site, and analyzes the linkages between pages instead.
The mechanism is worth understanding rather than accepting, because it explains everything downstream. A crawler arrives at a page, extracts every anchor element with an href, and adds those URLs to a queue. It does that recursively. What it ends up with is a directed graph of your site, and the shape of that graph is the only structural information it has. A folder name is a string in a URL. It is not an edge.
Google’s own statement on discovery makes the same point from the other direction: the vast majority of new pages it finds every day are found through links. Not through folders, and not primarily through your sitemap. Links are the transport layer for the entire index, and your architecture is either made of them or it is a drawing.
Here is the same nine URLs twice. Identical folders, identical slugs, identical sitemap. On the left nothing links along them, on the right everything does. To a URL audit those two columns are the same site.
Same folders, same URLs, and only one of them is an architecture
Nine real addresses from this site, drawn twice. Nothing in the left column is misspelled, inconsistent or too deep. Every slug is readable, every folder is logical, and the sitemap for both sides is the same file. The difference is the arrows, and Google says the arrows are what it reads.
Folders only A tree nobody linked along
Scroll sideways
Every URL is fine. The only route in is the directory grid, so the site has said nothing about which of these matters, and a crawler has nothing to infer a hierarchy from.
Linked The same tree, with a graph on it
Scroll sideways
Identical addresses. Now there is a hierarchy to infer, a relative importance to read, and one page that already earns attention routing some of it downward.
This is not hypothetical, it is the left-hand column
On 22 August 2026 the link-graph tool found 97 of my 103 tool listings with no
editorial link pointing at them. Every one of them sits in a consistent
/seo-tools/{tool}/ folder, two clicks from the home page, in the sitemap, with a
readable slug and a canonical tag. By every measure a URL audit reports on, that section is
immaculate. By the measure Google says it actually uses, four fifths of it is a list.
Now the receipt, because a page arguing this has no business making you take it on faith. I ran the link graph of this site twice, a day apart, with a tool that lives in a public repository. Both readings are below and one of them is embarrassing.
The receipt, on my own site
Two link-graph readings, one day apart, both published
node tools/authority.mjs against this site’s content collections. It reads every markdown file,
extracts the hand-written internal links and reports who points at what. The first reading is why
this page exists. The second is a day later, after the cheap half was fixed and the expensive half
was not.
21 August 2026 the reading that started it
148
content pages
13 posts, 103 listings, 19 spokes, 13 resources
269
hand-written internal links
Between content pages. Navigation and footer links are excluded on purpose.
118 / 122
money pages with no link from a post
97% of the commercial pages on the site.
1 / 13
posts linking to their own category hub
Against a written rule saying all of them.
The reading that started this. Every draft the writing system had produced passed its own quality gate, and the site still had 103 commercial pages that no argument anywhere on it pointed at.
22 August 2026 one day later
147
content pages
11 posts, 103 listings, 19 spokes, 14 resources
228
hand-written internal links
Between content pages. Navigation and footer links are excluded on purpose.
118 / 122
money pages with no link from a post
97% of the commercial pages on the site.
11 / 11
posts linking to their own category hub
Against a written rule saying all of them.
One day later. Every post now links to its own category hub, which took an afternoon. The money-page number did not move at all, which is the part that takes months and the reason this page exists.
Fixed in an afternoon
1 of 13 → 11 of 11
Every post now links to its own category hub. This was one sentence added to eleven files, and it is the kind of fix that makes an architecture audit feel productive.
Did not move at all
118 of 122 → 118 of 122
Identical. Routing authority to 118 commercial pages means finding, for each one, a page that already gets attention and has an honest reason to mention it. There is no version of that which takes an afternoon, and no plugin that does it without producing noise.
Pages nothing on the site points at
-
97 of 103 Tool listings
Reachable through the directory grid and the sitemap. Nothing on the site argues for them.
-
4 of 19 Tool spokes
All four are Semrush depth pages that the hub lists but no post reaches for.
-
3 of 11 Blog posts
Reachable from the category index. No other post recommends them.
-
2 of 14 Resources
One of the two is the free keyword research template, which is the most valuable thing on the site to give away.
All of them are reachable. Every one is in the sitemap, most are two clicks from the home page, and a click-depth report would give this section a clean bill of health. Reachable and routed to are different states, and only one of them is a claim about importance.
Where the links actually go
- 28 /resources/technical-seo-checklist/ resource
- 22 /seo-tools/semrush/pricing/ spoke
- 21 /resources/seo-report-template/ resource
- 15 /seo-tools/semrush/ listing
- 10 /seo-tools/semrush/tutorial/ spoke
This is the useful half of the same export. The most linked-to page on the domain is a free checklist and the second is a pricing page, and both are correct. Then nothing commercial appears again for a long way, which is the gap between what the graph says matters and what I would say.
Depth pages per tool listing
-
semrush18 -
google-search-console1 -
the other 101 listings0
That asymmetry looks like neglect and is mostly correct. One tool has eighteen sub-topics with their own demand. Forcing symmetry across 103 listings would produce a hundred thin pages, which is the mistake in the other direction.
Run this against your own site before you read another chapter. Any crawler has an orphan or "no inbound internal links" report and almost nobody opens it. Screaming Frog is free to 500 URLs. If you have Search Console and nothing else, the indexing report plus a sitemap count gets you most of the way. The number will be worse than you expect, and that is the finding.
The important thing about that pair of readings is not the 97. It is that the post-to-hub number went from 1 of 13 to 11 of 11 in a day, and the money-page number did not move by a single page. Fixing relationships you already have is an afternoon. Routing attention to 118 commercial pages means finding, for each one, a page that already earns attention and has an honest reason to mention it. There is no plugin for the second thing, and every plugin that claims to be one produces links nobody meant.
One clarification before the rest of the page, because I am about to spend a chapter on
URLs and I do not want it read as a contradiction. Folders are not worthless. Google
documents one specific benefit: grouping similar topics into directories helps it learn how
often the URLs in each directory change, so a /policies/ folder and a
/promotions/ folder can be crawled at different rates. That is real, it is
useful, and it is about crawl scheduling. It is not a relevance mechanism and it is not
worth a migration.
Five different things share the word architecture
If you remember one thing Five different objects share one word. Information architecture, site architecture, URLs, navigation and internal linking. Most architecture arguments are two people naming different layers.
Not in this path.
Information architecture, site architecture, URL structure, navigation and internal linking are five different objects, and almost every argument about architecture is two people diagnosing different rows.
The reason to separate them is not tidiness, it is diagnosis. "Our site architecture is fine, look at the URL structure" is not an argument, it is a category error, and it is the sentence that gets a team to spend six weeks renaming folders while the pages nobody links to stay unlinked. Each layer owns different decisions and fails in a different way, and the failure column is the one worth memorizing.
Five layers, one word
What people mean when they say "architecture"
Top to bottom, from the most abstract to the most specific. Each row owns a different set of decisions and fails in a different way, and the last column is the useful one: it tells you which layer a symptom belongs to. Most architecture arguments are two people diagnosing different rows.
- 01
Information architecture
What information exists, and the rules used to classify it.
What are the categories, and what does each one mean?
It owns
The taxonomy. The vocabulary. The decision that "guides" and "tutorials" are one thing or two.
When this layer is wrong and the others are right
You end up with five category names that mean the same thing and no rule for choosing between them. Every new page is a judgement call, and two people make it differently.
- 02
Site architecture
How the pages that exist are arranged and connected.
What is the parent of this page, and what links to it?
It owns
The tree, the hubs, the parent and child relationships, the link graph.
When this layer is wrong and the others are right
Pages exist and nothing points at them. This is the failure this whole page is about, and it is invisible in every tool that reports on URLs rather than on links.
- 03
URL structure
How the addresses are written.
Does this URL read like the thing it points at?
It owns
Folders, slugs, parameters, trailing slashes, casing.
When this layer is wrong and the others are right
Ugly addresses, and mildly worse click-through from a visible URL in a result. It is the layer people fix first and the one that carries the least weight.
- 04
Navigation
The permanent, site-wide routes between major areas.
What is on every page of the site, and why that?
It owns
Header, footer, sub-menus, breadcrumbs, the mobile menu.
When this layer is wrong and the others are right
The most important pages are reachable but not surfaced, and the site tells every crawler that a legal page matters as much as the service page.
- 05
Internal linking
The contextual, in-content links between individual pages.
When I finish this page, where should I be able to go next?
It owns
Anchor text, related links, next-step links, hub-to-spoke and spoke-to-hub.
When this layer is wrong and the others are right
Dead ends. A reader finishes and leaves, and a page that ranks passes nothing to the page that needs it.
Rows two and five are the expensive ones and row three is the one that gets fixed first, which is most of what is wrong with how this subject gets taught. Ugly URLs are visible in a screenshot. An unlinked commercial page is visible in an export nobody runs.
One more distinction, and it is the one beginners most often collapse. An SEO site map is a planning diagram: the pages that should exist and how they relate. An XML sitemap is a machine-readable list of URLs. They are not the same document, they are not produced by the same activity, and the second one is not evidence that you did the first. I have watched teams generate the XML file, believe the architecture work was done, and never draw the diagram that would have shown them the orphans. This site’s own build even prunes noindex URLs out of its sitemap, because eighteen category pages were being advertised in there while asking Google to ignore them.
Why this comes after intent and before any content
If you remember one thing Demand tells you what people want, intent tells you what would satisfy them, and this step decides where that lives. Any other order and you are placing pages before you know what they are.
Not in this path.
Architecture belongs between deciding what would satisfy a searcher and deciding what to write, because it is the last step that can still say no cheaply.
Run the chain out. Keyword research gives you what people want. Search intent gives you what would satisfy them, as a page type with a count behind it. Architecture asks where that page belongs, what will link to it, and how important the site should treat it. Content then fills it in. Every step narrows, and the cost of reversing a decision goes up at every stage: deleting a row from a spreadsheet is free, deleting a published page costs a redirect and a small amount of trust.
The usual order is keyword research, then write, then technical SEO. That order has one structural flaw and it is fatal: the page and its place are decided by the same person at the same moment, which means the place is decided by whatever folder was convenient. Do it a hundred times and you have a redirect file. Mine has 47 rules in it.
And a caveat I would rather give you than have you discover: this is not the highest-impact step in the roadmap for most sites. Google says in its own starter guide, immediately after the site organization advice, that search engines will likely understand your pages as they are right now regardless of how the site is organized. That sentence is almost never quoted by anybody selling a restructure. It does not mean architecture is worthless. It means the gains are in discovery, prioritization and human navigation rather than in Google failing to comprehend you, and a restructure sold on comprehension is sold on the wrong thing.
A cluster is not a page
If you remember one thing A cluster is not a page. The unit that becomes a URL is a need, and several clusters routinely share one.
Not in this path.
The unit that becomes a URL is a need, and several clusters routinely share one. That sentence is the whole bridge between the previous lesson and this one, and getting it wrong is the most common way a keyword list turns into architecture debt.
Watch the failure happen. Six keywords come out of research. Six rows in a spreadsheet, six lines in a content calendar, six articles. Nobody decided that; the spreadsheet decided it, because a row looks like a page. Six weeks later there are six URLs, four of them answering the same question slightly differently, and Google is picking between them on evidence you did not choose.
The architectural version of the same six keywords looks like this instead:
SEO Tools Comparisons │ │ ├── Keyword Research Tools ├── Ahrefs Alternatives ├── Rank Tracking Tools └── Ahrefs vs Semrush ├── Technical SEO Tools └── Free SEO Tools
Same demand. Now it is a structure, and two decisions have been made that the spreadsheet could not make: that four of the six are slices of one taxonomy, and that the two comparisons are a different job and therefore a different parent. Neither decision came out of a volume column.
What arrives from step four is a page type with a count on it. "Tool, five of eight" is a usable input to this step. "Commercial" is not, and neither is a volume number, because neither tells you how many destinations the demand supports.
One page or several, and the free second opinion
If you remember one thing Four questions decide it: same job, same page type, overlapping results, and can one page do both naturally. Ahrefs will hand you a fifth for free in the parent topic column.
Not in this path.
Four questions decide it, and a keyword tool will hand you a fifth for free in a column most people scroll past.
The rule is not one keyword, one page. It is one intent, one destination, and the difference is worth thousands of pounds a year to anybody publishing at volume. Two phrasings of one need share a page. Two needs that happen to share vocabulary do not.
Commit before you look
One page or two: six real pairs
Every number here came out of Ahrefs on 22 August 2026. Read each pair, decide, then open the answer. Two of the six are cases where I overrule the parent topic column, and both say so, because the interesting part of this skill is knowing when the tool is describing the market rather than your argument.
- 1
Same job?
Not the same topic. The same thing the person is trying to finish.
- 2
Same page type?
A guide and a tool are never one page, however related the subject.
- 3
Do the results pages overlap?
Search both. If six of ten URLs are the same, Google already thinks it is one need.
- 4
Could one page satisfy both naturally?
Naturally. Not "if we add a section", which is how a page becomes two pages stapled together.
- 5
What does the parent topic column say?
The free fifth test. Same parent means the page currently winning both is one page.
Pair 01
seo site architecture 250 vol TP 2,800
parent website structure
seo website structure 150 vol TP 2,700
parent website structure
The column says same parent, one destination
What I would do
one page
Same parent, near-identical traffic potential, and the results pages overlap heavily. This is the easy case and it is on the list so the hard ones have something to be compared against.
Pair 02
website structure 3,900 vol TP 4,100
parent website structure
seo site architecture 250 vol TP 2,800
parent website structure
The column says same parent, one destination
What I would do
one page
The machine says one page, and it is right about the demand and wrong about the market. Four of the eight results for the bigger term belong to design and website-builder companies and its AI Overview cites Figma, MDN and a sitemap tool with no SEO source at all. So: one page, and a decision about which of the two audiences it is for. I picked the SEO one, which means giving up fifteen times the volume on purpose.
Pair 03
orphan pages 450 vol TP 1,200
parent orphan pages seo
click depth 60 vol TP 10
parent click depth
The column says different parents, two destinations
What I would do
one page, one section overruling the column
Two different parents, and only one of them earns a URL. Orphan pages has a traffic potential of 1,200 and a real diagnostic workflow behind it. Click depth has a traffic potential of 10. A distinct parent topic is permission to consider a page, not an instruction to build one.
Pair 04
topic clusters 600 vol TP 1,500
parent topic cluster
silo structure seo 150 vol TP 150
parent seo silos
The column says different parents, two destinations
What I would do
one page overruling the column
The machine sees two parents because two vocabularies grew up around one idea, ten years apart. A reader searching either one wants to know how to group related pages without cutting them off from each other. I am overruling the column with an argument, which is allowed as long as you say you are doing it.
Pair 05
keyword mapping 1,400 vol TP 1,400
parent keyword mapping
seo site architecture 250 vol TP 2,800
parent website structure
The column says different parents, two destinations
What I would do
two pages
Different parents, difficulty 23 against 28, and a genuinely different deliverable: a mapping page is a spreadsheet method, this page is a system of decisions. This is the row on my own list that I would build next, and it is chapter eight here in the meantime, which is a deliberate temporary answer rather than a permanent one.
Pair 06
subdomain vs subfolder 400 vol TP 1,000
parent subdomain vs subdirectory
url structure 1,400 vol TP 30,000
parent what is a url
The column says different parents, two destinations
What I would do
two pages
Different parents and, more usefully, different difficulty by a factor of eight: 8 against 68. And look at what the parent column does to the second one. "url structure" rolls up to "what is a url" at 186,000 searches, which tells you the number one page for it is a general web-literacy page and that a technical SEO article is entering somebody else’s market.
The fifth test is the one nobody teaches, so here it is properly. Ahrefs computes a parent topic by taking the number one ranking page for your keyword and finding the keyword that sends that page the most traffic. So it is not a clustering heuristic invented by a tool. It is an observation about what currently wins, which makes it a direct answer to the question this chapter asks, published free, in a column next to the volume figure everybody reads instead.
Run it on the subject of this page and it makes my argument better than I can.
The one-page-or-many test, run live
8 ways of saying one thing, and Ahrefs says 4 pages
One Keywords Explorer request, US database, 22 August 2026. Eight of the twenty keywords I pulled are phrasings of the same subject and they carry 8,550 US searches a month between them. Read the parent topic column rather than the volume column: it collapses those eight into 4 destinations, and it did that without being asked.
Parent topic website structure 3,900 vol 4 phrasings
-
website structure3,900 KD 25 TP 4,100 -
website architecture900 KD 49 TP 2,800 -
seo site architecture250 KD 28 TP 2,800 -
seo website structure150 KD 52 TP 2,700
Parent topic seo site structure 700 vol 2 phrasings
-
site structure1,400 KD 41 TP 700 -
site architecture seo500 KD 30 TP 2,600
Parent topic site architecture 1,000 vol 1 phrasing
-
site architecture1,100 KD 50 TP 200
Parent topic website architecture 1,200 vol 1 phrasing
-
website architecture seo350 KD 16 TP 2,700
What the column is telling you
- Four destinations, not eight. Half of the eight roll up to
website structure, which means the page currently winning those four is one page and it is the same page. - The parent is not always the biggest sibling.
site architectureis its own parent at 1,100 searches, and its traffic potential is 200. Volume 1,100, potential 200. Whatever the number one result for that query is, it barely earns anything, and it is the only keyword of the twenty that also came back commercial. That combination is usually a query about software architecture rather than about SEO. - An inconsistency worth printing. Three rows report their parent as
website structurewith a parent volume of 4,900, and the row forwebsite structureitself reports 3,900. One response, two numbers for the same keyword. It changes nothing about the decision, and a page that asks you to read this column carefully should say when the column disagrees with itself. - You are allowed to overrule it. Twice on this page I do, and both times the reason is written down. A tool that observes what currently wins cannot know that two vocabularies grew up around one idea ten years apart.
The other 12 keywords from the same request all 20 of 20 informational
| Keyword | Vol | KD | Parent topic | TP | Intent labels | Results page crawled |
|---|---|---|---|---|---|---|
information architecture | 5,400 | 60 | information architecture | 2,200 | informational | 11 Aug 2026 |
internal linking | 1,600 | 68 | internal links seo | 11,000 | informational | 2 Aug 2026 |
url structure | 1,400 | 68 | what is a url | 30,000 | informational | 17 Aug 2026 |
keyword mapping | 1,400 | 23 | keyword mapping | 1,400 | informational | 12 Aug 2026 |
faceted navigation | 1,200 | 42 | faceted navigation | 500 | informational | 19 Aug 2026 |
keyword cannibalization | 1,100 | 53 | keyword cannibalization | 3,300 | informational | 3 Aug 2026 |
breadcrumbs seo | 900 | 44 | breadcrumbs seo | 1,900 | informational | 2 Aug 2026 |
topic clusters | 600 | 51 | topic cluster | 1,500 | informational | 7 Aug 2026 |
orphan pages | 450 | 20 | orphan pages seo | 1,200 | informational | 4 Aug 2026 |
subdomain vs subfolder | 400 | 8 | subdomain vs subdirectory | 1,000 | informational | 7 Aug 2026 |
silo structure seo | 150 | 4 | seo silos | 150 | informational | 29 Jul 2026 |
click depth | 60 | 0 | click depth | 10 | informational | 9 Aug 2026 |
Every one of the twenty came back informational. One, site architecture, also came back
commercial. Zero transactional, zero navigational, zero local, zero branded. The intent column is a
filter for sorting ten thousand rows and it is not an answer, which is the same conclusion the
previous lesson reached from a completely different keyword set.
Eight phrasings, 8,550 searches a month, four parent topics. A volume-first process produces eight pages from that list and every one of them competes with three siblings. The parent topic column produced four destinations without being asked, and it took one request.
You are allowed to overrule it, and I do twice in the drill above. What you are not allowed to do is overrule it silently. A tool that observes what currently ranks cannot know that two vocabularies grew up around one idea a decade apart, which is exactly the case with "topic clusters" and "silos". Write the override down with the reason, and the next person who opens your map does not undo it.
Cannibalization is a decision you already made
If you remember one thing Nothing is being punished. You built two destinations for one need and now the site cannot say which one it means, so Google decides on its own evidence.
Not in this path.
Nothing is being punished. You built two destinations for one need, so your site cannot say which one it means, and Google decides on its own evidence rather than on your intentions.
The mechanism is canonicalization, not penalty. When several URLs answer the same need, Google picks one to represent the set, using signals it can see: internal links, external links, sitemap presence, redirects, content similarity. Your preference is one input among several and it is often the weakest one. So the page it picks is the page your architecture actually endorsed, which is frequently not the one you would have chosen.
Which makes this an architecture chapter rather than a technical one. The four fixes are all structural: merge into the strongest page and redirect, differentiate so each answers a genuinely different job, redirect the weaker one and stop, or reposition one of them onto a need you were not covering. Canonical tags are how you tell Google about the decision. They are not the decision.
This site prevents the case rather than fixing it, and it is the most useful thing in this chapter. Every depth page under a tool listing declares in its frontmatter the one query it exists for, and the build compares every claim across every collection at once. Two pages claiming the same query is a build error naming both owners, so it cannot ship. That is a twenty-line function and it has caught more cannibalization than any amount of editorial care would have at page fifty.
primaryQuery: "semrush pricing" ← /seo-tools/semrush/pricing/
primaryQuery: "semrush pricing" ← a new post
build fails, both owners named You do not need my stack to do this. A shared spreadsheet with one row per query and one owner per row does the same job, as long as somebody checks it before publishing rather than after. The mechanism that matters is that the check happens before the URL exists.
Not every query wants an article
If you remember one thing Search intent picks the page type before content picks a word. Twenty-nine types exist and most content operations can only produce one, which is why so many queries get an article they did not ask for.
Not in this path.
Search intent picks the page type before content picks a single word, and there are far more types than most content operations can produce.
This is the highest-cost mistake in the whole subject, and it hides behind a process that feels rigorous. Research produces a cluster. The cluster goes into a content calendar. The content calendar produces a brief. The brief produces an article. At no point does anybody ask whether the market wanted an article, and if it wanted a free tool or a category page or a comparison, no length of article gets into it. The previous lesson has the receipt for that: ten of ten results for "mortgage calculator" are calculators.
So here is the register. 29 types, grouped by what they are for, with the query shape that asks for each one.
The register
29 page types, and the query shapes that ask for them
Filter by group, or read it as a list. The middle column is the one that makes this a decision table: it is the shape of query that wants this kind of page. 15 of the 29 carry a live example from this site so you can check the claim by clicking. The ones with no example are types I do not currently have, which tells you something about what this site can produce.
Core The pages a business has because it is a business. 3
Commercial Pages whose job is to sell the thing. 6
-
Service page
One service, one buyer, one outcome.
Wanted by "<service>", "<service> agency", "hire a <role>".
/services/saas-seo-consultant/ -
Product page
One buyable thing.
Wanted by A product name, a model number, "buy X".
Not on this site
-
Feature page
One capability of a product, in depth.
Wanted by "<product> <feature>", "software that does X".
Not on this site
-
Use case page
The product framed for one job.
Wanted by "X for Y", "how to do Y with X".
Not on this site
-
Industry page
The product framed for one sector.
Wanted by "<thing> for <industry>".
Not on this site
-
Pricing page
What it costs and what changes with tier.
Wanted by "<brand> pricing", "how much does X cost".
/seo-tools/semrush/pricing/
Evaluation Pages for somebody choosing between options. 5
-
Comparison page
Two named things, side by side.
Wanted by "X vs Y".
Not on this site
-
Alternatives page
One named thing, and what replaces it.
Wanted by "X alternatives", "sites like X".
/seo-tools/semrush/alternatives/ -
Review
One thing, tested, with a verdict.
Wanted by "X review", "is X any good".
Not on this site
-
Best-of list
A ranked set, with criteria.
Wanted by "best X", "top X for Y".
Not on this site
-
Directory
Many things, filterable, with no ranking claim.
Wanted by "X tools", "X software", "list of X".
/seo-tools/
Educational Pages for somebody learning. 5
-
Guide
A subject explained end to end.
Wanted by "what is X", "X guide", "X for beginners".
/seo-basics/ -
Tutorial
One procedure, start to finish.
Wanted by "how to X", "X tutorial", "X step by step".
/seo-tools/semrush/tutorial/ -
Definition
One term, answered in a paragraph.
Wanted by "X meaning", "X definition", "what does X mean".
Not on this site
-
Original research
A number nobody else has.
Wanted by "X statistics", "X study", "X data".
Not on this site
-
Case study
One engagement, with the exports.
Wanted by "X case study", "X results", "does X work".
/seo-case-studies/saas-seo/
Utility Pages that do something rather than say something. 4
-
Free tool
Input, output, no reading.
Wanted by "X checker", "X generator", "X calculator".
/seo-tools/free/word-counter/ -
Calculator
Numbers in, a number out.
Wanted by "X calculator", "how much X".
Not on this site
-
Template
A file somebody opens and fills in.
Wanted by "X template", "X spreadsheet", "free X".
/resources/keyword-research-template/ -
Checklist
An ordered set of things to tick.
Wanted by "X checklist", "X audit checklist".
/resources/technical-seo-checklist/
Support Pages for somebody who already bought. 3
-
Documentation
Reference for somebody using the product.
Wanted by "how do I X in Y", exact feature names.
Not on this site
-
Help article
One problem, one fix.
Wanted by "X not working", "X error", "fix X".
Not on this site
-
FAQ page
Grouped short answers, when they genuinely belong together.
Wanted by Rarely its own query. Usually a section.
Not on this site
Discovery Pages whose job is to route to other pages. 3
-
Category page
One slice of a taxonomy, listing its members.
Wanted by "<category> tools", "<category> for X".
/seo-tools/category/rank-tracking/ -
Hub page
A subject, and the routes into it.
Wanted by The head term of a topic.
/seo-roadmap/ -
Tag or filter page
A cross-cut of the taxonomy. Useful rarely, indexable rarely.
Wanted by Almost never its own query. Assume no until proven.
Not on this site
The exercise that hurts. Skim a hundred of your own URLs and write one word from this register next to each. Most sites come back with two or three types covering ninety percent of the list. Then take the last ten clusters your team decided to target and ask what type each one wanted. Every mismatch is either a reading of the market or a description of what your process can make, and it matters enormously which.
The column that turns this from a taxonomy into a decision table is "wanted by". A query beginning "best" wants a ranked list with stated criteria. A query ending "checker" wants an input box. A query of the shape "X vs Y" wants two named things side by side, and a guide that mentions both is not the same object.
Nine of the 29 types have no example on this site, and I left that visible rather than filling the gaps with something plausible. It tells you what this site can currently produce, which is the same finding the exercise below will produce about yours.
The keyword-to-URL map, and the verb that makes it a plan
If you remember one thing The deliverable of this step is not a tree, it is a table: cluster, intent, page type, existing URL, parent, verb. The verb is what makes it a plan.
Not in this path.
The output of this step is a table, not a tree. Cluster, intent, page type, existing URL, parent, verb. The tree comes later and it is generated from the parent column.
Two columns do almost all the work and both get skipped. The existing URL column is asked for before the verb, which is the only thing that reliably stops somebody creating a page the site already has. And the verb column is what makes the document a plan: a row with no verb is research, and research does not get built.
Here is mine, from this site, with the rows that say no in it.
The deliverable of step three, worked
My own keyword-to-URL map, with the rows that say no
Fifteen real rows from this site. 13 of the 15 are not "create", which is the whole reason it is worth showing you mine rather than a clean example. Two of them are me declining traffic on purpose, one is a post I merged, one is a pillar I redirected, and one is a page I deleted four days after publishing it.
| Cluster | Intent | Page type | Existing URL | Parent | Verb | Why |
|---|---|---|---|---|---|---|
learn seo in order | Educational | Hub | /seo-roadmap/ | Home | Keep | The parent of the whole curriculum, and it passes the parent test: arriving there tells you what to read first. |
seo basics | Educational | Guide | /seo-basics/ | /seo-roadmap/ | Keep | Step two. Nothing to decide. |
keyword research | Educational | Guide | /seo-keyword-research/ | /seo-roadmap/ | Keep | Step three. |
search intent | Educational | Guide | /search-intent/ | /seo-roadmap/ | Keep | Step four, published the day before this page. |
seo site architecture | Educational | Guide | None | /seo-roadmap/ | Create | This page. Step four explicitly handed off to it, and a handoff to nothing is not a handoff. |
website structure | Educational | Guide | None | n/a | Ignore | Fifteen times the volume of the row above and lower difficulty. Four of its eight results are design and site-builder companies. It is a planning market, not an SEO one, and entering it would mean writing for somebody I am not. |
keyword mapping | Educational | Guide | None | /seo-roadmap/ | Create | Its own parent topic, difficulty 23, and a distinct deliverable. Queued, not built. Chapter eight of this page is the interim answer and it says so. |
seo tools | Commercial research | Directory | /seo-tools/ | Home | Keep | The commercial spine of the site. 103 listings under it. |
rank tracking tools | Commercial research | Category | /seo-tools/category/rank-tracking/ | /seo-tools/ | Keep | A real slice of the taxonomy with its own demand. It also absorbed a retired best-of post, which is the merge row below. |
best seo rank tracker tools | Commercial research | Best-of list | /seo-tools/best-seo-rank-tracker-tools/ | n/a | Merge | A list post competing with its own category page for one need. Merged into the category and redirected. It is line 36 of the redirect file. |
accuranker review | Commercial | Listing | /seo-tools/accuranker/ | /seo-tools/ | Improve | Two clicks from the home page and zero editorial links. One of the 97 listings nothing on the site argues for, and the fix is a link from a post rather than a rewrite. |
keyword research template | Utility | Template | /resources/keyword-research-template/ | /resources/ | Improve | The most valuable giveaway on the site and one of only two resources with no inbound editorial link. Routing, not writing. |
keyword research | Educational | Guide | /keyword-research/ | n/a | Redirect | The old pillar. Two URLs for one need, and the lesson version has the sequence around it. Redirected to /seo-keyword-research/, last line of the redirect file. |
seo digital marketing | Educational | Guide | /seo-digital-marketing/ | n/a | Remove | It was roadmap step four for one day and step three before that. On the day it was pulled Ahrefs showed 0 organic keywords and 0 traffic. Redirected to /seo-basics/ and archived. Deleting a page you shipped four days ago is cheaper than defending it for four years. |
serp analysis | Educational | Guide | None | n/a | Ignore | More volume and lower difficulty than "search intent", and the top two organic results are free rank checkers. An article does not inherit a tool’s traffic. Decided on the previous lesson and still true. |
Scroll sideways
The seven verbs, and the mistake each one usually replaces
Keep
When: The page exists, matches the intent, and has a parent and a path.
Cost: Nothing. Write the row down anyway so the next person does not re-decide it.
Usually chosen instead of: Rewriting a page that was fine, because it was on the list.
Create
When: Real demand, a distinct intent, no existing destination, and a business reason.
Cost: The build, plus the maintenance forever, plus a slot in somebody’s attention.
Usually chosen instead of: Creating because the cluster had volume, which is how you get architecture debt with a content calendar’s name on it.
Improve
When: The right page exists and is losing on format, depth or angle.
Cost: Hours, and it is almost always the cheapest verb on the list.
Usually chosen instead of: Creating a second page about the same thing, which is cannibalization you chose.
Merge
When: Two or more URLs answer one need.
Cost: A redirect, a rewrite, and the nerve to delete something you wrote.
Usually chosen instead of: Leaving both and hoping Google picks the right one. It will pick one, and not necessarily yours.
Move
When: The page is right and its parent is wrong.
Cost: A redirect, and every internal link updated rather than left to redirect.
Usually chosen instead of: Moving for tidiness. If the parent is defensible, leave the URL alone.
Redirect
When: The page should not exist and something else answers its need.
Cost: One rule, pointed at the most relevant page rather than the home page.
Usually chosen instead of: A bulk redirect to the home page, which Google treats as a soft 404 and discards.
Remove
When: Nothing answers its need because the need was not real.
Cost: A 410, or a noindex if it still serves a human purpose.
Usually chosen instead of: Keeping it because deleting feels like losing. An unlinked page you will not maintain is already gone.
"Ignore" is not one of the seven verbs and it appears twice in my map, which is a deliberate eighth. The seven are for pages that exist or will exist. Ignore is for a cluster with real demand that you are choosing to leave to somebody else, and writing it down matters more than the other seven, because otherwise somebody rediscovers the cluster next quarter and builds the thing you already decided against.
An invented example would have been tidier and would have taught you nothing, because the only interesting rows in a map like this are the ones where the answer is not "write it". Two of mine decline traffic on purpose. One merges a post I wrote. One redirects a pillar I spent weeks on. One deletes a lesson page I shipped four days earlier, which had zero organic keywords and zero traffic on the day I pulled it, and which was roadmap step four for exactly one day.
A note on the honest cost of "keyword mapping" being a chapter here rather than its own page. It has its own parent topic, difficulty 23, and a genuinely different deliverable from this page, which by my own rules in chapter five means it should be a separate URL. It is queued and not built, so this chapter is a deliberate interim answer, and saying that out loud is cheaper than pretending the decision was principled.
The pages you should not build
If you remember one thing Six questions, and any one of them can kill a page. Five hundred clusters is not five hundred pages, and in the AI era the cost of a page has fallen far faster than the cost of maintaining one.
Not in this path.
Five hundred clusters is not five hundred pages. A page now costs almost nothing to produce and the same as ever to maintain, which makes the filter worth more than it was two years ago.
Six questions. Any one of them can kill a row, and a map where none of them ever does is a map nobody applied a filter to.
- Is there distinct demand? Not demand for the topic. Demand for this page, after the parent topic column has collapsed the phrasings. "Click depth" has sixty US searches a month and a traffic potential of ten, and it is chapter twenty-four here rather than a URL.
- Is the intent meaningfully different from a page I have? If the job is the same job, this is an improvement to an existing page wearing a new URL.
- Can I add something? Not "cover it thoroughly". Add. A number nobody has, a test nobody ran, a decision nobody published. If the honest answer is that you would be summarizing the current top three, the page is a retelling and retellings are what the results page already has too many of.
- Does the business need it? The question no results page can answer. You can match an intent perfectly and gain nothing, and the pages where that happens are usually the ones with the best volume.
- Do I already have something that satisfies it? Search your own site for the job rather than for the words. This is the question the existing URL column exists to force.
- Will this create architecture debt? A page with no obvious parent, no obvious inbound link and no obvious next step is debt on the day it publishes. If you cannot name its parent and one page that will link to it, you are not ready to build it.
The sixth question is the one this era needs most. AI made the marginal page cheap to write and did not make it cheaper to maintain, link, route or eventually delete. A hundred pages generated in a week is a hundred parents to assign, a hundred routes to design, and eventually a hundred lines in a redirect file. The generation was the cheap part.
Four architectures, four business models
If you remember one thing The shape follows the business model. Service, software, catalog and curriculum sites are four different trees, and picking the wrong one is a mistake internal linking will not repair.
Not in this path.
There is no universal tree. A local service business, a software product, a catalog and a curriculum are four different shapes, and the shape follows the business model rather than a best practice.
Semrush says as much: structure depends on site size, goals and users. That sentence is correct and it is usually the last line of the section, immediately before one pyramid diagram that assumes an ecommerce catalog. "It depends" is only useful advice if somebody tells you what it depends on, so here are four shapes with the load-bearing decision named in each.
Architecture follows the business model
Four shapes, and the decision that makes or breaks each one
There is no universal tree. A service business, a software product, a catalog and a curriculum are four different structures, and the crux row is the one worth reading: it is the decision inside each shape that people get wrong, and it is different in every one.
Fits Few pages, high value each, a buyer who wants proof and a phone number.
The shape
-
Home -
Services -
Service A -
Service B -
Locations -
City A -
City B -
Case studies -
Guides -
About -
Contact
The load-bearing decision
Services and locations are two axes, and the temptation is to multiply them. Service A in City A, Service B in City A, and so on. Do that only where the service genuinely differs by place, which is rarer than the template implies.
How it fails
A grid of near-identical pages, one per service and city pair, none of which any human wrote and all of which compete with each other.
Fits One product, many jobs, a buyer comparing you against three alternatives.
The shape
-
Home -
Product -
Features -
Feature page -
Use cases -
Integrations -
Solutions -
By industry -
By role -
Pricing -
Comparisons -
Us vs competitor -
Competitor alternatives -
Resources -
Docs
The load-bearing decision
The comparison and alternatives branch is where the revenue is and where it is most often missing. It is also the branch a content team is least comfortable owning, because it means naming competitors on your own domain.
How it fails
A features list nobody reads, a blog with a thousand posts, and no page that answers "why you rather than them".
Fits Many similar things, a buyer narrowing by attribute.
The shape
-
Home -
Category -
Subcategory -
Product -
Product -
Subcategory -
Category -
Buying guides -
Support
The load-bearing decision
Google documents this exact chain: links from menus to category pages, category to subcategory, subcategory to every product you want indexed. It also says Googlebot generally does not submit searches into a search box, so anything only reachable by searching is not reachable.
How it fails
Filters generating an unbounded URL space while the products themselves are three levels below anything a crawler was invited to follow.
Fits Many pages, one subject, a reader who arrives mid-way and needs the next step.
The shape
-
Home -
Topic hub -
Lesson -
Lesson -
Lesson -
Topic hub -
Tools -
Resources -
About the author
The load-bearing decision
The hub has to be a real destination rather than a table of contents, and every lesson has to point at the next one. This site is this shape, and the sequencing is the whole product.
How it fails
Two hundred posts in reverse chronological order, a category page nobody links to, and no way for a reader to tell what to read second.
The catalog shape is the one Google documents directly, and the chain is specific: links from menus to category pages, category to subcategory, and subcategory to every product you want indexed. It also says Googlebot generally does not submit searches into a search box while crawling, which is the sentence that quietly invalidates a lot of "you can find it by searching" architectures.
This site is the fourth shape, which is why the sequencing between lessons is the product rather than a nicety. A curriculum whose lessons do not know their own order is a blog.
The one test a parent page has to pass
If you remember one thing One test: if somebody lands on this parent, does having these children collected here help them finish something? A parent that exists because a framework wanted a pillar is a filing cabinet with a URL.
Not in this path.
One question decides whether a parent deserves to exist: if somebody lands on it, does having these children collected here help them finish something?
That test fails a surprising number of pillar pages, and it fails them for a specific reason. A page created because a framework says a topic needs a pillar has no job of its own, so nobody links to it voluntarily, so it accumulates no authority to pass down, so it becomes a table of contents with a URL. The children were fine. The parent was scaffolding.
A parent that passes the test is doing one of three things. It is orienting somebody who does not yet know what they need, like a roadmap. It is helping somebody choose between siblings, like a category or a directory. Or it is answering the head question while routing the specific ones, like a good guide. If your parent is doing none of those, its children should be siblings and the parent should be a nav item.
Now the part that is not transferable by explanation. Placement instinct comes from being wrong about a placement and hearing why, so here are eight pages and three of them have a tidy wrong answer.
Where should this page go?
Eight placements, and the tidy answer is wrong three times
Every option below is a real section of this site, so you can check the answers by clicking rather than by trusting me. Commit to one before you open the explanation. The first two items are the whole lesson: two pages that share almost every word and belong in different places, because the job is different.
0
of 8
Pick a home for the first one
-
Show the reasoning
Where it goes B. SEO Tools, as a category page
The job is choosing a tool, not learning a skill. Somebody arriving on this query has already decided to do keyword research and wants to know what to buy. Filing it under Learn puts a purchasing decision inside a curriculum, where the reader has to leave the section to act on it.
The trap It contains the words "keyword research", so it looks like it belongs with the keyword research lesson. Topic overlap is not the same as job overlap.
-
Show the reasoning
Where it goes B. Learn SEO
Same subject, different job, and therefore a different parent. This one is somebody acquiring a skill in sequence, which is exactly what the learning section is for. The pair above and this one are the clearest example on the page of why you group by job rather than by topic.
-
Show the reasoning
Where it goes B. Under the Semrush listing as a depth page
It belongs to one product, so its parent is that product’s listing. On this site it lives at /seo-tools/semrush/pricing/ and it is the second most linked-to page on the whole domain with 22 inbound editorial links, which is what happens when a page has an obvious parent and everything in the section has a reason to reference it.
-
Show the reasoning
Where it goes B. Services
Obvious, and it is on the list because the obvious ones are how you check the rule rather than the instinct. The buyer is transacting, so it goes where the transacting happens. Note that this one needs no research: the job is legible in the query.
-
Show the reasoning
Where it goes A. Resources
A template is a utility page. Somebody wants a file, not an argument, and the resource library is the section whose promise is "here is a working file". This site’s copy of it carries 21 inbound links, second only to the technical SEO checklist, because every post that mentions reporting has an obvious reason to point at it.
-
Show the reasoning
Where it goes C. Its own comparison area, parented to neither
A comparison filed under one of the two brands quietly asserts that the other is subordinate, and it forces every future comparison to pick a side. A comparison area parented to the directory keeps both brands peers. This site enforces the underlying rule at build time: every depth page declares the one query it exists for, and the build fails if two pages claim the same one, so "ahrefs vs semrush" can only have one home.
The trap Both single-brand answers feel tidy, and both create a rule you will break the next time you write a comparison.
-
Show the reasoning
Where it goes B. A section inside the site architecture lesson
On 22 August 2026 "click depth" carried 60 US searches a month with a traffic potential of 10. Ten. That is not a page, and building it would add a URL, a maintenance obligation and an internal link decision in exchange for nothing. It is chapter twenty-four of this page instead.
The trap It is a distinct concept with its own name and its own parent topic, and all three of those things are true of plenty of subjects that should never get a URL.
-
Show the reasoning
Where it goes B. The case studies section
A case study is a page type with its own job: proving the service works. Filing it in the blog puts your strongest commercial evidence in a reverse-chronological list where it ages out of view, and it makes the service pages link into the blog to find their own proof.
If you got the first two the same way round, that is the single most transferable thing on this page. Grouping by topic puts a purchasing decision inside a curriculum. Grouping by job puts it where somebody can act on it. 3 of the eight carry a named trap, and every trap is a version of the same mistake: two pages sharing a subject and not sharing a reader.
The first two items are the whole chapter. "Best keyword research tools" and "how to do keyword research" share almost every word and belong in different sections, because one is somebody acquiring a skill and the other is somebody about to spend money. Group by job. Topic overlap is not job overlap, and a taxonomy built on topic overlap puts purchasing decisions inside curricula and wonders why they do not convert.
A taxonomy is a rule, and if you cannot write it you do not have one
If you remember one thing A taxonomy is the rule you use to classify a new page. If you cannot write the rule, you do not have a taxonomy, you have a habit, and the bill arrives as a redirect file.
Not in this path.
A taxonomy is the rule you use to decide where a new page goes. If you cannot write that rule down in a sentence, you do not have a taxonomy, you have a habit, and two people will apply it differently within a month.
Test it on your own site right now. Take two of your categories and write the rule that decides which one a new page belongs to. If the honest rule is "whichever feels more relevant", those are one category with two names, and the version of that failure I lived through is in the next chapter with a price on it.
The other failure is mixing classification axes inside one hierarchy. Topic, format, audience, industry, location and product are all legitimate ways to classify pages, and stacking three of them as though they were levels of one tree produces a structure nobody can navigate or extend. Under a "Blog" folder, "SEO" is a topic, "Marketing" is a discipline and "Guides" is a format. Three different kinds of thing pretending to be a hierarchy. One axis per level.
Here is the same twelve pages filed badly and filed well. Look at the left column first and try to name what is wrong before you read the list. There are five distinct problems and none of them is a spelling mistake.
Twelve pages, filed badly, then filed well
Same twelve pages in both columns. Look at the left one first and try to name what is wrong before you read the list underneath. There are five distinct problems in it and none of them is a spelling mistake, a slow page or a missing meta description.
Before Four years of adding a folder
-
Home -
Blog -
SEO -
Marketing -
Guides -
Keyword Research -
Resources -
SEO -
Keyword Research Guide -
Uncategorized -
Keyword Tips -
Best Keyword Tools
Every page here exists, is indexed, and has a readable URL. The problem is that the structure makes four contradictory claims about where keyword research lives.
After Grouped by job, one destination per need
-
Home -
Learn SEO -
Keyword Research -
Search Intent -
Site Architecture -
SEO Tools -
Keyword Research Tools -
Rank Trackers -
Resources -
Keyword Research Template -
Services -
About
Nothing deeper than two. One destination per need. Learning and buying separated because they are different jobs, and the commercial page has a parent.
The five problems, and what each one actually costs
-
01 Four destinations for one need 4 rows
Keyword Research at depth 5, Keyword Research Guide at depth 3, Keyword Tips under Uncategorized, and a tools list at the root. Four URLs, one job.
Fix Pick the one with links and history, merge the rest into it, redirect. Then decide whether the tools list is genuinely a different job, which it is.
-
02 Depth for no reason 1 row
Blog, SEO, Marketing, Guides, Keyword Research. Five levels, and three of them are synonyms for each other.
Fix Collapse the synonym levels. Depth should record a real narrowing, not the fact that nobody wanted to delete a folder.
-
03 Categories that classify nothing 3 rows
Marketing, Guides and Uncategorized. Write the rule that decides whether a new page goes in Marketing or Guides. You cannot, which means they are not categories.
Fix Delete them. If a taxonomy level has no rule, every new page is a coin toss and two people will toss it differently.
-
04 Two classification axes in one level 1 row
Under Blog, "SEO" is a topic. Under it, "Marketing" is a discipline and "Guides" is a format. Three different kinds of thing stacked as if they were a hierarchy.
Fix Pick one axis per level. Topic, then subtopic. Format is a page type, not a folder.
-
05 The commercial page at the root, unlinked 1 row
Best Keyword Tools sits directly off the home page and nothing in the blog tree points at it. It is the most commercially valuable page on this diagram and the structure says nothing about it.
Fix Give it a parent it belongs to and at least one contextual link from a page that already gets attention.
What the rebuild does not do: it does not move any URL that already works. The right-hand column is a
statement about parents and links, and it is entirely compatible with keyword research staying at
/keyword-research/ rather than moving to /learn-seo/keyword-research/. That
distinction is chapter twenty, and it is the difference between a restructure and a migration.
Three of those five defects were on this domain in the week I wrote this page, which is why the left column is not a strawman. And notice what the rebuild does not do: it does not move a single working URL. It changes parents and links. That distinction is chapter twenty and it is the difference between a restructure and a migration.
Then the working example, which is this site with the arguable groupings argued rather than asserted. Any diagram looks obvious once drawn; the four questions underneath it are the part worth reading.
The working example
This site, and the four groupings that could have gone either way
Seven top-level areas. Every node below is a live link, so you can check the claim rather than trust the diagram. The diagram is the easy half: the four questions underneath are the ones somebody could argue with, and each one has a rule under it you can apply to your own site.
-
Learn SEO5 lessonsA sequence, not a category. Each lesson assumes the one above it and says so.
-
SEO Tools103 listingsA directory. Its job is helping somebody choose, which is a different job from learning.
- Categories by job
- Free tools utility pages, not reading
- Depth pages under a listing 19 spokes
-
Separate from the directory because a deal expires and a listing does not.
-
Resources14 filesTemplates and checklists. Somebody wants a file, not an argument.
-
SEO Case Studies6 studiesPromoted to the top level because it is the page a prospect checks before reading anything else.
-
Blog2 categoriesTwo categories, down from five. The other three are in the redirect file.
-
Who is behind it and what you can buy. Kept together because a buyer checks both in one visit.
Four decisions, and the rule under each one
-
01 Why does "Keyword Research Tools" sit under SEO Tools rather than inside the learning sequence?
Because the user job is different. Somebody in the learning sequence is acquiring a skill in order. Somebody searching for keyword research tools has already decided to do keyword research and wants to know what to buy. Filing the second one inside a curriculum means a person with money in their hand has to leave the section to spend it.
Group by job, not by topic. Two pages can share every word in their titles and belong in different sections.
-
02 Why is the learning sequence a hub with five children rather than five root-level pages?
Because the sequence is the product. Every lesson is useful alone and considerably more useful in order, and the only thing that can state the order is a parent. Without it, a reader arriving on step four has no way to know there are three steps above it, which was exactly the situation before the roadmap page existed.
A parent earns its place when arriving there helps somebody decide what to do next.
-
03 Why do the lessons live at the root rather than under /learn-seo/?
Because /seo-basics/ and /seo-keyword-research/ already have history at those addresses, Google says it works out hierarchy from links rather than from URL folders, and a migration would spend real risk to make a folder tree match a diagram. The parent relationship is stated by the roadmap page, the breadcrumbs and the nav panel. The address is not carrying it and does not need to.
Conceptual hierarchy and URL hierarchy are separate. Only one of them is expensive to change.
-
04 Why did case studies get promoted to the top level when there are only six of them?
Because volume is the wrong axis for a nav slot. It is the page a prospect opens before deciding whether to read anything else, and buried two levels down under About it was reachable only on hover. Six pages that decide a sale beat sixty that do not.
Nav slots are allocated by what the page does for the business, not by how many pages sit behind it.
What this diagram does not show, and should: the link graph. Seven tidy areas, every node reachable, and 97 of the 103 pages inside the second one have no editorial link pointing at them. A structure diagram is a statement of intent. The link export is the statement of fact, and on this site the two still disagree.
Who wrote this, and how to check me on it
Placed here rather than at the end, because you have just read two chapters built entirely on measurements of my own site and you are entitled to know who took them and what they cost me to publish.
Not in this path.
Who is teaching this, and how to check me on it
The worst number on this domain is on this page, dated
This lesson argues that a site structure is only real if you can measure it, and that most sites are claiming something their own link graph contradicts. It would be a poor lesson if I asked that of your site and not of mine. So here it is in Google's four categories, then the method behind every data set on the page, including the one that makes me look worst.
Senior Digital Marketing Manager, Brainstorm Force · SEO since 2010 · Coimbatore, Tamil Nadu
01
Experience Has this person actually made these decisions, and lived with the consequences?
Since 2010, across more than a hundred sites I have owned and run, two of which I lost entirely to Panda and Penguin. Every mistake in the architecture debt chapter is one I have personally shipped: duplicate categories, a best-of post competing with its own category page, an old pillar and a new lesson answering the same need, and a roadmap step that lasted four days. The redirect file on this domain has 47 rules in it and I wrote all of them.
The version with the failures in it02
Expertise Do they know the mechanism, or only the vocabulary?
Senior Digital Marketing Manager at Brainstorm Force, MSc Computer Software Engineering with Distinction from University of Greenwich, and a dissertation that was an automated SEO management system. Professional Member of BCS, The Chartered Institute for IT since 2012. Which is why the load-bearing claims on this page are quoted from Google's documentation with the sentence intact, why four of them are graded invented rather than argued with, and why this site enforces one query per page at build time instead of writing a policy about it.
Credentials, dated and checkable03
Authoritativeness Does anyone else say so, or only them?
30,000+ students taught across six courses and free programs, 617 published videos, and a 15,000 member lifetime-deal community. The part I can prove is on the case studies with the raw exports attached. The part I cannot prove and will not claim: that a tidy structure caused any of it, which is exactly the claim this page grades as inference when other people make it.
Six case studies, with the exports04
Trust What happens when the evidence is embarrassing?
It goes on the page. The central receipt here is my own link graph, and on the day this published 118 of 122 money pages had no link from any post, which is exactly what it was the day before. I could have waited a month and published a better number. The tool that produced it is in a public repository and it is the same one this page tells you to run against your own site.
What this site earns from, and howThree data sets on this page. Here is how to go and get a different answer
The link graph of this site
node tools/authority.mjs, run against the content collections on 21 August 2026 and again on 22 August 2026. 147 content pages, 228 hand-written internal links, and 97 of 103 tool listings with no editorial link pointing at them. Run any crawler against your own site and open the orphan report; the finding is the same shape and the numbers will be yours.
The auditTwenty keywords, four parent topics
One Keywords Explorer request, US database, 20 keywords. All 20 came back informational. 8 of them are ways of saying the same thing, carrying 8,550 US searches a month between them, and Ahrefs assigns them 4 parent topics rather than 8. Put the same eight into any tool with a parent topic column.
The collapseThe URL decision, published
Two SERP Overview requests. On the query with fifteen times the volume, 4 of 8 organic results belong to design and website-builder companies. On the query this page targets, 0 of 8 ranking pages were built for it. Search both yourself and count the publishers rather than reading the titles.
Why not the bigger queryThe experience part, as numbers you can go and check
74%
ZipWP organic click growth from a zero baseline
Brainstorm Force, Search Console · 2026
28
ZipWP keyword variations taken to #1 from nothing
Brainstorm Force, Search Console · 2026
302,037
Bing Copilot citations earned by owned properties
Bing AI Performance · 6-month windows
30,000+
Students taught across all platforms
Udemy plus direct and free courses · Aug 2026
500+
SaaS products personally bought and tested
Since 2019
Those are counts from properties I own, with the tool and the window named. They are evidence that I have done this at some scale. They are not evidence that a structural change caused any of them, and nobody presenting architecture work as the cause of a traffic number has isolated that variable either.
Hubs, without the mythology
If you remember one thing A hub earns its place by being somewhere useful to arrive, not by sitting above things. Three pages is a legitimate cluster and so is fifty.
Not in this path.
A hub earns its place by being somewhere useful to arrive, not by sitting above things. Three pages is a legitimate cluster and so is fifty, and any framework that arrives with the number already filled in has not looked at your demand.
The useful version of a hub is concrete. Take keyword research. A hub for it might route to a guide, a page on search volume, one on difficulty, one on long-tail queries, one on competitor research, one on clustering and one on tools. Seven children, and each one earns its URL because it has its own demand and its own answer. The hub earns its URL because somebody arriving on the head term needs to be shown the shape of the subject before they can pick.
The mythological version is a pillar page of four thousand words that mentions each subtopic in a paragraph and links to a full article about it, published because a framework said a topic needs a pillar. That page competes with its own children, gets linked to by nobody, and exists to satisfy a diagram.
My own site has the asymmetry that this argument predicts. One tool listing has eighteen depth pages under it. The other 101 have one between them. That looks like neglect and it is mostly correct: one tool has eighteen sub-queries with their own measurable demand and the rest do not, and forcing symmetry would produce a hundred thin pages that each need a parent, a route and eventually a redirect.
Topic clusters are an organizing model, not a ranking formula
If you remember one thing Topic clusters are an organizing model, not a ranking formula. The pages have to genuinely relate, and no number of semantically adjacent URLs adds up to authority.
Not in this path.
A topic cluster is a hub plus its spokes, treated as one unit. It is a good way to arrange pages. It is not a mechanism by which arranging pages produces authority.
The version that works is boring: the pages genuinely relate, each one answers something distinct, the hub routes to all of them and every spoke links back. Technical SEO with crawling, indexing, canonical tags, robots.txt and XML sitemaps under it is a real cluster, because a person working on any one of those five is plausibly about to need another.
The version that does not work is a hundred pages generated because a tool called them semantically related. Semantic adjacency is not the same as a reader needing both, and no number of adjacent URLs adds up to expertise. The claim that it does is graded inference on this page, which is chapter twenty-four.
The honest test for a cluster: could you write one sentence explaining why somebody reading spoke three might want spoke five? If yes, they are in the same cluster. If the only connection is that both contain the same noun, you have a keyword group rather than a cluster, and keyword groups do not deserve a parent page.
Silos leak, and that is the point
If you remember one thing Architecture creates organization, not walls. Search Engine Land warns plainly that tightly isolated silos hide useful content and can leave pages effectively orphaned.
Not in this path.
Architecture creates organization, not walls. A link between two sections that a reader genuinely needs is not a leak, and refusing it because the pages are in different folders is siloing turned into superstition.
Search Engine Land says this directly, and it is worth quoting because it comes from a guide that otherwise recommends silos: tightly isolating silos can backfire, because if related topics are not linked across categories, users miss out on helpful content and search engines may struggle to index orphaned pages. Two failures, one of them about humans and one about crawlers, from over-tightening the thing you were told to tighten.
The useful half of siloing is that a topic has a home, which means a reader knows where they are and a crawler can infer a grouping. The harmful half is treating the folder boundary as a rule about links. On this site, keyword research and search intent are separate lessons and they link to each other repeatedly, because you cannot read a results page without a cluster and you cannot brief a page without both.
The practical rule I use: a link is justified by a reader needing it, and by nothing else. Not by folder proximity, not by anchor text opportunity, and not by a related-posts algorithm. If you cannot write the sentence the link sits in, it is not a link, it is a widget.
What earns a permanent navigation slot
If you remember one thing The header is the most repeated statement on your domain. Everything in it is on every page, so putting fifty things there says nothing is important.
Not in this path.
Everything in your primary navigation appears on every page of your site, which makes it the most repeated statement your domain makes about what matters. Fifty links in a header says nothing is important.
The mechanism is the one from chapter one. Google infers relative importance partly from how many internal links point at a page, and a header link points at that page from every single URL you own. That is a lot, and the total is fixed: a twentieth item does not create a twentieth share of something new, it takes a twentieth share of what the other nineteen were getting. Which is why "put important things in the header" is not advice. Everybody thinks their thing is important. A budget is advice.
What earns a permanent slot
Four tiers of navigation, and a budget for each
Everything in your primary navigation appears on every page of the site, which makes its anchor text the most repeated text on your domain and its length a statement about what matters. Fifty links in a header says nothing is important. The budget column is the part most advice leaves out, because "put important things in the header" is not advice when everybody thinks their thing is important.
-
01 Primary navigation
The highest-level journeys, on every page of the site.
Is this one of the five or six things somebody comes here to do?
Budget
Five to eight items. Every one appears on every page, so the anchor text is the most repeated text on your domain.
The mistake
Fifty links, because every department asked for one. The header stops being a statement of priority and becomes a directory nobody reads.
-
02 Secondary navigation
The level below a primary item, usually a panel or a sub-menu.
Does this earn reach without needing its own top-level slot?
Budget
Four to eight per parent. Enough to show the shape of the section.
The mistake
Making the panel the only route to a page. A hover menu is not a link a reader can find twice, and it is a poor place for anything commercially important to live alone.
-
03 Footer
Utility destinations, plus the small number of pages worth a site-wide link that do not fit the header.
Would somebody go looking for this at the bottom of the page?
Budget
Generous, and still a list somebody chose. A footer of two hundred links is a site-wide link farm and reads as one.
The mistake
Treating it as a second taxonomy the header does not share, so the site describes itself two different ways depending on where you look.
-
04 Contextual links
Links inside the content, in a sentence, pointing at the next useful thing.
Having read this, what would somebody want next?
Budget
As many as the argument genuinely supports. This is the only tier where the anchor text can be specific to the page it sits on.
The mistake
Leaving it to a plugin. Automated related-links blocks are the reason so many sites have a perfect nav and a link graph nobody designed.
This site’s header, read live from the file both menus render from
-
AI SEO Expert -
SEO Tools -
SEO Deals -
SEO Resources -
SEO Roadmap6 in the panel -
SEO Blog3 in the panel -
SEO Case Studies4 in the panel -
About4 in the panel
8 items, 4 of them with a second level. Every one is a real link rather than a hover-only parent, so the panel adds reach without ever becoming the only route to a hub. That was a deliberate change: an earlier version buried the whole learning sequence inside a Resources panel, where somebody trying to learn SEO in order would never have found it.
The tier most sites neglect is the fourth one, contextual links inside the content, and it is the only tier where the anchor text can be specific to the page it sits on. Leaving that tier to a plugin is how a site ends up with a considered header and a link graph nobody designed, which is precisely the state mine was in when I measured it.
The mobile menu that hides half your links
If you remember one thing Google indexes the mobile rendering. Two hand-maintained menus always drift, and the links only in the desktop one are links Google does not have.
Not in this path.
Google indexes the mobile rendering. If your desktop menu links to twenty category pages and your mobile menu links to ten, ten of those links effectively do not exist.
The mechanism is unglamorous and total. Mobile-first indexing means the mobile version is the version Google crawls and evaluates. A link present only in a desktop-width menu is not a link Google has, so whatever crawl route and importance signal it was carrying is gone, and nothing in any report tells you, because both menus render fine to a human checking them one at a time.
The cause is nearly always the same, and it is an engineering cause rather than an SEO one:
two hand-maintained lists. Somebody built a desktop nav component and a mobile nav
component, each with its own array of links, and the two drifted. They always drift. This
site had exactly that bug. There was a NAV_MOBILE array separate from
NAV, it had already fallen out of step with the desktop panels, and the fix was
to delete it and render both menus from one list. Now a link cannot exist in one menu and
not the other, because there is only one menu described in one place.
If you take one thing from this chapter, take the check rather than the principle. Open your site’s source, find both menus, and count the anchors in each. If the numbers differ, you have found something worth more than the rest of this page.
A link Google cannot follow is not a link
If you remember one thing Google states it can only crawl a link that is an anchor element with an href. A navigation built on click handlers is a navigation only humans can use.
Not in this path.
Google states it can only crawl a link that is an anchor element with an href attribute. Everything else in your navigation is a route for humans only.
Its own documentation gives the bad examples, and they are the exact patterns a modern front end produces by default: an anchor carrying a framework routing attribute instead of an href, and an element with an onclick handler that navigates. Both work perfectly in a browser. Neither puts a URL into a crawler’s queue, which means the pages behind them enter the index only if something else happens to link to them.
crawlable <a href="/technical-seo/">Technical SEO</a>
not <a routerLink="/technical-seo/">Technical SEO</a>
not <a onclick="goto('/technical-seo/')">Technical SEO</a>
not <span data-href="/technical-seo/">Technical SEO</span>
not a "load more" button that fetches the next twenty products That last one is the expensive case, because it is where the products live. A category page that shows twenty items and loads the rest on click has, as far as a crawler is concerned, twenty products in it. The fix is not clever: paginate with real URLs, and let the button be an enhancement on top of links that already exist.
I am deliberately stopping here rather than going into JavaScript rendering, because that is technical SEO and it is a different lesson. The architectural point is narrow and worth having on its own: your structure is made of anchors in the HTML, so read the HTML.
URL architecture, which is smaller than you think
If you remember one thing Readable words, hyphens, no pointless parameters, one trailing-slash shape, and then stop. The documented benefit of a folder tree is crawl scheduling, and it is smaller than the effort people spend on it.
Not in this path.
Readable words, hyphens, no pointless parameters, one trailing-slash shape, and then stop. This is the shortest chapter on the page on purpose, because URLs are the layer people fix first and the one carrying the least structural weight.
What Google actually recommends is modest and worth following: readable words rather than long ID numbers, hyphens rather than underscores, and trimming parameters that do not change the content. It also warns that additive filtering, session identifiers and dynamic calendars create unnecessarily high numbers of URLs and can generate infinite spaces, which is a crawling problem rather than a naming one and belongs to chapter twenty-three.
The address layer
Seven URL rules, and which ones Google actually publishes
5 of the seven come out of Google’s URL structure documentation with the recommendation intact. 2 are habits worth keeping on operational grounds that I am not going to dress up as ranking advice. This is the shortest chapter on the page for a reason: the documented benefit of a folder tree is that it helps Google learn how often a directory changes, which is real, useful and much smaller than the effort people spend on it.
- 01
Readable words, not IDs
Google says so/seo-tools/semrush/pricing//page?id=235897&type=4Google recommends readable words rather than long ID numbers. A URL is shown in results and pasted into messages, so it is read by people more often than most of the page is.
- 02
Hyphens, not underscores
Google says so/keyword-research-template//keyword_research_template/Google recommends hyphens to separate words. This is genuinely small and it is also free.
- 03
No unnecessary parameters
Google says so/seo-tools/category/rank-tracking//tools/?category=seo&sort=name&sess=8f21aGoogle says to trim parameters that do not change the content, and warns that additive filters, session IDs and tracking parameters create unnecessarily high numbers of URLs.
- 04
One trailing-slash shape, enforced
Sensible, not documented/about/ everywhere, in links, sitemap and canonicalNav links to /about, canonical says /about/Both work. Mixing them means every click from your own navigation goes through a redirect. This site pins it in astro.config.mjs with `trailingSlash: "always"` so the build, the sitemap and the internal links cannot disagree.
- 05
Do not repeat the folder in the slug
Sensible, not documented/shirts/blue/sparkles//shirts/blue-shirts/blue-shirts-with-sparkles/It reads as manipulation and it makes the URL longer for no gain. The mechanism people quote for this is invented; the habit is still right.
- 06
Stop changing URLs that work
Google says soA good URL, left alone for four yearsA migration to make the folder tree match the navGoogle has a video specifically on whether URL structure changes affect SEO, and every migration spends real risk. Prettiness is not a reason. A URL that misrepresents the page is.
- 07
Map redirects one by one
Google says so/seo-guides/semrush-tutorial/ to /seo-tools/semrush/tutorial/Every retired URL to the home pageA redirect to an unrelated page is treated as a soft 404, so the equity you were trying to preserve is discarded anyway. This site has 47 rules and every one names a specific destination.
The last two rules in that table are worth more than the first five and neither is about the words in the URL. Stop changing URLs that work, and map retired URLs one at a time. A bulk redirect of dead pages to your home page is treated as a soft 404, so the equity you were preserving is discarded and you have added a rule to your redirect file that does nothing except make the file harder to read.
One habit worth the five minutes: pick a trailing-slash shape and enforce it in the build
rather than in a review. This site pins trailingSlash: 'always' in its Astro
config, so the sitemap, the canonical tags and every internal link agree by construction. The
alternative is a navigation whose every click goes through a redirect, which nobody notices
and everybody pays for.
A page can sit at the root and still be a child
If you remember one thing A page can sit at the root and still be a child. Navigation, breadcrumbs and links state the parent relationship; the address does not have to, and moving it costs real risk for tidiness.
Not in this path.
Conceptual hierarchy and URL hierarchy are separate things, and only one of them is expensive to change. This is the chapter that saves people from unnecessary migrations, and it follows directly from Google reading links rather than folders.
The situation is common enough to be predictable. You draw a curriculum, and the folder tree
does not match it. Keyword research lives at /keyword-research/ and your diagram
says it should be at /learn-seo/keyword-research/. The diagram is right about
the relationship. It does not follow that the address has to carry it.
Navigation states the parent. Breadcrumbs state the parent. A contextual link from the hub states the parent. All three are things Google reads, and all three cost nothing. The URL is the one mechanism that costs a migration, and it is the one that Google says it does not use to work out your hierarchy.
This site does exactly that, and I would rather show you than tell you. The five roadmap
lessons all live at the root: /seo-basics/,
/seo-keyword-research/, /search-intent/, and this page. Their
parent is /seo-roadmap/ and every one of them is a child of it in the
navigation, in the breadcrumbs and in the links between them. Nothing about the addresses
says so and nothing needs to.
When is a migration justified? When the URL actively misrepresents the page, when you are consolidating genuine duplicates, or when a folder has to exist for a technical reason like a different template or a different language. "So the folders match the nav" is not on that list. Google has a video specifically on whether URL structure changes affect SEO, and the honest summary is that every migration spends real risk to buy something, so the something had better not be tidiness.
Subfolders, subdomains, and the claim nobody has isolated
If you remember one thing Use a subfolder unless an engineering constraint forces a subdomain. Google says it handles both; practitioners keep reporting otherwise, and nobody has isolated the variable.
Not in this path.
Use a subfolder unless an engineering constraint forces a subdomain. I believe that, I would do it every time, and I am following practitioner consensus rather than evidence, which I would rather say than dress up.
Here are both sides properly. Google has said repeatedly, including in a video from its own channel on exactly this question, that it handles subdomains and subdirectories and that the choice should be made on operational grounds. Practitioners keep reporting real gains from moving a blog off a subdomain onto a path. Both of those things can be true, and the reason nobody has settled it is that almost every one of those migrations rebuilt the internal linking at the same time. Nobody has published a version with the linking held constant, so the variable is not isolated and everybody quoting either side as settled is overstating it.
The reasons I would still use a subfolder are ones I can defend without a ranking claim. Analytics is one property rather than two, so you can see a journey from a guide to a pricing page without stitching sessions. Internal linking between the blog and the commercial pages is unremarkable rather than cross-domain. And nobody has to have the conversation again in two years.
A subdomain is genuinely right for a separate application, where you do not want a marketing CMS and a product sharing a deployment, and for large user-generated platforms that behave like their own site. Those are engineering answers, which is what Google said in the first place.
The related habit worth more than either: stop putting every page at the root. Not because
folders concentrate anything, but because a directory is a unit you can measure, audit and
prune. Filter analytics to /seo-tools/ and you learn something about a section.
A hundred pages at the root are a hundred pages.
Breadcrumbs, and what they are not
If you remember one thing Breadcrumbs show where you are and let you go up. Ahrefs says plainly they are not a replacement for main navigation, and a trail that disagrees with the menu is worse than no trail.
Not in this path.
A breadcrumb tells somebody where they are and lets them move up. Ahrefs says the rest plainly: it is not a replacement for your main navigation.
What they are genuinely good at is the case where somebody arrives from a search result three levels deep with no idea what the site around them is. A trail reading Home, SEO Tools, Semrush, Pricing does two useful things at once: it orients the reader, and it states a parent relationship in crawlable anchors on a page where the URL might not.
Home › SEO Tools › Semrush › Pricing
Two failure modes worth naming. The first is a trail that disagrees with the navigation, which is worse than no trail because the site is now making two contradictory claims about where a page belongs. The second is treating breadcrumb markup as the architecture: the markup describes a hierarchy, it does not create one, and a beautiful trail on an orphan page is a page describing a parent that never links to it.
A page can legitimately have two parents, and Google has a short video on whether you can place multiple breadcrumbs on a page for exactly that case. It is a real architectural situation rather than a mistake, and it usually means you should pick a primary parent for the trail and express the second relationship as a contextual link.
Faceted navigation and pagination, as architecture rather than implementation
If you remember one thing Filters multiply. Google warns that additive filtering and stray parameters create unnecessarily high numbers of URLs and can generate infinite spaces. The architecture makes the problem; technical SEO contains it.
Not in this path.
Filters multiply. Three filters with five options each is a hundred and twenty-five URL combinations before anybody has sorted anything, and the architecture created that problem before technical SEO had a chance to contain it.
Google warns about this in its URL documentation in almost those words: additive filtering, irrelevant parameters such as session identifiers, and dynamic calendars can create unnecessarily high numbers of URLs, and in some cases an infinite space. Which is the architectural point, and it is the only part that belongs on this page. You are choosing, when you design a filter set, how large a URL space your site will generate.
/tools/?category=seo&price=free&platform=windows&sort=name
↑ ↑ ↑ ↑
real slice real slice real slice changes nothing The decision to make here is which filter combinations correspond to a real need somebody searches for. "Free SEO tools" is a real need and deserves a real, linked, indexable URL. "Free SEO tools for Windows sorted by name descending" is not, and generating it as a crawlable link is volunteering your crawl budget for nothing.
Pagination is the same shape of problem with a different cause. Large categories need it, the pages need real URLs, and the architecture has to make sure the items stay discoverable without generating enormous depth. Chapter twenty-four has the example of what happens when pagination becomes the navigation.
Everything past this point is technical SEO: parameter handling, robots rules, canonical strategy, the PRG pattern. I am deliberately not covering it, and I would be suspicious of an architecture lesson that did, because the implementation follows the decision rather than replacing it. Google’s crawl budget guide is the reference, and read its opening line before you lean on it: it says it applies to sites with over a million pages, or over ten thousand pages changing daily. Almost nobody reading this is in either bracket, and citing it anyway is one of the ways this subject gets oversold.
Click depth, and the three-click religion
If you remember one thing Three clicks is a good target and a bad law. What Google documents is that required clicks and inbound link counts help it infer relative importance, which is a gradient rather than a cliff.
Not in this path.
Important pages should generally be easy to reach. That is the whole documented claim, and the version you have heard is a threshold that no Google document contains.
What Google actually says is about two quantities, not one: the number of links needed to reach a page, and the number of links pointing at it, both contribute to inferring that page’s relative importance within your site. A gradient, not a cliff, and it names inbound links alongside depth. Semrush recommends three clicks or fewer as a target, which is sensible advice and a different sentence from "anything deeper will not rank".
Four real paths from this site, and the fourth is why the threshold version is a distraction.
Depth, measured properly
Four real paths on this site, and the one depth cannot see
Google’s documented statement is that the number of links needed to reach a page, and the number of links pointing at it, help it infer that page’s relative importance within your site. Two quantities, and depth is only one of them. Read the fourth row: two clicks, immaculate depth, and nothing on the site has ever argued for it.
- depth 3 fine
-
Home -
/seo-tools/ -
/seo-tools/semrush/ -
/seo-tools/semrush/pricing/
The shape this site is built for. Three clicks, and every step is a page somebody would want anyway.
-
- depth 2 fine
-
Home -
/seo-roadmap/ -
/seo-site-architecture/
This page. Two clicks, and it is also linked from step four, from the footer and from the roadmap card, so the depth number understates how reachable it is.
-
- depth 5 buried
-
Home -
/blog/ -
/seo/ -
page 3 -
page 7 -
a post
Pagination as navigation. The post is reachable and the site has said nothing about it except that it is old.
-
- depth 2 looks fine
-
Home -
nothing -
/seo-tools/accuranker/
Two clicks through the directory grid, and zero editorial links. Depth looks fine. This page is one of the 97 nothing on the site argues for, which is a different problem that depth cannot see.
-
What Google documents
"The more links a page has to it within a site, the higher the relative importance of the page to other pages on your site." A gradient, about ordering pages within one site, and it names inbound links first.
What gets said instead
"Anything more than three clicks deep will not rank." No source, a hard threshold, and it sends people to flatten a structure when the actual problem is that nothing links to the page at any depth. Semrush recommends three or fewer as a target, which is a different sentence.
The practical version, which is boring and correct: important pages should be easy to reach and genuinely linked. If you have to choose which to fix first, fix the links. A page at depth four with forty inbound internal links is being treated as more important than a page at depth two with none, and I have both.
A page at depth two that nothing links to has perfect depth and no architecture. I have 97 of those. Depth measures the shortest route and says nothing about how many routes exist, which means the metric everybody optimizes is the weaker half of a claim nobody reads carefully.
Since this is the chapter where the evidence gets thin, here is every load-bearing claim on this page with a grade on it. Seven are documented and quoted. Four have no source at all, and those four are the ones you will hear stated most confidently.
How to argue about this subject
Twelve claims, graded by what is actually behind them
Documented means Google published the sentence and it is quoted here rather than paraphrased. On the record means it has been acknowledged with no current published model. Inference means reasonable and unmeasured. Invented means no source and no mechanism, and four of the twelve are in that bucket. Notice which four, because they are the ones repeated most confidently.
7
Documented
Google published it. The sentence is quoted, not paraphrased.
1
On the record
Acknowledged publicly, with no current published model. Direction known, magnitude not.
2
Inference
Reasonable, unmeasured, and nobody outside the company can isolate it.
2
Invented
No source and no mechanism. Usually stated with the most confidence.
-
Google finds most new pages by following links from pages it already crawled.
Stated in the SEO starter guide: the vast majority of new pages Google finds every day are through links. This is the whole mechanical case for architecture and it is the only part that needs no argument.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Google works out a site’s hierarchy from links between pages rather than from URL structure.
Verbatim, in the ecommerce site structure guide: "Google generally doesn’t look at the structure of URLs to work out the structure of a site. Instead, it analyzes the linkages between pages." This is the sentence the folder-tree tutorial is arguing with.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
The more internal links a page has, the higher its relative importance on that site.
Also verbatim from the ecommerce guide. Note the word relative. It is a statement about ordering pages within one site, not about competing with anybody else’s site.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Every page you care about should have a link from at least one other page on your site.
From the crawlable links guide, and it is the closest thing this subject has to a rule. It also quietly kills the idea that a sitemap entry is a substitute for a link.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
A link is only crawlable if it is an <a> element with an href attribute.
Verbatim. Their own bad examples include a routerLink attribute and an onclick handler. If your navigation is built that way, your architecture is a drawing.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Grouping similar topics in directories helps Google learn how often those URLs change.
This is the real, modest, documented benefit of a folder tree, and it is about crawl scheduling rather than relevance. Worth having. Not worth a migration.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Search engines will probably understand your pages as they are, regardless of how the site is organized.
Google’s own hedge, in the starter guide, immediately after the organization advice. Almost no architecture tutorial quotes it, and it is the most important sentence in this chapter for anybody about to spend six weeks on a restructure.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Internal links pass PageRank, and anchor text describes the target.
The anchor-text half is documented. The PageRank half is a mechanism Google has acknowledged for decades without publishing a current model, and every number anybody puts on it is invented. Treat direction as known and magnitude as unknown.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
A page more than three clicks from the home page will not rank.
Semrush recommends three clicks or fewer as a target, which is reasonable advice. The threshold version is not in any Google document. What Google says is that link counts and required clicks contribute to inferring relative importance, which is a gradient, not a cliff.
Website Architecture: Best Practices for SEO Site Structures, by Dana Nicole Semrush, Published 30 Jan 2025, read 22 Aug 2026
-
The URL slug accounts for about half of a page’s relevance.
There is no source for this and no plausible mechanism. Google’s URL guidance is about human readability and crawl efficiency, and its own starter guide says search engines will likely understand your pages as they are. Delete the number, keep the habit of writing readable slugs.
-
Subfolders concentrate authority that subdomains would strand.
Google has said repeatedly that it handles both and that the choice should be made on operational grounds. Practitioners keep reporting gains from moving a blog off a subdomain, and almost every one of those moves changed the internal linking at the same time. Both things can be true and nobody has isolated the variable.
Should I structure my site using subdomains or subdirectories? Google Search Central, Checked 22 Aug 2026
-
A clean hierarchy makes an LLM more likely to cite you.
Search Engine Land now frames hierarchy and consistent signals as making topical expertise easier for people, algorithms and LLMs to interpret, which is a reasonable claim about clarity. It is not a measured retrieval effect and nobody outside a lab can isolate it. Build the clear structure because it is clear.
Site Architecture for SEO: Structure That Ranks and Scales, by Kody Wirth and Sophie Atkinson Search Engine Land, Updated 27 Nov 2025, read 22 Aug 2026
Your link graph is your site saying what matters
If you remember one thing Your link graph is your site stating what matters. On most sites that statement contradicts what the owner would say out loud, and the contradiction is measurable in an afternoon.
Not in this path.
Every internal link is your site asserting that a page is worth reading. Add them up and you have a ranked list of what your site thinks is important, and on most sites that list contradicts what the owner would say out loud.
The useful thing is that the contradiction is measurable in an afternoon. List your five most linked-to pages. Separately, list the five pages your business most needs to work. Compare. Whatever gap you find is not an opinion, it is what your site is currently claiming, and Google is reading the claim rather than your intentions.
Mine, from the audit in chapter one. The most linked-to page on this domain is a free technical SEO checklist with 28 inbound editorial links. Second is a Semrush pricing page with 22. Both of those are correct and I am glad about them. Then nothing commercial appears again for a long way down, and 118 of 122 money pages have no link from any post at all. Two out of five agreeing is not a disaster and it is not an architecture either.
The comparison that makes this vivid is a page’s route rather than its depth:
Home → SEO Tools → Semrush → Pricing 22 links in Home → Blog → Archive → 2019 → page 14 → a post 0
Both pages are reachable. Only one of them is being argued for. The second is a page the site has said nothing about except that it is old, and no amount of flattening the folder tree changes that.
Orphans, dead ends, and the sitemap that hides both
If you remember one thing An orphan has no route in. A dead end has no route out. Google says every page you care about should have a link from at least one other page, which is also why a sitemap entry is not a fix.
Not in this path.
An orphan page has no route in. A dead end has no route out. Google says every page you care about should have a link from at least one other page on your site, which is also the sentence that kills the idea that a sitemap entry is a fix.
The mechanism for orphans is the one from chapter one. Discovery runs on links, so a page nothing links to enters the index only if you asked for it directly, and it accrues no importance signal from anywhere because there is nothing pointing at it to signal with. It is not penalized. It is unargued for, which on a large site is functionally the same thing.
A sitemap does help discovery and you should have one. What it does not do is create a route, pass anything, or tell Google that a URL matters relative to your other URLs. My own build has a small object lesson in this: it strips noindex URLs out of the sitemap automatically, because eighteen category pages were being advertised in there while simultaneously asking Google to ignore them. Two contradictory signals, and a crawler resolves that by trusting neither fully.
Dead ends are the mirror image and they get almost no attention, because they cost you the reader rather than the crawler. A page with inbound links, a good answer and no useful next step is a page where somebody finishes and leaves. Architecture is not only "can this be found". It is also "can this person continue", and the second question is the one that decides whether a well-ranked page earns anything.
The check is one export and almost nobody runs it. Every crawler has an orphan or "no inbound internal links" report. Screaming Frog is free to 500 URLs. If you have nothing but Search Console, the indexing report plus a sitemap count gets you most of the way.
Internal links do exactly three jobs
If you remember one thing Three jobs: discovery, relationship, importance. Any link advice that does not serve one of those three is decoration, and anchor text is the only part of it Google documents.
Not in this path.
Discovery, relationship and importance. Any advice about internal linking that does not serve one of those three is decoration, and only one of the three has a documented lever.
Discovery. A crawlable link is how a URL enters the system. This is the job with no ambiguity and no debate, and it is why the orphan report matters more than any other architecture metric.
Relationship. Anchor text and the sentence around it tell both a reader and Google what the target page is about. Google documents this directly: anchor text helps people and Google understand what the linked page is about. "Click here" throws that away for nothing, on both audiences, forever.
Importance. More internal links to a page means higher relative importance within your site. Google's words, and note the word relative: it is about ordering your own pages, not about competing with somebody else's domain.
What I am deliberately not going to do is put a number on the third job. PageRank flow inside a site is real in direction and unpublished in magnitude, and every internal-link equity model in circulation is an inference dressed as arithmetic. The honest version: linking from a page that already gets traffic to a page that needs it is a sound bet, and you should measure the target page rather than the model.
Two practical notes that come straight from Google rather than from folklore. Do not nofollow your internal links; there is a video on it and the answer is short. And varying anchor text is fine and mostly about readability, because the question of whether repeated internal anchors hurt has also been answered on that channel and the answer is less dramatic than the advice built on top of it.
The rest, internal PageRank analysis, link opportunity automation, anchor distribution modelling, belongs to a dedicated internal linking lesson. Here you need the three jobs and the habit of routing one link deliberately, with the date written next to it.
Design for the two hundredth page, not the twentieth
If you remember one thing Design for the two hundredth page, not the twentieth. Architecture gets exponentially more expensive to change after growth, which is exactly when people notice they need to.
Not in this path.
Architecture gets exponentially more expensive to change after growth, and growth is exactly when people notice they need to change it.
Search Engine Land makes the point about designing taxonomy and navigation for future scaling, and Backlinko makes the complementary observation from the other end: complicated architectures usually emerge from sites adding categories, subdomains and pages over time rather than from one bad initial plan. Nobody designs a mess. It accretes, one convenient decision at a time.
So ask the questions at twenty pages that you will be forced to ask at two hundred. What is the rule for adding a category? What happens when a category has ninety children? Which filter combinations get a real URL? Who decides, and where is that decision written down?
The specific case worth planning for is programmatic. If you intend to publish ten thousand tool pages, fifty categories, a comparison matrix and an alternatives page per product, every one of those decisions has to exist before the generator runs, because a generator applies your taxonomy ten thousand times whether or not the taxonomy was any good. Architecture precedes programmatic SEO, and the order is not negotiable in the way most of this page's advice is.
My own version of learning this: /seo-tools/{tool}/ was the right
call and I made it before there were a hundred listings, which is the only reason the
section is coherent today. The part I got wrong was not planning how a listing earns depth
pages, which is why one tool has eighteen and the rest have one, and why the routing
problem in chapter one exists at all.
A competitor’s structure is evidence, not instructions
If you remember one thing A competitor’s structure encodes their business model. On the bigger of the two queries behind this page, four of eight ranking pages belong to companies that sell website builders.
Not in this path.
Look at competitor architecture after keyword research, the way Ahrefs recommends, and read it as evidence about a market rather than as a template. A competitor’s structure encodes their business model, and copying it means inheriting a business model you do not have.
What to look for: their categories, which commercial page types they bothered to build, where the hierarchy is deep and where it is flat, what is in their navigation, and which patterns repeat across many URLs. That last one is the most informative, because a repeated pattern is a decision somebody made deliberately at scale.
And now the demonstration, because I made exactly this call about this page and the evidence is worth showing at full size. Two results pages. One is the query this page targets. The other has fifteen times the volume at lower difficulty, and I declined it.
The URL decision, on the record
Why this page is not at /website-structure/
Two SERP Overview requests, US database, 22 August 2026. The query on the right has fifteen times the volume and lower difficulty than the one on the left, and I am not targeting it. Read the publisher column on both, not the titles: the titles are nearly interchangeable and the audiences are not.
seo site architecture
| # | Who published it | URL | DR | Traffic | Its own top keyword |
|---|---|---|---|---|---|
| 2 | SEO software | semrush.com/blog/website-structure/ | 92 | 2,769 | seo site structure 700 |
| 3 | SEO publisher | backlinko.com/hub/seo/architecture | 90 | 637 | site architecture seo 500 |
| 4 | A subreddit | reddit.com/r/SEO/comments/1dv0j7j/website_architecture_for_seo/ | 95 | 272 | seo architecture 1,000 |
| 5 | An SEO agency | terakeet.com/blog/website-architecture/ | 72 | 295 | site architecture for seo 450 |
| 6 | A crawler vendor | lumar.io/learn/seo/site-architecture/ | 75 | 194 | seo architecture 1,000 |
| 7 | Trade press | searchengineland.com/guide/website-structure | 91 | 360 | seo structure 1,100 |
| 8 | SEO software | moz.com/learn/seo/website-architecture-internal-links-video | 91 | 884 | moz’s link explorer 500 |
| 9 | Website builder | wix.com/blog/what-is-website-architecture | 95 | 65 | site architecture seo 500 |
AI Overview cited 6
semrush.com/blog/website-structure/searchengineland.com/guide/website-structuremagiclogix.com/theories/seo-site-architecture/wix.com/blog/what-is-website-architectureseosherpa.com/website-architecture/rankmath.com/seo-glossary/site-architecture/
Not one of the eight ranking pages was built for this query. Every single one has a different phrasing as its own top keyword, and one of the eight is a page whose top keyword is Moz’s own product name. Eight pages, eight different primary targets, and all of them collecting the same need. This results page is the argument of chapter five, made by the market rather than by me.
website structure
| # | Who published it | URL | DR | Traffic | Its own top keyword |
|---|---|---|---|---|---|
| 2 | Design software | figma.com/resource-library/website-structure/ | 92 | 4,185 | website structure 3,900 |
| 3 | Sitemap software | slickplan.com/blog/types-of-website-structure | 74 | 1,771 | website structure 3,900 |
| 5 | SEO software | semrush.com/blog/website-structure/ | 92 | 2,769 | seo site structure 700 |
| 6 | Website builder | squarespace.com/blog/understanding-website-structure | 95 | 740 | website structure 3,900 |
| 7 | A university IT help desk | help.ithaca.edu/TDClient/36/Portal/KB/Article/965/Website-Structure | 75 | 179 | website structure 3,900 |
| 8 | Website builder | webflow.com/blog/website-structure | 92 | 741 | website structure 3,900 |
| 9 | An independent creator | youtube.com/watch?v=db4CoweZIJE | 99 | 240 | website structure 3,900 |
| 10 | SEO software | yoast.com/site-structure-the-ultimate-guide/ | 91 | 923 | site structure 1,400 |
AI Overview cited 3
figma.com/resource-library/website-structure/developer.mozilla.org/en-US/docs/Learn_web_development/Core/Structuring_content/Structuring_documentsslickplan.com/blog/types-of-website-structure
The bigger term, and it is not an SEO query. Four of the eight organic results are design and website-builder companies, two are SEO, one is a university help desk and one is a video. The AI Overview cites Figma, MDN and a sitemap tool, and no SEO source at all. Whoever owns this query owns a planning market, not a search market.
The decision, and what it cost
What I gave up
3,900 US searches a month against 250, at difficulty 25 against 28. On a volume-first process this is not a close call, and a content calendar built from a keyword export would have picked the bigger one every time.
Why I gave it up
4 of the 8 organic results on the bigger query belong to design software, sitemap software and website builders. Only 2 are SEO publishers, at positions five and ten, and its AI Overview cites Figma, MDN and a sitemap tool with no SEO source at all. That is a planning market for people about to build a website. I am writing for somebody five steps into an SEO curriculum.
The finding I did not expect
Not one of the 8 pages ranking for the query I did pick was built for it. Every single one has a different phrasing as its own top keyword, 6 distinct ones between them, and position eight is a page whose top keyword is Moz’s own product name. Eight pages, eight different primary targets, all collecting the same need. That is chapter five happening on a live results page.
What I will be wrong about
If this page ends up ranking better for website structure than for the query it
targets, the decision was wrong and the results page told me so. That is a testable prediction
with a date on it, and it will appear in the page history either way.
Four of the eight results on the bigger query belong to design software, sitemap software and website builders, and its AI Overview cites Figma, MDN and a sitemap tool with no SEO source at all. The titles on the two results pages are nearly interchangeable. The publishers are not, and the publisher column is where the audience lives. Copying that architecture as an SEO consultant would mean copying somebody else’s product.
The finding I did not expect is the one that makes chapter five real. Not one of the eight pages ranking for "seo site architecture" was built for that phrasing. Every single one has a different phrasing as its own top keyword, 6 distinct ones between them, and position eight is a page whose top keyword is Moz’s own product name. Eight pages, eight different primary targets, all collecting the same need. The market has already made the one-page-per-need decision and none of those eight built the extra pages.
One structure for humans, crawlers and retrieval
If you remember one thing Do not build a second architecture for AI. Build one clear information system, and note that on this page’s own query only three of the six AI Overview citations were in the visible top ten.
Not in this path.
Do not build a second architecture for AI. Build one clear information system, and be honest that the case for it is clarity rather than a measured retrieval effect.
Here is what is actually true and useful. A well-organized site states relationships: product has feature, product integrates with X, product serves job Y, this person wrote that page, this case study proves that service. Those relationships are asserted by having both pages and a link between them, which is the same mechanism as the rest of this lesson. A buyer comparing two products needs exactly that information and most sites make them assemble it from a features page and a pricing table.
Relationships, stated by structure
Eight relationships a site asserts just by linking two pages
No markup involved. Each row is a claim your site makes about the world, and it is made by having both pages and a link between them. The third column is what turns this into work: an edge with no page behind it is not an assertion, and an edge whose two pages do not link to each other is an assertion nobody can follow.
-
Product has feature Feature
Feature page, linked from the product page and back
-
Product integrates with Third-party tool
Integration page, linked from both directions where you own both
-
Product serves Job to be done
Use case page, linked from the feature that does the job
-
Product costs Price
Pricing page, linked from every commercial page
-
Product competes with Named competitor
Comparison page, and an alternatives page for the other direction
-
Product is documented at Documentation
Docs, linked from features rather than only from the footer
-
Person wrote Page
An author page that actually exists and is linked from every piece
-
Company proved it at Client outcome
Case study, linked from the service it proves
What this buys, definitely
A buyer comparing two products needs exactly this information, and most sites make them assemble it from a features page and a pricing table. Building the relationships once serves the buyer, the crawler and anything summarizing you, which is cheaper than building three versions of it.
What it does not buy
A measured increase in citations. Search Engine Land frames a clear hierarchy as making topical expertise easier for people, algorithms and LLMs to interpret, which is a claim about clarity rather than a retrieval finding. On the query this page targets, three of the six AI Overview citations were not even in the visible top ten, so whatever is happening in there is not a simple function of your folder tree.
What I will not claim is that this makes you more likely to be cited. Search Engine Land now frames a clear hierarchy and consistent signals as making topical expertise easier for people, algorithms and LLMs to interpret, which is a reasonable claim about clarity and not a retrieval finding. This page grades it as inference and it stays inference until somebody isolates the variable, which nobody outside a lab can.
And there is one measurement on this page that argues against the simple version. On the query this page targets, the AI Overview cited 6 sources and only 3 of them were in the visible top ten. Whatever is happening inside that box is not a straightforward function of your folder tree, and anybody selling you an "AI-ready architecture" should be asked what they measured.
The practical conclusion is boring and I stand behind it. Build one information system that humans, crawlers and retrieval systems can all follow, because you would have wanted it for the humans anyway and building it twice is how a good idea becomes a budget line.
SEO architecture debt
If you remember one thing Architecture debt is what you owe for every page published without deciding where it belonged. It compounds quietly and it is paid in redirects, mine included.
Not in this path.
Architecture debt is what accrues when pages get published without anybody deciding where they belong. It compounds quietly, it is invisible in any report that counts pages rather than links, and it comes due as a redirect file.
Nobody designs a bad architecture. The mechanism is always the same: a page does not fit an existing category, so somebody makes a new one, and the new one has no rule attached. Do that eight times over four years and you have eight categories, no rules, and every new page is a coin toss.
Here is my bill, from the file in this repository, dated.
Architecture debt
Ten symptoms, and my own bill for eight of them
Nobody designs a bad architecture. It accrues, one page at a time, every time somebody publishes without deciding where the page belongs. It is invisible in any report that counts pages rather than links, and it comes due as a redirect file.
-
01
Duplicate categories
How you know Two category names you cannot write a rule to choose between.
Fix Merge, redirect the loser, and write down what the survivor means.
-
02
Several URLs for one need
How you know Search your own site for a job and get three pages back.
Fix Pick the one with links and history. Merge the rest into it and redirect.
-
03
Orphans
How you know A page in the sitemap that nothing on the site links to.
Fix One contextual link from a page that already gets traffic, or delete it.
-
04
Dead ends
How you know A page with inbound links and no useful outbound next step.
Fix Add the next step. Not a related-posts widget, a sentence.
-
05
Abandoned folders
How you know A directory holding two pages from three years ago.
Fix Move them, redirect the folder, and stop pretending it is a section.
-
06
Thin filter URLs
How you know Indexable tag or facet pages with one item on them.
Fix Stop generating them, or stop indexing them. Preferably both.
-
07
Nav that lies
How you know The header names areas that are not where the value is.
Fix Rebuild the header from what people actually come to do.
-
08
Breadcrumbs that disagree with the nav
How you know The trail says one parent, the menu implies another.
Fix Decide which is true and change the other.
-
09
Pages nobody can name the purpose of
How you know Ask what job a page finishes and get a topic instead of an answer.
Fix Merge it into the page that does have a job, or remove it.
-
10
A structure you cannot draw
How you know You cannot fit the site on one page without a legend.
Fix That is the finding. Start from the six steps rather than from the existing tree.
The bill, from this repository
public/_redirects, replaced 21 August 2026, amended 23 August 2026
Where the redirects came from
Where they point now
Seven folders stopped existing:
/seo-guides//seo-experiments//tools//resources/seo//resources/guides//resources/analytics-and-reporting//resources/campaigns-and-conversion/ Every one of them was created because a page did not fit anywhere, which is the whole
mechanism of architecture debt in one sentence.
The nine redirects out of /seo-guides/ are the most expensive line here. It was a
category that could not answer the question "what goes in this and not in the other one", so for
four years the answer was whatever the author felt like, and the correction was nine URLs moving
and a category being deleted. The fix was not a project. It was writing down, once, what each of
the two surviving categories means.
Nine of the 47 rules retire one category, /seo-guides/, which
existed for four years and could never answer the question "what goes in this one and not the
other one". The blog category enum went from 5 values to
2 in three days. That correction was not a project. It was
writing down, once, what each of the two survivors means.
The reassuring half: this is normal, and the fix is a habit rather than a rebuild. Before publishing anything, name its parent and one page that will link to it. That is the whole habit. Sites that do it accumulate structure. Sites that do not accumulate a redirect file, and I have both a redirect file and the habit now, in that order.
The mistakes I would kill first
If you remember one thing Twelve claims about this subject, four of them invented, and the invented ones are the ones repeated with the most confidence.
Not in this path.
Twelve claims I have been told about this subject, four of them with no source anywhere, and those four are the ones stated with the most confidence.
The kill list
Twelve things people believe about site architecture
Verdict first, then the reason, then where the reason comes from. 8 of the twelve cite a primary source and the rest are answered by measurements of this site. Six of the citations are Google’s own documentation, which on this subject says something noticeably less exciting than the tutorials built on top of it.
-
01
"A tidy folder structure is site architecture."
It is the least load-bearing layer
Google’s own ecommerce guidance says it generally does not look at URL structure to work out the structure of a site, and analyzes the links between pages instead. My own site is the demonstration: 103 tool pages in a perfectly consistent /seo-tools/{tool}/ tree, and on 22 August 2026 the link-graph tool found that 97 of them had no editorial link pointing at them. Immaculate folders, no architecture.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
02
"Anything more than three clicks from the home page will not rank."
A useful target sold as a law
Semrush recommends three or fewer, which is sensible. Google’s documented statement is different and softer: the number of links needed to reach a page, and the number of links pointing at it, help it infer the page’s relative importance within your site. That is a gradient. A page at depth four with forty internal links is treated as more important than a page at depth two with none, and I have both on this site.
Website Architecture: Best Practices for SEO Site Structures, by Dana Nicole Semrush, Published 30 Jan 2025, read 22 Aug 2026
-
03
"One keyword, one page."
The most expensive sentence in SEO
On 22 August 2026 I pulled twenty keywords in one Ahrefs request. Eight of them are ways of saying "how should a website be arranged", they carry 8,550 US searches a month between them, and Ahrefs assigns those eight four parent topics rather than eight. The results page agrees: not one of the eight pages ranking for "seo site architecture" was built for that phrasing, and all eight rank anyway.
Answered from my own data: the link graph of this site, run on 21 and 22 August 2026, plus the twenty keywords and two results pages pulled from Ahrefs on the second date. Every row of all of it is shown in full further up this page.
-
04
"Every topic needs a pillar page and twenty cluster articles."
A template, not a finding
Sometimes the right architecture for a subject is three pages. Sometimes it is fifty. The number comes from how many genuinely distinct destinations the demand supports, and any framework that arrives with the number already filled in has not looked at your demand. On this site one tool has eighteen depth pages under it and 101 others have none, and that asymmetry is correct: one tool has eighteen sub-topics with their own search demand and the rest do not.
Answered from my own data: the link graph of this site, run on 21 and 22 August 2026, plus the twenty keywords and two results pages pulled from Ahrefs on the second date. Every row of all of it is shown in full further up this page.
-
05
"Silos work because they keep authority inside a topic."
Over-tightened silos cost you discovery
Search Engine Land warns directly that tightly isolating silos can backfire, that users miss helpful content and search engines may struggle to index orphaned pages. The useful half of siloing is that a topic has a home. The harmful half is refusing a link between keyword research and search intent because they are in different folders, when a reader needs both.
Site Architecture for SEO: Structure That Ranks and Scales, by Kody Wirth and Sophie Atkinson Search Engine Land, Updated 27 Nov 2025, read 22 Aug 2026
-
06
"An XML sitemap is how Google finds your pages."
It is a hint, and it is not a route
Google says every page you care about should have a link from at least one other page on your site, and that the vast majority of new pages it finds each day come through links. A sitemap entry with no inbound link is a page you have asked to be indexed and given no reason to be. This site’s build even prunes noindex URLs out of its own sitemap, because eighteen category pages were being advertised there while asking not to be indexed.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
07
"The URL slug is about half of a page’s relevance."
Invented, with a suspiciously round number
There is no source and no mechanism. Google’s URL documentation is about readability and crawl efficiency, and its starter guide says search engines will likely understand your pages as they are, regardless of how the site is organized. Write readable slugs because a human reads them in the result. Do not budget half your relevance to them.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
08
"Move the blog off the subdomain and the authority flows."
Probably right, for reasons nobody has isolated
Google has said repeatedly that it handles subdomains and subdirectories and that the choice is an operational one. Practitioners keep reporting gains from the move. Almost every one of those migrations also rebuilt the internal linking, and nobody has published a version with the linking held constant. I would still use a subfolder, and I would say out loud that I am following consensus rather than evidence.
Should I structure my site using subdomains or subdirectories? Google Search Central, Checked 22 Aug 2026
-
09
"You need a separate architecture for AI search."
No, and the pitch usually reveals itself
The mechanism people describe for an AI-friendly architecture is the same mechanism as a human-friendly one: clear groupings, stated relationships, reachable pages. On the query this page targets, the AI Overview cited six sources on 14 July 2026 and only three of the six were in the visible top ten. That is worth knowing and it is not an argument for a second site.
Answered from my own data: the link graph of this site, run on 21 and 22 August 2026, plus the twenty keywords and two results pages pulled from Ahrefs on the second date. Every row of all of it is shown in full further up this page.
-
10
"Copy the architecture of the site that outranks you."
Evidence, not instructions
Ahrefs recommends looking at competitor structures after keyword research, and that is the right use for it. The trap is that a competitor’s structure encodes their business model. On "website structure", four of the eight ranking pages belong to design and website-builder companies whose architecture exists to sell a site builder. Copying it as an SEO consultant would be copying somebody else’s product.
How to Structure Your Website Architecture for SEO, by Chris Haines Ahrefs, Published 23 May 2023, read 22 Aug 2026
-
11
"Cannibalization is Google punishing you for duplicate keywords."
It is a decision you already made
Nothing is being punished. You built two destinations for one need and now the site cannot tell anybody which one it means, so Google picks, and it picks on its own evidence rather than on your intentions. The fix is architectural: merge, differentiate, or redirect. This site prevents it at build time instead, by making every depth page declare its one query and failing the build when two pages claim the same one.
Answered from my own data: the link graph of this site, run on 21 and 22 August 2026, plus the twenty keywords and two results pages pulled from Ahrefs on the second date. Every row of all of it is shown in full further up this page.
-
12
"Get the architecture right at launch and you are done."
It is a maintained system, like a codebase
Backlinko makes the practical observation that complicated architectures emerge from sites adding categories and pages over time rather than from a bad initial plan. The bill arrives as a redirect file. Mine was replaced on 21 August 2026 with 46 rules in it, nine of them retiring a single category that had never meant anything a reader could use. It was 47 two days later.
Website Architecture: How to Setup an SEO-Friendly Structure Backlinko, Read 22 Aug 2026
Test yourself
If you remember one thing Sixteen questions. Every answer is either quoted from a document or measured on my own site.
Not in this path.
Sixteen questions. Six of them quote a sentence Google published, four are answered by measurements of this site, and none of them is a definition you could have looked up.
Test yourself
Sixteen questions that separate a link graph from a folder tree
Every answer is sourced to Google’s own documentation, or to a measurement of this site you can reproduce with a free tool. Four of the sixteen have Google’s sentence quoted verbatim in the explanation. If an answer disagrees with something you were taught, check the source before you check me.
0
of 16
Answer the first question to start
-
Show the answer
Correct B. The links between your pages
Verbatim from Google’s ecommerce site structure documentation: it generally does not look at the structure of URLs to work out the structure of a site, and analyzes the linkages between pages instead. This one sentence reorders the whole subject, because it means a perfect folder tree with no links is not an architecture.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct B. Build one page and check how many parents the eight collapse into
Ahrefs computes parent topic by finding the keyword that sends the most traffic to the current number one page, which makes it a machine estimate of how many destinations the demand supports. On 22 August 2026 those eight phrasings collapsed to four parents, and the results page for one of them contained eight pages none of which was built for it.
-
Show the answer
Correct B. An orphan page
An orphan is a page with no internal path pointing at it. Google says every page you care about should have a link from at least one other page on your site, which is why sitemap presence does not fix it. A dead end is the opposite problem: links in, nowhere useful to go out.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct B. <a onclick="goto('/technical-seo/')">Technical SEO</a>
Google’s crawlable links documentation states that it can only crawl a link if it is an <a> element with an href attribute, and its own bad examples include an onclick handler and a framework routerLink attribute. If your main navigation is built that way, the structure exists for users and not for crawlers.
SEO Link Best Practices for Google Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct B. It contributes to inferring the page’s relative importance on your site
The documented claim is about relative importance within one site, alongside the number of links pointing at the page. It is a gradient rather than a threshold, and the three-click rule is a reasonable target that got promoted to a law somewhere along the way.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct B. Leave the URL and express the hierarchy through navigation and links
The folder tree is not how Google infers hierarchy, so the migration buys you tidiness and spends real risk. Navigation, breadcrumbs and contextual links can state the parent relationship perfectly well while the address stays where its history is.
Ecommerce Website Navigation Structure Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct B. One is a planning diagram of pages and relationships, the other is a machine-readable URL list
The confusion is worth killing early because it hides a real failure: teams produce the XML file, believe they have done architecture, and never draw the diagram that would have shown them the orphans. The deliverable at the end of this lesson is the diagram.
-
Show the answer
Correct B. It helps Google learn how often URLs in that directory change
This is the documented benefit, from Google’s SEO starter guide, and it is about crawl scheduling. A /policies directory that rarely changes and a /promotions directory that changes daily can be crawled at different frequencies. Worth having. Not worth a migration, and not a relevance mechanism.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct C. Redirect the weakest to the home page to preserve equity
A redirect to an unrelated page is treated as a soft 404, so the equity you were preserving is discarded and you now have a rule in your redirect file that does nothing. Map each retired URL to its most relevant destination one at a time.
-
Show the answer
Correct B. Google crawls mobile-first, so ten of those links may not be seen at all
Mobile-first indexing means the mobile rendering is the one that counts. Two hand-maintained menus always drift; this site had exactly that bug, a separate NAV_MOBILE array that had already fallen out of step, and the fix was to render both menus from one list.
-
Show the answer
Correct B. Help somebody finish a task by having its children collected there
That is the test, and it fails a surprising number of pillar pages. If the only argument for a parent is that a framework says a topic needs one, it is a filing cabinet with a URL, and it will get no links because nobody has a reason to send anyone there.
-
Show the answer
Correct B. A very large number of URLs from combining filters
Google’s URL documentation warns specifically that additive filtering, irrelevant parameters and dynamic calendars create unnecessarily high numbers of URLs and can produce infinite spaces. The implementation controls belong in technical SEO. Understanding that the architecture generates the problem belongs here.
URL Structure Best Practices for Google Search Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
-
Show the answer
Correct D. None
None. All twenty came back informational and exactly one also came back commercial. That is the same finding as the previous lesson from a different keyword set, and it is why the intent column sorts rows rather than deciding anything.
-
Show the answer
Correct B. Export your internal links and find the pages nothing points at
Depth is computed from the links you have, so it looks healthy on a page that is two clicks away through a grid and has no editorial link at all. I have 97 pages exactly like that. Listing the pages with no inbound link is the check that finds them.
-
Show the answer
Correct B. One category with two names
A taxonomy is the rule, not the labels. If the honest rule is "whichever feels more relevant", two people will apply it differently within a month and the correction arrives as a redirect. Mine did: nine of the 47 rules in this site’s redirect file retire one category that could never answer this question.
-
Show the answer
Correct B. That search engines will likely understand your pages as they are, regardless of how the site is organized
Almost no architecture tutorial quotes it, because it deflates the sales pitch. It does not mean structure is worthless. It means the gains are in discovery, prioritization and human navigation rather than in Google failing to comprehend your pages, and a restructure sold on comprehension is being sold on the wrong thing.
SEO Starter Guide: the Basics Google Search Central, Doc updated 10 Dec 2025, read 22 Aug 2026
The six steps, run on the worst section of my own site
Not a hypothetical and not a client. /seo-tools/ is 103 listings, 19 depth pages and 97 pages nothing on the site argues for. Eight steps, and the last one ends with the problem still open.
Not in this path.
The six steps, run on a real section
Rebuilding /seo-tools/, and it is not finished
103 tool listings, 19 depth pages, and 97 pages nothing on the site argues for. This is the six-step spine applied to that, in order, on the day this page published. Step eight ends with the problem still open, which is the honest state of it and the reason it is worth reading rather than a case study written backwards from a good number.
- 01 Setup
Write down what is actually there before deciding anything
What I did Ran the link-graph tool against the content collections rather than reading the sitemap. 22 August 2026.
What came back 147 content pages, 228 hand-written internal links, and a list of 106 pages with no editorial link pointing at them. 97 of the 106 are tool listings.
The transferable part The sitemap would have told me I had 147 pages and nothing else. The link export told me what the site is actually claiming, which is that four fifths of its commercial pages do not matter.
- 02 Demand
Ask which of the 103 listings has demand worth routing to
What I did Sorted the listings by whether the tool has its own search demand and its own affiliate or commercial value, rather than alphabetically.
What came back A short head and a very long tail. One tool already carries 18 depth pages; 101 carry none.
The transferable part The 97 unrouted listings are not one problem. Maybe fifteen of them deserve an editorial link this year and the rest are a directory, which is a legitimate thing for a page to be. Deciding that is step one, and skipping it turns a routing job into 97 obligations.
- 03 Intent
Check what each surviving listing is for
What I did Read the results pages for the head tools. Some want a review, some want a pricing page, some want an alternatives page, and one wants a free trial page.
What came back Four page types across fifteen tools, not one type across all 103.
The transferable part This is why step four comes before this one. A routing plan that treats every listing as the same kind of destination will link to the wrong page on the tools where the money is.
- 04 Page
Decide how many URLs each head tool actually needs
What I did Applied the parent topic test to the sub-queries under two tools.
What came back One tool supports eighteen distinct destinations. The next most demanded supports one. Nothing in between currently justifies more than two.
The transferable part The eighteen-to-one asymmetry looks like neglect and is mostly correct. A hub gets the number of spokes its demand supports, and forcing symmetry across 103 listings would produce a hundred thin pages.
- 05 Parent
Fix the parent relationships that were only implied
What I did Checked whether every listing and spoke states its parent, and whether every post states its category hub.
What came back On 21 August, 1 of 13 posts linked to its own category hub. On 22 August, 11 of 11 did.
The transferable part That fix took an afternoon and it is the cheapest thing in this entire example. It also did not move the money-page number by a single page, which is the honest limit of tidying relationships you already had.
- 06 Path
Give the surviving listings a route that is not the grid
What I did Picked the pages that already get traffic and asked which listing each one has a genuine reason to mention.
What came back A shortlist. It is short because most posts on this site have no honest reason to link to most tools, and manufacturing one is how a link graph becomes noise.
The transferable part This is the slow part and there is no version of it that a plugin does. A related-tools widget would have produced 103 links today and told a crawler nothing about which of them anybody meant.
- 07 Priority
Check whether the structure agrees with the business
What I did Compared the five most linked-to pages against the five pages the business most needs to work.
What came back The most linked-to page on the site is a free checklist with 28 inbound links. The second is a pricing page with 22. Both are correct. Nothing commercial appears again until much further down.
The transferable part Two of five agreeing is not a disaster and it is not an architecture either. The gap between what the link graph says matters and what I would say matters is the actual work item, and it is measurable, which is the only reason it will get done.
- 08 Decision
Write the verbs, including the ones that admit it is unfinished
What I did Put one verb on each row and dated the list.
What came back Keep the directory as a directory. Improve routing on roughly fifteen head listings. Ignore the tail. Merge one best-of post into its category. Redirect the old keyword research pillar. Remove one lesson page entirely.
The transferable part On the day this published, 118 of 122 money pages still had no link from a post, which is exactly what it was the day before. I could have waited and published a nicer number. A page arguing that architecture is a maintained system should show it mid-maintenance.
Still open, on the day this published
118 of 122 money pages have no link from a post. That is the same number it was the day before, and the same number it was the day the writing system first measured it. The hub links took an afternoon and are done. The routing is a quarter of work and it is not started, and if it were finished I would have published a better-looking page and taught you less.
The next entry in this page’s history will carry the new number, whatever it is. If it has not moved in three months, that is worth knowing too, and it will be on the page rather than quietly absent.
Now build the thing this page exists to produce
Not in this path.
Reading this page does not make you able to do any of it, and I would be selling you something if I implied otherwise. This next block is the part that does, and it produces one document the next step of the roadmap assumes you have.
The deliverable
Build your SEO site map
Not an XML sitemap. A planning document: every destination your demand supports, what it is, where it sits, how anybody reaches it, and one verb saying what happens to it. Add a row at a time below and the tree draws itself from the parent field. If your rows do not produce a tree, your parents are wrong, and that is the fastest possible way to find out.
Start from
Typed into your own browser and stored there. Nothing is sent anywhere.
Add a destination
Your site map 0 rows
Add a row and the tree appears here.
The rows, with the verbs
The full schema, if you would rather use a spreadsheet
Fourteen fields, grouped by the step each one serves. The field order is the argument: the reading conditions come first so the map is reproducible, the existing URL is asked for before the verb so nobody can decide to create something the site already has, and priority sits last because it is the only field that requires comparing this row against every other row.
Setup
-
Site
The domain this map is for
-
Date
The day you drew it. It will be wrong within a quarter
Demand
-
Cluster
From step three. The name of the need, not a keyword
-
Demand evidence
Volume, or the reason volume is not the evidence here
Intent
-
Intent and page type
From step four. "Tool, 5 of 8"
Page
-
Existing URL
A URL, or None. Fill this in before the verb
-
Destination URL
Where it will live. Same as existing, if you are keeping it
Parent
-
Parent
The page above it, or Home
-
Children, if any
And the reason grouping them here helps somebody
Path
-
In navigation?
Primary, secondary, footer, or none
-
Links in, links out
Which existing page links to it, and where it sends people next
Priority
-
Priority
High, medium or low, compared against your other rows
-
Decision
Keep, create, improve, merge, move, redirect or remove
-
Why
One sentence, so nobody has to rediscover the reasoning
Two rows from my own map, filled in
this page Create
The row that produced the page you are reading, including the fact that it deliberately declines a query with fifteen times the volume.
- Site
- alstonantony.com
- Date
- 22 August 2026
- Cluster
- seo site architecture
- Demand evidence
- 250 US searches, KD 28, traffic potential 2,800. Parent topic is "website structure" at 4,900
- Intent and page type
- Educational. Guide, 8 of 8 organic results are guides or hubs
- Existing URL
- None
- Destination URL
- /seo-site-architecture/
- Parent
- /seo-roadmap/
- Children, if any
- None yet. A keyword mapping page is queued and would sit under this one
- In navigation?
- Secondary, under SEO Roadmap
- Links in, links out
- In from /search-intent/, /seo-roadmap/ and the footer. Out to content SEO when it exists
- Priority
- High
- Decision
- Create
- Why
- Step four hands off to it explicitly, and the roadmap had a named fifth step with no page behind it
a row that says no Ignore
The same map, one row down. Higher volume, lower difficulty, and the answer is still no. Most maps have no rows like this, which is how you can tell nobody applied a filter.
- Site
- alstonantony.com
- Date
- 22 August 2026
- Cluster
- website structure
- Demand evidence
- 3,900 US searches, KD 25, traffic potential 4,100. Fifteen times the row above
- Intent and page type
- Educational, and the audience is web designers. 4 of 8 results are design or site-builder companies
- Existing URL
- None
- Destination URL
- n/a
- Parent
- n/a
- Children, if any
- n/a
- In navigation?
- None
- Links in, links out
- None
- Priority
- Low
- Decision
- Ignore
- Why
- The AI Overview cites Figma, MDN and a sitemap tool and no SEO source at all. Winning it would mean writing for somebody who is not my reader
One test before you close it. Give the map to somebody who was not in the room and ask them to explain your site back to you from it. If they can, it is an architecture. If they ask which of two pages a query belongs to, you have found a row that needs merging, and finding it now is worth more than the whole document.
Then the seven assignments. One per step, in the order the steps come, and each one produces something real rather than a note in a document.
The seven assignments
One assignment per step
Bigger than the exercises and meant to be done once, properly. Finish all seven and you have a site map for a real business, with a date on it and a verb under every row.
0/7
Ticks are stored in your browser and nowhere else. Nothing is sent anywhere, which matters because three of these ask you to write down where your own site is weak.
And if you would rather have a schedule than a list, the same work spread across a week. About half an hour a day, and day seven is the one people skip.
If you want a schedule
The seven-day architecture audit
Same material, paced. Days one and two measure, days three and four decide, days five and six route, and day seven adds exactly one link and writes the date next to it. Skipping day seven leaves you with a diagram, which is what this lesson is trying to talk you out of.
0/7
Put the site name and the date at the top of whatever you end up with. In six months it is the only thing that will tell you whether the structure moved or your judgement did.
Then the checkpoint. Eleven questions, and if you can answer all eleven about your own site you understand architecture better than the person who told you to keep everything within three clicks.
Before you call it finished
The architecture checkpoint
Answerable yes or no about a real site. Any no is a work item, and the last one is the one that catches people who have done everything else.
0/11
Question eleven is the honest test. If you cannot draw the whole site on one page, the answer is not a bigger sheet of paper.
The six decisions, one more time
Not in this path.
If you keep one thing from this page, keep these. Every tactic, framework and tool in site architecture serves one of these six decisions, and any advice that does not answer one of them is decoration.
The framework
Demand, Intent, Page, Parent, Path, Priority
Six questions, in order, and the order is not decorative. You cannot pick a parent before you know what kind of page it is, you cannot design a route to a page you have not decided to build, and you cannot rank six things by importance until all six exist on paper. Steps one and two arrive from the two lessons before this one. Publishing comes after all six, which is the only genuinely controversial thing on this card.
- 01
Demand
Is there something here worth satisfying?
It arrives from step three as a cluster with a number on it. Architecture does not create demand and it cannot rescue a cluster that has none. The only work at this step is throwing rows away, which is the step people skip because it produces nothing to show.
You end up holding A shortened list, with the rows you deleted still visible and a reason next to each.
- 02
Intent
What would actually satisfy it?
It arrives from step four as a brief: a dominant page type with a count, a format, an angle. If you skipped that step you are about to decide where a page lives without knowing what kind of page it is, which is the most expensive order to do this in.
You end up holding One brief per cluster, with a page type and a verb on it.
- 03
Page
Which destination satisfies it, and does it already exist?
The first genuinely architectural question. Not "what do I write" but "how many URLs does this need, of what kind, and do I have one already". Two clusters can share a destination. Twelve clusters can share a destination. Occasionally one cluster needs three.
You end up holding A keyword-to-URL map: cluster, intent, page type, existing URL, verb.
- 04
Parent
Where does that page belong, and what does it belong under?
Grouping, and the test every group has to pass: if somebody lands on the parent, does having these children collected there help them finish something? A parent that exists because a framework said to have a pillar page is a filing cabinet with a URL.
You end up holding A tree. Named parents, assigned children, and the overlaps marked rather than hidden.
- 05
Path
How does anybody, or anything, get there?
Navigation, breadcrumbs, contextual links, URLs. This is the step where the drawing becomes a website, and the step most architecture advice skips by assuming that a folder implies a route. It does not. Google reads the links.
You end up holding At least one crawlable route to every page you care about, written down.
- 06
Priority
How prominently should the structure treat it?
Everything cannot be important. Nav slots, link counts and click depth are the site saying out loud what it thinks matters, and on most sites what it says contradicts what the owner would say. Fixing that contradiction is the cheapest hour in SEO and it produces no content.
You end up holding A ranked list, and the pages you demoted named alongside the ones you promoted.
Then publish. Not before. The whole argument for putting this lesson fifth is that a page whose parent, path and priority were decided after it shipped is a page whose architecture is an accident.
The test I would give a beginner for evaluating any advice about this subject, including mine: ask which of the six it answers, and ask what it would tell you not to build. A framework that only ever produces more pages is a content calendar wearing a method’s clothes.
And the one sentence, if the whole page has to reduce to one. Stop drawing folder trees and start reading your link graph, because that is what Google says it reads and it is the only version of your architecture that exists. Export your internal links tonight. The list of pages nothing points at will be longer than you expect, and it will be the most useful document anybody hands you this quarter, because you handed it to yourself.
Next comes content SEO: what those pages should actually contain, now that you know which ones deserve to exist and where they sit. That lesson is being written. In the meantime the roadmap has the rest of the sequence, the free Excel template from step three has room for the map columns if you would rather keep everything in one file, and the technical SEO checklist covers the implementation half of chapters eighteen and twenty-three.
Glossary
Every word this lesson uses, defined once, in the plainest phrasing that is still correct. Where a term is normally taught with a ranking promise attached, the promise has been left out.
Not in this path.
Reference
Every word this lesson uses, in plain terms
Several are defined against Google’s own documentation rather than the industry’s retelling of it, which on this subject changes the definition and not just the wording. Where a term is usually taught with a ranking claim attached, the claim has been left out on purpose.
26 terms
- Anchor text
- The visible text of a link. Google says it helps both people and Google understand what the target page is about, which makes it the one documented lever in internal linking.
- Architecture debt
- What accumulates when pages are published without deciding where they belong. Paid later in merges, redirects and deletions.
- Breadcrumb
- A trail showing where a page sits and letting a reader move up. Ahrefs states plainly that it is not a replacement for main navigation.
- Click depth Crawl depth
- The number of link clicks from the home page to a page. Google says required clicks and inbound link counts contribute to inferring relative importance within a site.
- Crawlable link
- An <a> element with an href attribute. Google states this is the only form it can crawl, which rules out click handlers and framework routing attributes.
- Dead end
- A page with inbound links and no useful next step out. The reverse of an orphan, and it costs you the reader rather than the crawler.
- Faceted navigation
- Filters that combine into URLs. Powerful for users and capable of generating an unbounded URL space, which Google warns about explicitly.
- Hub page Pillar page
- A parent page that collects and routes to related pages. It has to be somewhere useful to arrive, not just somewhere above other things.
- Information architecture IA
- What information exists and the rules used to classify it. The layer above site architecture, and the one that decides what your categories mean.
- Keyword cannibalization
- Two or more of your pages competing for one need. Not a penalty. An architecture decision you already made and can unmake.
- Keyword-to-URL map Keyword mapping
- A table pairing every cluster with the destination that will own it, plus a verb. The main working document of this lesson.
- Link graph
- The set of links between your pages, considered as a structure. The actual object of this lesson, and the thing a folder tree only describes if somebody linked along it.
- Orphan page
- A page with no internal link pointing at it. Google says every page you care about should have a link from at least one other page on your site.
- Page type
- The kind of thing a page is: guide, tool, category, comparison, product, directory. Chosen by intent, before any words are written.
- Parent and child
- The relationship between a page and the page above it in the hierarchy. Stated by links and navigation, optionally reflected in the URL.
- Parent topic
- Ahrefs’ estimate of the broader keyword that the current number one page for your query mainly ranks for. In this lesson, a free second opinion on how many URLs a subject needs.
- SEO site map
- A planning diagram of the pages a site should have and how they relate. Not an XML sitemap, and the confusion between the two hides a lot of orphan pages.
- Silo
- A section deliberately kept internally cross-linked and externally separate. Useful as a grouping, harmful when tightened until related pages cannot link to each other.
- Site architecture Site structure
- How the pages of a site are grouped and connected. Google says it works this out from the links between pages rather than from URL folders.
- Soft 404
- A page Google treats as missing despite a 200 or a redirect, typically because a retired URL was redirected somewhere irrelevant. The reason bulk redirects to the home page preserve nothing.
- Spoke page Cluster page
- A child page covering one sub-topic in depth, linked from its hub and linking back.
- Subdomain
- blog.example.com. Treated by Google as manageable either way, and by practitioners as a place authority goes to be lonely. Use a subfolder unless engineering forces otherwise.
- Subfolder Subdirectory
- example.com/blog/. The default choice, mainly for operational and analytical reasons that are easier to defend than the ranking argument.
- Taxonomy
- The classification rules themselves. If you cannot write the rule that decides where a new page goes, you have a habit rather than a taxonomy.
- Topic cluster
- A hub plus its spokes, treated as one organizational unit. A model for arranging pages, not a ranking formula.
- XML sitemap
- A machine-readable list of URLs submitted to search engines. A discovery aid. It is not a route and it does not replace a link.
Nothing matches that. If it is a real term and it is not here, it is either something this lesson deliberately leaves to technical SEO, or one of the acronyms somebody coined last quarter for work that already had a name.
Watch the people who own the surface
Not in this path.
Where I would send you instead of a course. All free, filterable by channel, ordered by the chapter they belong to. Every channel here either owns a search surface or owns a crawler you might use to find your own orphan pages, and there are no independent SEO channels on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. That rule costs something here, because several of the best architecture explanations on YouTube are by individual practitioners and none of them is in this list.
Watch the people who own the surface
68 videos, 8 channels, ordered by chapter
Every channel here either owns a search surface or owns a database that measures one. No independent SEO channels, on purpose: a page arguing that you should read what Google published is a poor place to send you to a retelling. Filter by channel, or read it top to bottom as a syllabus. Every id was checked against the YouTube oEmbed endpoint on 22 August 2026. Of 31 candidates a web search attributed to these channels, 24 belonged to individual creators and one was gone entirely, so the list below was rebuilt by searching YouTube directly and verifying every id the same way.
01 Architecture is the link graph
- What’s a preferred site structure? embedded above Google being asked the question this whole page answers, and declining to give a shape. Watch it before chapter one, then notice that the answer is about links and usefulness rather than about a tree.
- What is THE most important thing for SEO? Short, and useful as a check on anybody selling you a restructure as the most useful work available right now.
- How Google Search Works (in 5 minutes) Five minutes on the machine your structure is being read by. Step two of the roadmap in one clip.
- How does Google Search work? The same ground from the developer channel, and franker about how much of discovery is link following.
- Serving and ranking web pages on Search The stage where your architecture stops mattering and your page starts. Worth watching to see how late in the process that is.
- How Google Search serves pages Retrieval through presentation. Pair it with chapter three, which argues architecture is a discovery and prioritization lever rather than a relevance one.
- SEO in 2025: How I’d Learn it if I Were Starting Over A vendor’s ordering of the same curriculum. Compare it against the roadmap you are on and notice where architecture appears in each.
04 Cluster, intent, destination
- Keyword Research Pt 1: How to Analyze Searcher Intent. 1.2. SEO Course by Ahrefs The input to this whole page. If you skipped step four, this four-minute version is the minimum before chapter four makes sense.
- What are Keywords and How to Choose Them? 1.1. SEO Course by Ahrefs Where a cluster comes from, from the company whose parent topic column this page leans on twice.
06 Cannibalization is an architecture decision
- Canonical URLs: How Does Google Pick the One? embedded above What actually happens when you leave two destinations for one need. Google picks, on its own evidence, and this explains what that evidence is.
- How to avoid duplicate content Older and still the clearest short statement that this is a consolidation problem rather than a penalty.
07 Page types, and the queries that want them
- How do I optimize an e-commerce site without rich content? A useful corrective for anybody whose only page type is "article". The answer is not "add a blog to your category page".
- On-Page SEO Pt 2: How to Optimize a Page for a Keyword. 2.2. SEO Course by Ahrefs Deliberately out of scope for this lesson and here so you can see the boundary. Page type is decided here; everything in this video happens afterwards.
10 Four architectures, four business models
- Welcome to Ecommerce Essentials! The opener of Google’s own ecommerce series, which is where the documented menu to category to subcategory to product chain comes from.
- How to SEO optimize your ecommerce website (8 Tips) The applied version. The structural tips are the first half; the rest belongs to technical SEO.
- SEO for ecommerce, Google Search Console Training Where to look in Search Console once a catalog structure is live. Useful for seeing which architecture problems the reports can and cannot see.
- Ecommerce SEO, Get Traffic To Your Online Store A vendor take on the catalog shape. Watch it for the category page treatment and ignore the tool pitch.
18 A link Google cannot follow is not a link
- How Google Search crawls pages embedded above The mechanism the whole page rests on: URLs enter the system by being linked from pages already known.
- How does Google Search crawl a web page? The shorter version of the same thing, useful for sending to somebody who will not watch ten minutes.
- How Browsers Really Parse HTML (and What That Means for SEO) Technical, and it is the best explanation of why your navigation has to be anchors in the source rather than anchors after hydration.
- 3 Tips for Crawling Errors Older. The relevant part is how many crawl errors turn out to be a navigation that generates links to nothing.
- How Alamy’s website migrated to React A real migration where the routing was the risk. Worth watching before you let anybody rebuild your navigation as a single-page app.
22 Breadcrumbs, and what they are not
- What are some best practices for indicating breadcrumbs? embedded above Short and specific. Watch this rather than reading a listicle about breadcrumbs.
- Can I place multiple breadcrumbs on a page? The answer matters if your page genuinely has two parents, which is a real architecture situation rather than a mistake.
- Why aren’t breadcrumbs displaying in search results for my site? Useful for separating "my breadcrumbs are wrong" from "Google chose not to show them", which are different problems.
- Rich Snippets: Breadcrumbs The oldest video in this library and still the clearest on what the trail is for.
23 Faceted navigation and pagination
- Faceted Navigation and PRG. Lesson 12/34, Semrush Academy embedded above The most systematic free explanation of the filter problem I could find from a vendor. Chapter twenty-three stops where this video starts.
- Pagination. Lesson 13/34, Semrush Academy Pagination as an architecture problem rather than a canonical problem. Watch it if you have a large category or an archive.
- Pagination and SEO Google’s own take. Shorter, and less prescriptive than most third-party advice on the same subject.
- Should you block your Search result pages? The internal-search half of the facet question, and a reminder that Googlebot generally does not use your search box.
26 Orphans, dead ends and the sitemap that hides them
- How can I get Google to index more of my Sitemap URLs? embedded above The exact question somebody with 97 orphan pages asks, and the answer is not "resubmit the sitemap". Watch it before chapter twenty-six.
- Help! Google Search isn’t indexing my pages The longer treatment, and structure comes up early in it, which is not what most people expect from an indexing video.
- 3 tips for setting up a sitemap Do the sitemap properly, and then notice that none of these three tips is a substitute for a link.
- Sitemaps in Search Console, Google Search Console Training Where to check whether the file is even being read, which is a different question from whether it is working.
- How Google Search indexes pages The stage between crawling and serving. Relevant here because a page with no inbound link often never reaches it.
27 Internal links do three jobs
- How to use internal linking for SEO embedded above Google on the three jobs a link does, from the people whose crawler is doing two of them. If you watch one video from this library, watch this one.
- Will multiple internal links with the same anchor text hurt a site’s ranking? The anchor text question, answered by Google rather than by the anchor-variation folklore.
- Should internal links use rel="nofollow"? A short no, and the reasoning is a useful demonstration that link sculpting is not the lever people think it is.
28 Building for 200 pages, not 20
- How long does SEO take for new pages? Worth watching before you commit to two hundred pages. The timeline is the constraint your architecture has to survive.
- SEO for a New Website, Do This the First 30 Days A reasonable sequencing for a site with nothing on it yet. Compare its ordering against the six steps on this page.
29 Competitor structure is evidence, not instructions
- Keyword Research Pt 3: Understanding Ranking Difficulty. 1.4. SEO Course by Ahrefs How to read a competitor’s page rather than their score. The architectural version of the same skill is chapter twenty-nine.
- Official Ahrefs Tutorial: How to use Ahrefs to Improve SEO Watch for the site structure report specifically. It is the fastest way to see somebody else’s architecture as a tree.
- How to Create Content that’s "Better" than Your Competitor’s Useful counterweight. Structural imitation is the version of this mistake nobody warns you about, because the structure is invisible.
30 One structure for humans, crawlers and retrieval
- How I Structure Content so AI Actually Cites It embedded above The most concrete version of the AI structure argument from a vendor. Watch it, then check chapter thirty, which grades the claim as inference rather than measurement.
- AI Websites, Crawling and Search Console updates (Q1 ’26) What actually changed on Google’s side, from Google. Shorter and less exciting than the average AI SEO video, which is the point.
- How AI Is Changing Google Search and SEO Google’s own framing. Note how much of it is the same discovery and clarity advice with a new name on it.
- Introducing AI Mode: A New Experiment in Google Search The surface itself. Worth seeing so the retrieval discussion is about something you have used.
- Introducing Grounding with Google Search The developer view of retrieval, which is the clearest look at what "cited" actually means mechanically.
- Level up AI with Grounding with Google Search More applied. Useful for understanding why a well-grouped set of pages is easier to ground against than a blog archive.
- Search, 12 Days of OpenAI: Day 8 The launch of the surface where your structure is consumed as retrieval rather than navigation. Straight from the people who built it.
- A quick guide: how to search in ChatGPT Two minutes, and it will change how you think about what a "page" is to a user who never sees your navigation.
- Updates to deep research in ChatGPT Relevant because a deep research run reads many pages from one site, which is the first use case where your internal linking is being followed by a machine on purpose.
- Perplexity: A New Search Perspective The other surface. Worth watching for how the citation is presented, since that presentation is what a structured site is competing for.
- Introducing: Deep Research on Perplexity Same argument as the OpenAI deep research clip, from a company with a different index. Watch one of the two.
- SEO in 2026: How I’d Rank in Google in the AI Era A vendor forecast. Useful for spotting which parts of it are architecture advice wearing an AI label.
- How to Perform a Site Audit. Lesson 1/7, Semrush Academy Where to start if the debt chapter described your site and you do not know what to run first.
- How to Use SEMrush Site Audit for Your Technical SEO. Lesson 1/9, SEMrush Academy The follow-on. The orphan and internal link reports are the two worth opening for this lesson; the rest is technical SEO.
What this lesson deliberately refuses to teach
Not in this path.
Ten subjects belong elsewhere, and cramming them in here would wreck the one thing this page exists to produce. The failure mode of a lesson like this is turning into an everything lesson, and every item below is real, learnable, and worth nothing to somebody who cannot yet export their own internal links and name the pages nothing points at.
What to put on the pages
That is content SEO, and it is the next step. This lesson decides what exists and how it connects. Deciding both at once is how a plan becomes a first draft.
Title tags, headings and on-page optimization
On-page SEO. The architecture decides which page owns a query; on-page decides how that page says so.
Internal link auditing at scale
A dedicated internal linking lesson. Here you need the three jobs a link does and the habit of routing one deliberately. Anchor distribution modelling can wait.
Canonical tags, robots directives, noindex
Technical SEO. Architecture creates the situations that need them, which is why faceted navigation appears here as a cause and not as an implementation.
The fundamentals lesson covers the basicsPageRank sculpting and link equity modelling
Nobody outside Google can measure it and the published models are twenty years old. What you can do is route a link from a page that gets traffic to a page that needs it, and measure the page rather than the model.
Structured data and schema markup
Technical and AI SEO. Breadcrumb markup gets a mention because it belongs to a navigation decision. The rest is a different lesson.
Faceted navigation implementation
Ecommerce and technical SEO. Parameter handling, robots rules and canonical strategy for a filter space are a project, and they follow from the architectural decision rather than replacing it.
Site migrations
A discipline of its own. This lesson’s advice on migration is mostly do not, and where you must, map every URL by hand.
Programmatic SEO at scale
Ten thousand generated pages is an architecture question answered before the generator runs. The scale chapter says what to decide first; building the generator is elsewhere.
CMS and framework specifics
Routing in Astro, WordPress permalinks, Shopify collections. All real, all different, and none of them changes a single decision in the six steps.
The learner only needs six things from this page. Architecture is the link graph rather than the folder tree. One intent gets one destination. Page type is decided before content. A parent has to help somebody finish something. Every page you care about needs a crawlable route in and a useful route out. And the deliverable is a map with a verb on every row, not a diagram.
Page history
What changed on this page, and when, because the central receipt on it is a measurement of my own site that moves every time I publish anything.
Not in this path.
Page history
What changed, and when
The central receipt on this page is a measurement of this site, taken twice on two consecutive days, and it will move every time anything is published here. So the changes get listed rather than absorbed silently into an "updated" stamp, and the next link-graph run appears here as its own entry with the new numbers in it, including the ones that got worse.
-
23 August 2026
First publication
- A 301 added from /seo-tools/semrush-tutorial/ to /seo-tools/semrush/tutorial/. That is the third address the Semrush tutorial has had: a flat slug under the tools folder, then a slug under /seo-guides/, and now a spoke under its own hub. Two of the three old URLs are in the redirect file and both point at the same page.
- The redirect file went from 46 rules to 47 the day after this page published, which is the debt chapter happening to the page while it was live. Every count on the page that reads from that file moved with it, and the two bar charts in the debt receipt now scale to the largest row instead of to a hard-coded divisor, because the hard-coded one was already wrong by one.
- Nothing else on the page changed. The link-graph numbers, the keyword pull and the two results pages are still the 22 August readings and are stamped as such.
-
22 August 2026
- First publication, as step five of the SEO roadmap. It sits between search intent and content SEO because intent tells you what kind of page deserves to exist and this step decides where it lives, what links to it and how prominently the site treats it.
- The central claim is quoted rather than asserted. Google’s ecommerce site structure documentation, read on this date and last updated 10 December 2025, states that it "generally doesn’t look at the structure of URLs to work out the structure of a site" and analyzes "the linkages between pages" instead. Every architecture tutorial that opens with a folder diagram is arguing with that sentence without quoting it.
- The receipt is my own site, measured with a tool anybody can run. `node tools/authority.mjs` against the content collections, twice: once on 21 August 2026 and again on 22 August. Both readings are published, including the fact that the money-page routing number did not improve at all between them.
- The specific number: 103 tool listings sit in a consistent /seo-tools/{tool}/ folder tree and 97 of them have no editorial link pointing at them. Perfect folders, no architecture. That is the whole page in one figure and it is a figure about my own site rather than a hypothetical one.
- Twenty keywords pulled from Ahrefs Keywords Explorer in a single request against the US database on this date. All twenty returned informational intent and exactly one also returned commercial. Eight of the twenty are ways of saying the same thing, they carry 8,550 US searches a month between them, and Ahrefs assigns them four parent topics rather than eight.
- Two results pages pulled from Ahrefs SERP Overview on this date. Not one of the eight organic results ranking for "seo site architecture" has that phrasing as its own top keyword; all eight were built for a sibling phrasing and rank for this one anyway. One of the eight is a page whose top keyword is Moz’s own product name.
- The URL decision is published in chapter five rather than hidden in a commit message. "website structure" carries fifteen times the volume at lower difficulty and four of its eight results belong to design and website-builder companies, with an AI Overview citing Figma, MDN and a sitemap tool and no SEO source at all. The bigger query was declined on purpose.
- The taxonomy-debt chapter is costed with this site’s own redirect file: public/_redirects, replaced 21 August 2026, 46 rules, nine of them retiring a single category that never meant anything a reader could use. The blog category enum went from five values to two in three days.
- The claim-grading table exists because this subject has the worst evidence hygiene in SEO. Seven claims are documented and quoted, two are on the record without a published model, two are inference, and four appear in the myth list as invented, including the "URL slug is 50% of relevance" number and the hard three-click threshold.
- All 68 video ids verified against the YouTube oEmbed endpoint on this date, and every channel name in the library is the one oEmbed returned rather than the one a search result implied. That check did most of the curation: of 31 candidates a web search attributed to one of these eight channels, 24 belonged to individual creators and one returned an HTTP 401 because the video no longer exists, so six survived and the rest of the library was found by searching YouTube directly and verifying every id the same way.
- Google’s crawl budget guide was read on this date and deliberately not leaned on: it states that it applies to sites over one million pages, or over ten thousand pages changing daily. Almost nobody reading this page is in either bracket, and citing it anyway is one of the ways this subject gets oversold.
If you find something here that a current Google, Ahrefs or Semrush document contradicts, that is a bug in this page and I would rather hear about it than have it quietly rot. The contact page works.
Sources
Every claim above, traced to the document it came from, with the date it was read and the date the document says it was last updated. Two of the entries are files in my own repository.
Not in this path.
Everything on this page, sourced
18 sources, each read on 22 August 2026. Where a document prints its own last-updated date, both dates are shown, because the document’s own date is the one that tells you whether the guidance has moved since I quoted it. Where an article has an author, the author is named rather than the brand.
Google Search Central 9
- Ecommerce Website Navigation Structure Doc updated 10 Dec 2025, read 22 Aug 2026
- URL Structure Best Practices for Google Search Doc updated 10 Dec 2025, read 22 Aug 2026
- SEO Link Best Practices for Google Doc updated 10 Dec 2025, read 22 Aug 2026
- SEO Starter Guide: the Basics Doc updated 10 Dec 2025, read 22 Aug 2026
- Large Site Owner’s Guide to Managing Your Crawl Budget Doc updated 22 Jul 2026, read 22 Aug 2026
- In-Depth Guide to How Google Search Works Read 22 Aug 2026
- Should I structure my site using subdomains or subdirectories? Checked 22 Aug 2026
- What’s a preferred site structure? Checked 22 Aug 2026
- How does URL structure affect PageRank? Checked 22 Aug 2026
Semrush 1
- Website Architecture: Best Practices for SEO Site Structures, by Dana Nicole Published 30 Jan 2025, read 22 Aug 2026
Ahrefs 2
- How to Structure Your Website Architecture for SEO, by Chris Haines Published 23 May 2023, read 22 Aug 2026
- Keywords Explorer and SERP Overview, US database, twenty keywords and two results pages Pulled 22 Aug 2026
Search Engine Land 1
- Site Architecture for SEO: Structure That Ranks and Scales, by Kody Wirth and Sophie Atkinson Updated 27 Nov 2025, read 22 Aug 2026
Backlinko 1
- Website Architecture: How to Setup an SEO-Friendly Structure Read 22 Aug 2026
Lumar 1
- Site Architecture: SEO Best Practices Read 22 Aug 2026
SE Ranking 1
- Website Structure for SEO Read 22 Aug 2026
Alston Antony, content-machine 1
- tools/authority.mjs, the editorial link graph of this site Run 21 and 22 Aug 2026
Alston Antony, alstonantony.com 1
- public/_redirects, 47 rules, replaced 21 August 2026 and amended 23 August 2026 Replaced 21 Aug 2026
Six of these are Google’s own documentation and four of them are quoted verbatim on this page, with the sentence rather than a paraphrase, because the paraphrases in circulation are how the folder tree became the whole subject. One is a document I read and deliberately did not lean on: the crawl budget guide states that it applies to sites over a million pages, or over ten thousand pages changing daily, and almost nobody reading this is in either bracket. If one of these sources now says something different from what I have written, the source wins and this page is out of date.
The last two entries are mine, and they are the fastest-rotting things here. The link graph is whatever ’node tools/authority.mjs’ said on 21 and 22 August 2026, and it changes every time I publish anything. The redirect file was 46 rules on the day it was replaced and 47 two days later, which is the chapter on debt happening to this page while it was live. Both are in my own repositories rather than in a screenshot, and the tool is the same one this page tells you to run against your own site. If your numbers disagree with mine, yours are about your site and they are the ones that matter.