Indexing
Indexing is the stage where a search engine analyses a crawled page and decides whether and how to store it. Google groups duplicate pages, picks one canonical URL to show, and may skip pages it considers low value. You influence this with redirects, rel="canonical", sitemaps and noindex, and you check it in Search Console’s Page indexing report.
What happens during indexing?
Google analyses the text, images and video on a page, works out whether it duplicates another page, and stores information about the canonical version in its index.[1] Officially documented Google states that it does not guarantee to crawl, index or serve any page.[1]
What is a canonical URL?
The canonical is the URL Google treats as the representative of a set of duplicate pages. Google lists the methods in order of strength: redirects are “a strong signal”, rel="canonical" link annotations are a strong signal, and sitemap inclusion is “a weak signal”.[2] Officially documented
<link rel="canonical" href="https://www.example.com/services/roof-repair">
How do you keep a page out of the index?
Add <meta name="robots" content="noindex"> or an X-Robots-Tag header. For the rule to work, the page must not be blocked by robots.txt, otherwise the crawler never sees it; Google also does not support noindex inside robots.txt.[3][4] Officially documented
Why is a page crawled but not indexed?
| Cause | Fix |
|---|---|
| Duplicate of another URL | Consolidate with a redirect or canonical |
| Thin or unhelpful content | Improve or merge the page ([5]) |
| Accidental noindex | Remove the rule; request reindexing |
| Soft 404 or error status | Return 200 for real content, 404/410 for removed pages ([6]) |
| Content only visible after JavaScript fails | Server-render key content ([7]) |
Search Console’s Page indexing report shows which URLs are indexed and the reason others are not.[8] Officially documented
Why does indexing matter for entities?
Your entity home, About page, author profiles and contact page should all be indexed, canonical and consistent, because these are the pages that state your core facts. The same pages are what answer engines retrieve; AEO.wiki covers access and rendering in rendering, speed and access.
Frequently asked questions
Is rel=canonical a directive?
Google treats it as a strong signal, not an absolute command. Redirects are a stronger signal.
How do I check if a page is indexed?
Use the URL Inspection tool or the Page indexing report in Google Search Console.
Should I noindex thin pages?
Improving or merging them is usually better. Use noindex for pages that should exist for users but not appear in search, such as internal search results.
See also
References
Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "In-depth guide to how Google Search works". Google Search Central.
- ^ "How to specify a canonical URL with rel="canonical" and other methods". Google Search Central.
- ^ "Block Search indexing with noindex". Google Search Central.
- ^ "Robots meta tag, data-nosnippet, and X-Robots-Tag specifications". Google Search Central.
- ^ "Creating helpful, reliable, people-first content". Google Search Central.
- ^ "How HTTP status codes, and network and DNS errors affect Google Search". Google Search Central.
- ^ "Understand JavaScript SEO basics". Google Search Central.
- ^ "Page indexing report". Search Console Help.