Crawling
Crawling is the process in which search engine bots such as Googlebot discover URLs and download their content. Pages are found mainly through links and sitemaps; robots.txt controls which paths crawlers may request. If an important page cannot be crawled, it cannot be indexed, and nothing on it (including entity markup) will count.
How do crawlers find pages?
Google finds most new pages by following links from pages it already knows, and from sitemaps site owners submit.[1][2] Google can generally only follow links that are <a> elements with an href attribute.[3] Officially documented
What is Googlebot?
Googlebot is the name of Google’s web crawler, which has smartphone and desktop variants.[4] Because of mobile-first indexing, Google uses the mobile version of a site’s content, crawled with the smartphone agent, for indexing and ranking.[5] Officially documented
What does robots.txt do?
A robots.txt file tells crawlers which URLs they may request. The format is standardized as RFC 9309.[6] Google warns that robots.txt is used mainly to avoid overloading a site and “is not a mechanism for keeping a web page out of Google”; use noindex or password protection instead.[7] Officially documented
User-agent: * Disallow: /cart/ Sitemap: https://www.example.com/sitemap.xml
AI crawlers use their own user agents; AEO.wiki lists them in AI crawlers and robots.txt.
What is crawl budget?
Crawl budget is the time and resources Google spends crawling a site. Google’s guide is intended mainly for large sites (1 million+ unique pages changing about weekly) or sites with 10,000+ pages changing daily.[8] Officially documented Most small sites do not need to worry about it.
How do you fix crawl problems?
- Link every important page from at least one other crawlable page.
- Keep an XML sitemap of canonical URLs only.[2]
- Check that robots.txt does not block CSS, JavaScript or key sections.
- Return correct status codes; persistent 5xx errors slow Google’s crawling.[9]
- Test a URL with Search Console’s URL Inspection tool.[10]
Next step: indexing.
Frequently asked questions
Does blocking a page in robots.txt remove it from Google?
No. Google says robots.txt is not a mechanism for keeping a page out of Google. A blocked URL can still be indexed from links. Use noindex, and do not block the page, so Google can see the rule.
Do I need an XML sitemap?
Not always, but Google recommends one for large sites, new sites with few links and sites with rich media. Small, well-linked sites can be crawled without one.
What is crawl budget?
The amount of crawling Google does on a site. It matters mainly for very large or rapidly changing sites.
See also
References
Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "In-depth guide to how Google Search works". Google Search Central.
- ^ "Build and submit a sitemap". Google Search Central.
- ^ "Link best practices for Google". Google Search Central.
- ^ "What is Googlebot". Google Search Central.
- ^ "Mobile site and mobile-first indexing best practices". Google Search Central.
- ^ "RFC 9309: Robots Exclusion Protocol". IETF (RFC Editor). Published September 2022.
- ^ "Introduction to robots.txt". Google Search Central.
- ^ "Crawl budget management for large sites". Google Search Central.
- ^ "How HTTP status codes, and network and DNS errors affect Google Search". Google Search Central.
- ^ "URL Inspection tool". Search Console Help.