EntitySEO.wikiThe Entity SEO Reference

Crawling

From EntitySEO.wiki, the entity SEO reference · Last reviewed · By · Published by Local Blitz · How we research

Crawling is the process in which search engine bots such as Googlebot discover URLs and download their content. Pages are found mainly through links and sitemaps; robots.txt controls which paths crawlers may request. If an important page cannot be crawled, it cannot be indexed, and nothing on it (including entity markup) will count.

How do crawlers find pages?

Google finds most new pages by following links from pages it already knows, and from sitemaps site owners submit.[1][2] Google can generally only follow links that are <a> elements with an href attribute.[3] Officially documented

What is Googlebot?

Googlebot is the name of Google’s web crawler, which has smartphone and desktop variants.[4] Because of mobile-first indexing, Google uses the mobile version of a site’s content, crawled with the smartphone agent, for indexing and ranking.[5] Officially documented

What does robots.txt do?

A robots.txt file tells crawlers which URLs they may request. The format is standardized as RFC 9309.[6] Google warns that robots.txt is used mainly to avoid overloading a site and “is not a mechanism for keeping a web page out of Google”; use noindex or password protection instead.[7] Officially documented

User-agent: *
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xml

AI crawlers use their own user agents; AEO.wiki lists them in AI crawlers and robots.txt.

What is crawl budget?

Crawl budget is the time and resources Google spends crawling a site. Google’s guide is intended mainly for large sites (1 million+ unique pages changing about weekly) or sites with 10,000+ pages changing daily.[8] Officially documented Most small sites do not need to worry about it.

How do you fix crawl problems?

Next step: indexing.

Frequently asked questions

Does blocking a page in robots.txt remove it from Google?

No. Google says robots.txt is not a mechanism for keeping a page out of Google. A blocked URL can still be indexed from links. Use noindex, and do not block the page, so Google can see the rule.

Do I need an XML sitemap?

Not always, but Google recommends one for large sites, new sites with few links and sites with rich media. Small, well-linked sites can be crawled without one.

What is crawl budget?

The amount of crawling Google does on a site. It matters mainly for very large or rapidly changing sites.

References

Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.

  1. ^ "In-depth guide to how Google Search works". Google Search Central.
  2. ^ "Build and submit a sitemap". Google Search Central.
  3. ^ "Link best practices for Google". Google Search Central.
  4. ^ "What is Googlebot". Google Search Central.
  5. ^ "Mobile site and mobile-first indexing best practices". Google Search Central.
  6. ^ "RFC 9309: Robots Exclusion Protocol". IETF (RFC Editor). Published September 2022.
  7. ^ "Introduction to robots.txt". Google Search Central.
  8. ^ "Crawl budget management for large sites". Google Search Central.
  9. ^ "How HTTP status codes, and network and DNS errors affect Google Search". Google Search Central.
  10. ^ "URL Inspection tool". Search Console Help.