How search engines understand entities
Search engines understand entities by combining several signals: the language on a page, structured data that names things explicitly, links and anchor text, and facts corroborated across many independent sources such as Wikipedia, licensed data and official sites. Google documents the inputs in general terms but not the weighting.
Which signals do search engines use?
| Signal | What Google says | Evidence |
|---|---|---|
| Page text | Language systems such as BERT and MUM help understand queries and content.[1][2] | Officially documented |
| Structured data | Gives “explicit clues” about a page’s meaning.[3] | Officially documented |
| Knowledge Graph sources | Facts come from “materials shared across the web, as well as from open source and licensed databases.”[4] | Officially documented |
| Links | Google uses links to discover pages and as a relevance signal; anchor text describes the target.[5][6] | Officially documented |
| Entity owners | Verified people and organizations can suggest knowledge panel changes.[7] | Officially documented |
| Consistency across the web | Widely recommended; weighting not published. | Practitioner practice |
How does recognition work in text?
Google does not publish its entity-recognition pipeline, but Google Cloud’s Natural Language API shows the general idea: it finds entity mentions in text, assigns a type, a salience score from 0 to 1, and, when it can, a Knowledge Graph machine ID and a Wikipedia URL.[8][9] Officially documented That is a separate Cloud product, not Google Search, so treat its output as an approximation (see finding entities with the Natural Language API).
How are entities reconciled across sources?
When many sources describe the same thing, a knowledge graph has to decide they are one entity. Shared identifiers make that easier: Wikidata items carry external identifiers such as official websites, social profiles and authority IDs,[10] and schema.org’s sameAs property points to “a reference Web page that unambiguously indicates the item’s identity.”[11] Officially documented
What can a site owner influence?
- Clear names and definitions in visible text.
- JSON-LD with stable
@idvalues and accuratesameAslinks (@id and entity graphs). - One authoritative entity home page.
- Consistent facts on profiles, directories and Wikidata, where notability allows.
- Editorial coverage from independent sources (third-party presence on AEO.wiki).
Frequently asked questions
Does Google read schema to understand entities?
Google says structured data gives it explicit clues about a page's meaning. It is one input among many, and it must match visible content.
Does Google use Wikipedia for the Knowledge Graph?
Google has said Wikipedia is a commonly cited source but that facts come from many sources, including licensed data.
See also
References
Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "Understanding searches better than ever before (Pandu Nayak)". Google (The Keyword). Published October 25, 2019.
- ^ "MUM: A new AI milestone for understanding information (Pandu Nayak)". Google (The Keyword). Published May 18, 2021.
- ^ "Introduction to structured data markup in Google Search". Google Search Central.
- ^ "A reintroduction to our Knowledge Graph and knowledge panels (Danny Sullivan)". Google (The Keyword). Published May 20, 2020.
- ^ "Link best practices for Google". Google Search Central.
- ^ "A guide to Google Search ranking systems". Google Search Central.
- ^ "Submit feedback on content about you". Google Knowledge Panel Help.
- ^ "Natural Language API basics". Google Cloud.
- ^ "Analyzing entities". Google Cloud.
- ^ "Wikidata: Identifiers". Wikidata.
- ^ "sameAs property". Schema.org.