EntitySEO.wikiThe Entity SEO Reference

Entities and Knowledge Graph: patents and papers

From EntitySEO.wiki, the entity SEO reference · Last reviewed · By · Published by Local Blitz · How we research

The patents and papers behind entity SEO describe how search systems store facts about things, assign entity types, link mentions in text to the right entity, and judge which entities a page is really about. Eight patents (seven from Google, one from Microsoft) and seven research papers are listed here, each checked against Google Patents, the ACL Anthology or the publisher’s DOI record.

Read this first: a patent shows that a company sought legal protection for a method. It does not show that the method is used in Google Search, used as written, or still used. Only systems Google itself names (for example in its ranking systems guide) are confirmed. Dates and legal status come from Google Patents, which notes that its legal status is “an assumption and is not a legal conclusion”. See how to read a patent.

What does this theme cover?

These documents deal with the raw material of the Knowledge Graph: objects (entities), facts about them, entity types, and entity linking (deciding which entity a word on a page refers to). Google publicly launched the Knowledge Graph in May 2012 to understand “things, not strings”.[1] Several patents below were filed years earlier, which shows the ideas were being researched long before launch. That does not show which of them shipped. Practitioner practice

Which patents describe entities, facts and salience?

Browseable fact repository

Patent US 7,774,328 B2 Patent: use unconfirmed · Assignee Google LLC · Inventors Andrew W. Hogue, Jonathan T. Betz · Priority Feb 17, 2006 · Filed Feb 17, 2006 · Granted Aug 10, 2010 · Status Expired - Fee Related

What it describes: A “fact repository” that stores objects together with facts about them. A search returns the objects associated with facts relevant to the query, and each object is shown with a selection of its facts, ordered by relevance to the query.

Why it matters for SEO: An early description of results built from entities and their facts instead of lists of documents: the same idea people now see in knowledge panels. It is a reason to publish clear, factual statements about your entity (founding date, location, people, products) on your entity home.

Learning facts from semi-structured text

Patent US 7,769,579 B2 Patent: use unconfirmed · Assignee Google LLC · Inventors Shubin Zhao, Jonathan T. Betz · Priority May 31, 2005 · Filed May 31, 2005 · Granted Aug 3, 2010 · Status Expired - Fee Related

What it describes: Bootstrapping new facts from semi-structured text. The system starts with a few known “seed” facts about an object, finds documents that contain several of them, learns the contextual patterns around those facts, and uses the patterns to extract more facts.

Why it matters for SEO: Consistent, labelled facts (tables, definition lists, “Founded: 2009” style rows) are the kind of semi-structured text this method learns from. Presenting key facts the same way across your site and profiles makes them easier to extract and corroborate. Practitioner practice

Entity type assignment

Patent US 7,970,766 B1 Patent: use unconfirmed · Assignee Google LLC · Inventors Farhan Shamsi, Alex Kehlenbeck, David Vespe, Nemanja Petrovic · Priority Jul 23, 2007 · Filed Jul 23, 2007 · Granted Jun 28, 2011 · Status Active

What it describes: Assigning an entity type (person, organization, place and so on) to objects of unknown type. Features are generated from each object’s facts, a model is trained on objects of known type, and the model assigns types to the rest.

Why it matters for SEO: Entity type drives how an entity is displayed and compared. Declaring the right schema.org @type and publishing type-typical facts (a person’s job title, an organization’s founding date) removes guesswork. See Organization and Person schema.

Anchor text summarization for corroboration

Patent US 9,208,229 B2 Patent: use unconfirmed · Assignee Google LLC · Inventors Jonathan T. Betz, Shubin Zhao · Priority Mar 31, 2005 · Filed Mar 31, 2006 · Granted Dec 8, 2015 · Status Active

What it describes: Corroborating a set of facts using links. When the anchor text of links pointing to a document matches the name of a set of facts (an entity), that document is treated as relevant and used to corroborate or refute the facts.

Why it matters for SEO: Links whose anchor text is your entity’s name can help associate the linked page with the entity. Consistent naming in citations and profile links supports that association; see links and NAP and citations.

Finding and disambiguating references to entities on web pages

Patent US 8,122,026 B1 Patent: use unconfirmed · Assignee Google LLC · Inventors Leonardo A. Laroco, JR., Nikola Jevtic, Nikolai V. Yakovenko, Jeffrey Reynar · Priority Oct 20, 2006 · Filed Oct 20, 2006 · Granted Feb 21, 2012 · Status Expired - Fee Related

What it describes: An iterative method for disambiguating references to entities in documents. An initial model identifies documents that refer to an entity from their features; feature counts from those documents build a better model, and the loop repeats.

Why it matters for SEO: This is the core problem of entity disambiguation. Features that co-occur with your entity (location, industry, founders, related entities) are what separate you from namesakes, so mention them consistently.

Identifying salient items in documents

Patent US 9,251,473 B2 Patent: use unconfirmed · Assignee Microsoft Technology Licensing LLC · Inventors Michael Gamon, Patrick Pantel, Xinying Song, Tae Yano, Johnson Tan Apacible · Priority Mar 13, 2013 · Filed Mar 13, 2013 · Granted Feb 2, 2016 · Status Active

What it describes: Training models to predict which items (such as entities) are salient on a web page. Salience labels come from search logs: of all the query-driven visits to a page, the share whose queries included the item. The models use page features, including page classification features.

Why it matters for SEO: A Microsoft patent, included because it defines salience by what searchers were looking for when they visited a page. Google’s Natural Language API also reports a salience score for each entity.[2] Make your main entity the clear subject of the page; see entity salience.

Ranking search results based on entity metrics

Patent US 10,235,423 B2 Patent: use unconfirmed · Assignee Google LLC · Inventors Hongda Shen, David Francois Huynh, Grace Chung, Chen Zhou, Yanlai Huang, Guanghua Li · Priority Dec 12, 2012 · Filed Dec 12, 2012 · Granted Mar 19, 2019 · Status Active

What it describes: Ranking search results for entity queries by combining several metrics with weights that depend on the entity type. The description names four metrics: relatedness, notable entity type, contribution and prize.

Why it matters for SEO: It illustrates that “good” can mean different things for different entity types (awards for a film, contributions for a scientist). Publish the attributes that matter for your type, with sources: awards, notable work, affiliations. Practitioner practice

Question answering using entity references in unstructured data

Patent US 9,477,759 B2 Patent: use unconfirmed · Assignee Google LLC · Inventors Dvir Keysar, Tomer Shmiel · Priority Mar 15, 2013 · Filed Mar 15, 2013 · Granted Oct 25, 2016 · Status Active

What it describes: Question answering from entity references in ordinary search results. For a query that seeks a type of entity, the system collects entity references found in the results, ranks them, selects an entity and returns it as the answer.

Why it matters for SEO: Answers can be assembled from mentions across many web pages, not only from structured databases. Being mentioned, by name, next to the attribute people ask about (“best X in Y”, “founder of Z”) is what such a system would count. This overlaps with citations and brand mentions on AEO.wiki.

Which research papers underpin entity understanding?

Freebase: a collaboratively created graph database for structuring human knowledge

Authors Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, Jamie Taylor · Published Proceedings of the 2008 ACM SIGMOD International Conference on Management of Data (2008) · ACM Digital Library Research paper

What it describes: Freebase, a shared, collaboratively edited graph database of general knowledge with a schema anyone could extend.

Why it matters for SEO: Google bought Metaweb, Freebase’s developer, in 2010,[3] and later shut the Freebase API and offered its data dumps and Wikidata mappings.[4] Freebase is the best-documented ancestor of the Knowledge Graph’s open data.

Knowledge Vault: a web-scale approach to probabilistic knowledge fusion

Authors Xin Dong, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Ni Lao, Kevin Murphy et al. · Published KDD ’14 (ACM SIGKDD) (2014) · ACM Digital Library Research paper

What it describes: A Google research system that extracts facts from web content and combines them with prior knowledge from existing knowledge bases, assigning each fact a probability of being true.

Why it matters for SEO: Facts are treated as probabilistic and corroborated across sources. Consistent facts about your entity across your site, profiles and third-party sources are easier to trust than a single claim. Practitioner practice

Wikidata: a free collaborative knowledgebase

Authors Denny Vrandečić, Markus Krötzsch · Published Communications of the ACM 57(10) (2014) · ACM Digital Library Research paper

What it describes: The design of Wikidata, the Wikimedia knowledge base that stores structured, referenced statements anyone can reuse.

Why it matters for SEO: Wikidata IDs are a common way to disambiguate entities in sameAs. See Wikidata and finding entities with Wikidata.

Large-Scale Named Entity Disambiguation Based on Wikipedia Data

Authors Silviu Cucerzan · Published EMNLP-CoNLL 2007 (2007) · ACL Anthology Research paper

What it describes: Linking names in text to Wikipedia entities at scale, using Wikipedia’s own pages, categories and link structure as knowledge.

Why it matters for SEO: A foundational paper on entity linking: the step that turns a mention of “Mercury” into the planet, the element or the singer.

Wikify!: linking documents to encyclopedic knowledge

Authors Rada Mihalcea, Andras Csomai · Published CIKM ’07 (ACM) (2007) · ACM Digital Library Research paper

What it describes: Automatically finding the important concepts in a document and linking each one to the right Wikipedia article.

Why it matters for SEO: Shows how a machine picks the key concepts of a page and resolves each one. Clear, specific wording makes that resolution easier.

Robust Disambiguation of Named Entities in Text

Authors Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Manfred Pinkal, Marc Spaniol et al. · Published EMNLP 2011 (2011) · ACL Anthology Research paper

What it describes: A disambiguation method that combines how often a name refers to an entity, how well the surrounding text matches the entity, and how coherent the chosen entities are with each other.

Why it matters for SEO: Coherence matters: entities mentioned together should make sense together. A page that names your city, industry and partners helps pick you over a namesake. See related entities.

A New Entity Salience Task with Millions of Training Examples

Authors Jesse Dunietz, Daniel Gillick · Published EACL 2014 (2014) · ACL Anthology Research paper

What it describes: A task and dataset for deciding which entities are central to a news article, not just mentioned, with a model that beats simple position and frequency baselines.

Why it matters for SEO: Salience is exposed publicly in Google Cloud’s Natural Language API.[2] Make your main entity clearly central: name it early, make it the subject of sentences, and keep side entities secondary. See entity salience.

What should you do with this?

Frequently asked questions

Did these patents create the Knowledge Graph?

Not provably. They describe fact repositories, entity typing and disambiguation, and some predate the 2012 launch, but a patent does not show that a system was built or used.

Which paper is most useful for entity SEO?

For daily work, the entity salience paper (Dunietz and Gillick, 2014) and the disambiguation papers are the most practical, because they explain how a page's main entity and its related entities are identified.

Is Knowledge Vault the Knowledge Graph?

No. Knowledge Vault was a Google research system described in a 2014 KDD paper. Google has not said it is the production Knowledge Graph.

References

Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.

  1. ^ "Introducing the Knowledge Graph: things, not strings (Amit Singhal)". Google (The Keyword). Published May 16, 2012.
  2. ^ "Analyzing entities". Google Cloud.
  3. ^ "Deeper understanding with Metaweb". Official Google Blog. Published July 16, 2010.
  4. ^ "Freebase data dumps (Freebase API has been shut down)". Google for Developers.