Neural IR, BERT, MUM and passage ranking: patents and papers
Neural information retrieval (neural IR) uses deep learning models to understand queries and passages and to match them by meaning. Google has confirmed using BERT, MUM for some tasks, RankBrain, neural matching and passage ranking. This page lists the verified research behind those systems (the Transformer, BERT and T5) plus key retrieval papers and three Google patents.
Which neural systems has Google confirmed?
| System | What Google says |
|---|---|
| BERT | Applied to ranking and featured snippets from 2019; Google said it would help Search “better understand one in 10 searches in the U.S. in English”.[1] |
| MUM | “Uses the T5 text-to-text framework and is 1,000 times more powerful than BERT”[2]; “not currently used for general ranking”.[3] |
| Passage ranking | Identifies individual sections or “passages” of a page to judge relevance.[3] |
| RankBrain, neural matching | Relate words to concepts and match concepts in queries and pages (see query understanding). |
Only the BERT and T5 papers below are directly tied to named systems by Google. The rest show how the field works. Officially documented
Which patents describe answer passages and Transformers?
Scoring candidate answer passages
What it describes: Scoring candidate answer passages for question queries. For each passage from responsive resources, the system computes a query-term match score and an answer-term match score, among other features, and selects an answer passage.
Why it matters for SEO: Passages that restate the question’s terms and contain the terms typical of good answers score well in this design. Put a direct, self-contained answer under each question heading. Practitioner practice
Context scoring adjustments for answer passages
What it describes: Adjusting answer-passage scores using context: a heading vector describes the path through the page’s heading hierarchy from the root heading to the heading the passage sits under.
Why it matters for SEO: Your heading structure is context for every passage. Use a clean H1>H2>H3 hierarchy where headings describe their sections. AEO.wiki covers this in content structure. Practitioner practice
Attention-based sequence transduction neural networks
What it describes: Attention-based sequence transduction neural networks: the encoder-decoder architecture built from attention layers, known as the Transformer.
Why it matters for SEO: The Transformer is the architecture behind BERT, T5 and today’s large language models; this is Google’s patent on it.
Which research papers underpin BERT, MUM and neural retrieval?
Attention Is All You Need
What it describes: The Transformer: a model built entirely on attention, which weighs how every word relates to every other word.
Why it matters for SEO: Understanding words in full context is why modern search handles long, conversational queries.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
What it describes: BERT, a Transformer pre-trained to read text in both directions, then fine-tuned for tasks such as question answering.
Why it matters for SEO: Google applied BERT to Search in 2019.[1] Small words like “to” and “for” now change meaning, so write naturally and precisely. Officially documented
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
What it describes: T5, which casts every language task as text in, text out.
Why it matters for SEO: Google says MUM uses the T5 text-to-text framework.[2] Officially documented
MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
What it describes: A large Microsoft dataset of real Bing queries with human-written answers and relevant passages.
Why it matters for SEO: The standard benchmark for passage ranking; much of the research below was measured on it.
Natural Questions: A Benchmark for Question Answering Research
What it describes: A Google dataset of real Google queries paired with Wikipedia pages annotated with long and short answers.
Why it matters for SEO: Answers are often a paragraph plus a short fact inside it: the same shape as an answer-first paragraph.
Passage Re-ranking with BERT
What it describes: Using BERT to re-rank candidate passages for a query, which sharply improved results on MS MARCO.
Why it matters for SEO: Shows the pattern of cheap retrieval followed by a smarter re-ranking step. Each passage is judged on its own.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
What it describes: A retrieval model that keeps BERT-level quality while being fast enough for large collections.
Why it matters for SEO: Makes passage-level neural retrieval practical at scale.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
What it describes: Producing sentence embeddings that can be compared by similarity, the basis of many semantic search tools.
Why it matters for SEO: Many SEO content tools use embeddings like these to measure topical similarity.
Dense Passage Retrieval for Open-Domain Question Answering
What it describes: Retrieving passages with learned dense vectors instead of word matching, outperforming BM25 on several QA benchmarks.
Why it matters for SEO: Passages can be found by meaning even without shared keywords, so clarity beats keyword stuffing.
Large Dual Encoders Are Generalizable Retrievers
What it describes: Google research showing that scaling up dual-encoder retrievers (GTR) improves how well they work on new domains.
Why it matters for SEO: Dense retrieval keeps improving, which rewards clearly written, self-contained passages.
What should you do with this?
- Write self-contained passages: each section should answer one question on its own.
- Use descriptive, hierarchical headings.
- Write naturally and precisely; modern models read context, not just keywords.
Frequently asked questions
Did BERT replace RankBrain?
No. Google lists both BERT and RankBrain as separate AI systems in its ranking systems guide.
Is MUM used for ranking?
Google's ranking systems guide says MUM is not currently used for general ranking, only for specific applications such as improving featured snippet callouts.
What is passage ranking?
An AI system Google uses to identify individual sections or passages of a page to better understand how relevant the page is to a search.
See also
References
Pages accessed September 29, 2026 unless a date is given. See all sources and our editorial policy.
- ^ "Understanding searches better than ever before (Pandu Nayak)". Google (The Keyword). Published October 25, 2019.
- ^ "MUM: A new AI milestone for understanding information (Pandu Nayak)". Google (The Keyword). Published May 18, 2021.
- ^ "A guide to Google Search ranking systems". Google Search Central.