What is BM25 Full-Text Search?
BM25 ranks full-text search results by combining term frequency, term rarity, and document length into a relevance score.
BM25 (Best Match 25) ranks documents by how well their indexed terms match a search query. It combines term frequency, term rarity, and document length into a relevance score. Search engines use that score to order matching results.
A search for kubernetes deployment error can match a troubleshooting guide, a release note, and a long discussion thread. Returning every match does not help the reader choose. BM25 estimates which documents focus most closely on the query terms.
BM25 is a ranking function, not a complete search engine. Tokenizers define searchable terms, inverted indexes find candidates, and query rules determine which candidates qualify. This distinction matters when diagnosing poor results: changing the scoring formula cannot recover a term that indexing discarded.
When Is BM25 a Good Starting Point?
BM25 is useful when the wording itself carries meaning. Examples include an API method, a legal clause name, a product model, or an incident identifier. A relevant document often repeats those terms because it describes the requested subject directly.
It also creates an interpretable baseline for a retrieval system. Engineers can inspect matched terms, their frequencies, and the effect of document length. This makes it easier to explain why a release note outranks a troubleshooting guide and which change could improve the ordering.
The ranker does not require an embedding model or a training pipeline. Index construction still consumes compute and storage, and maintaining relevance still requires evaluation. That combination makes BM25 a useful first measurement before adding more retrieval stages.
For example, establish whether keyword retrieval finds known support articles before introducing a semantic model. If exact identifiers already work well, preserve those results during later changes. If conceptual questions fail, the evaluation identifies where semantic retrieval could add value.
How Full-Text Search Finds Documents
Analyze text before indexing
An analyzer converts text into searchable tokens. Depending on its configuration, it can lowercase text, normalize accents, remove common words, or apply stemming. Indexing and query analysis must produce compatible tokens.
For example, an English stemmer can reduce running and runs to run. Stemming applies language-specific rules; it does not reliably resolve every irregular form or synonym. Elastic's stemmer documentation describes available language algorithms.
Technical identifiers need particular care. Splitting ERR_CONNECTION_RESET or customer-id changes what an exact match means. A product catalog often needs separate fields for analyzed descriptions and unchanged product codes. Inspect the actual tokens before adjusting relevance parameters.
Look up candidates in an inverted index
An inverted index maps each token to the documents containing it. Each posting list records document identifiers and term frequencies. Indexes can also store token positions for phrase queries.
kubernetes -> document_3, document_17, document_42
deployment -> document_3, document_8, document_17An AND query requires both terms. An OR query accepts either term. BM25 scores the resulting candidates; the query parser decides these matching rules.
The index avoids reading every document's raw text for each request. Query cost still depends on term frequency across the collection, filters, candidate counts, and index layout. A common term can produce a much larger candidate set than a rare identifier.
How BM25 Scores Documents
Term frequency and saturation
Term frequency counts a query term's occurrences within a document. Repeated mentions can indicate that the document focuses on that subject. However, repeating the same word twenty times should not make a document twenty times more useful.
BM25 applies saturation: each additional occurrence contributes less than the previous occurrence. The k1 parameter controls how quickly that contribution approaches its limit. Smaller values reduce the importance of repeated mentions. At k1 = 0, matching terms contribute independently of their frequency.
Inverse document frequency
Inverse document frequency, or IDF, measures how rare a term is across the indexed collection. A term found in ten documents carries more distinguishing information than one found in nearly every document.
Rarity depends on the collection. In a Kubernetes knowledge base, kubernetes can be common and contribute little. A specific error identifier can carry much more weight. Adding documents changes these statistics, so the same query can receive different scores after an index refresh.
Document length normalization
Long documents have more opportunities to mention a query term. BM25 adjusts term frequency using document length relative to the collection average. Length usually means analyzed token count, not bytes or characters.
The b parameter controls this adjustment. A value of zero disables length normalization; one applies the full length adjustment in the formula. This does not mean that short documents always win. A longer document can still rank higher through stronger matches across the query.
The formula and its parameters
The following common form sums a contribution for each query term. The Stanford information retrieval text explains BM25's frequency and length adjustments.
score(D, Q) = sum over query terms t of:
IDF(t) * tf(t, D) * (k1 + 1)
-----------------------------------------------
tf(t, D) + k1 * (1 - b + b * length(D) / avgdl)
IDF(t) = ln(1 + (N - df(t) + 0.5) / (df(t) + 0.5))| Symbol | Meaning |
|---|---|
tf(t, D) | Occurrences of term t in document D |
N | Number of documents in the collection |
df(t) | Number of documents containing t |
length(D) | Analyzed length of the document |
avgdl | Average analyzed document length |
k1 | Term frequency saturation parameter |
b | Document length normalization parameter |
This example uses the positive IDF variant documented by Lucene's BM25 implementation. Lucene defaults to k1 = 1.2 and b = 0.75. Other implementations can differ in constants, field statistics, and length encoding. Their raw scores need not match this calculation exactly.
A Worked BM25 Example
Consider a one-term query for kubernetes over 1,000 documents. Ten documents contain the term, and the average document length is 100 tokens. Use k1 = 1.2 and b = 0.75.
The formula gives an IDF of approximately 4.56. Three matching documents receive the following illustrative scores.
| Document | Length in tokens | Term occurrences | BM25 score |
|---|---|---|---|
| Short troubleshooting note | 50 | 3 | 8.02 |
| Longer overview | 200 | 3 | 5.90 |
| Detailed deployment guide | 200 | 10 | 8.29 |
The short note outranks the longer overview because both contain three occurrences, but the note concentrates them in less text. The detailed guide ranks slightly higher than the note because it contains more matching evidence. Saturation prevents its ten occurrences from producing a proportional score increase.
For a query containing several terms, repeat the calculation for each term and add the contributions. A rare error code can change the ordering substantially. The example isolates one term so the effects of frequency and length remain visible.
These scores are relative ranking signals, not confidence percentages. A score of 8.29 does not mean an 82.9% chance of relevance. Avoid applying one score threshold across unrelated queries or collections without testing it.
BM25 vs. TF-IDF and Vector Search
TF-IDF combines term frequency with inverse document frequency. It describes a family of weighting methods, not one fixed scoring formula. Implementations can use logarithmic frequency scaling and normalize document vectors.
BM25 adds an explicit saturation curve and a tunable adjustment against average document length. Saying that TF-IDF never normalizes length is therefore misleading. The useful comparison is between particular implementations on the same retrieval task.
| Approach | Main signal | Useful starting point | Common limitation |
|---|---|---|---|
| TF-IDF | Weighted lexical overlap | Text features and simple retrieval baselines | Behavior depends on weighting and normalization |
| BM25 | Lexical overlap with saturation and length adjustment | Documentation, names, identifiers, and technical queries | Vocabulary mismatch |
| Vector search | Similarity between learned embeddings | Paraphrases and conceptual questions | Exact identifiers can be difficult |
| Hybrid search | Combined lexical and vector candidates | Collections containing both query styles | More components to evaluate and operate |
BM25 cannot infer that account termination answers a query about canceling a subscription without matching terms or additional query processing. Synonyms and query expansion can improve lexical recall. Hybrid search adds semantic candidates from vector retrieval.
Reciprocal Rank Fusion combines result positions rather than raw scores. It therefore avoids requiring BM25 scores and vector similarities to share a numerical scale. Weighted score addition requires a separate calibration strategy. Neither method guarantees better results for every collection.
How to Tune and Evaluate BM25
Start with representative queries and documents labeled for relevance. Include product names, error identifiers, natural-language questions, ambiguous terms, and queries that should return nothing. Separate tuning queries from a held-out evaluation set.
Inspect failed queries before changing k1 or b. Missing results often come from tokenization, incorrect language selection, restrictive filters, or outdated indexes. Ranking changes cannot fix missing documents or access rules that exclude them.
Measure different parts of retrieval quality. Recall at ten measures how many relevant documents appear in the first ten results. Precision at ten measures the fraction of those results that are relevant. Normalized discounted cumulative gain, or nDCG, also considers result order and graded relevance.
Change one setting at a time and keep the corpus fixed during comparisons. Test lower b values when long, useful documents receive excessive penalties. Test k1 changes when repeated mentions receive too much or too little weight. Retain a setting only when it improves the chosen evaluation criteria.
Query-level inspection matters alongside aggregate scores. A small average improvement can hide regressions for critical identifiers or a particular language. Group results by query type and document source. For an internal help center, losing the exact incident runbook can matter more than improving several broad informational queries.
Measure latency, indexing time, and storage alongside relevance. Adding phrase information, more fields, or broader query expansion changes operational cost. For an application endpoint, test concurrent searches during ingestion rather than measuring an idle index alone.
Where Full-Text Ranking Runs
Full-text search does not imply BM25. PostgreSQL's built-in ranking functions, ts_rank and ts_rank_cd, use different ranking methods. A PostgreSQL deployment needs an appropriate extension or separate retrieval component when BM25 is specifically required.
Lucene-based systems, including Elasticsearch, use BM25 by default. Elasticsearch exposes similarity settings for tuning its parameters. Embedded search libraries can also implement BM25 without requiring a separate search service.
Choose the deployment based on index size, update rate, query features, and operational ownership. A separate service creates a synchronization boundary with the source database. An embedded index moves that responsibility into the application or data runtime. Both need explicit recovery and freshness policies.
Advanced Topics
Field weighting and phrase matching
A title match often deserves more weight than a body match. Independent field boosts multiply scores from separate fields. BM25F takes a different approach: it combines weighted, length-normalized field frequencies before applying saturation.
These methods are not interchangeable. Repeating text across a title, summary, and body can affect their rankings differently. Test field weighting with realistic documents instead of assuming a threefold title boost universally improves results.
BM25 itself does not enforce word order. A phrase query uses indexed positions to require adjacent terms or a permitted distance between them. Treat phrase matching, Boolean operators, and typo tolerance as query features surrounding the ranker.
Chunking changes the statistics
Retrieval systems often split long documents into passages. Each passage then becomes a searchable unit with its own length and term counts. Splitting also changes average length and document frequency across the index.
A parameter choice that works for complete manuals can behave differently for short passages. Reevaluate relevance after changing chunk size, overlap, or boundary rules. Deduplicate overlapping passages before returning context to an answer generator.
Preserve parent identifiers and section titles during chunking. Otherwise, several nearly identical passages can occupy the first results without adding useful information. Evaluate both passage relevance and coverage of distinct source documents.
Shards, filters, and candidate limits
Distributed indexes can calculate term statistics per shard. Uneven term distributions can then affect ranking across shards. Engines differ in whether they collect broader statistics before scoring, with extra coordination potentially increasing latency.
Filtering order also matters. Selecting ten global matches and then applying a tenant filter can return fewer than ten authorized results. Filtering during candidate retrieval searches the correct population before selecting the final results.
Access control must apply regardless of ranking or retrieval method. A high relevance score does not authorize a document. Test restrictive filters, empty result sets, and ranking during concurrent updates before exposing search to multiple tenants.
BM25 Full-Text Search with Spice
Spice hybrid SQL search combines lexical retrieval with SQL queries and vector search. Its built-in full-text engine uses Tantivy. The full-text search documentation defines index configuration and the text_search() SQL function.
This example requires a local Parquet file with unique id values and a text body column. Add the dataset to a Spicepod.
datasets:
- from: file:./knowledge_base.parquet
name: knowledge_base
acceleration:
enabled: true
columns:
- name: body
full_text_search:
enabled: true
row_id:
- idAfter the dataset and index load, query the indexed column.
SELECT id, body, _score AS score
FROM text_search(knowledge_base, 'kubernetes deployment error', body, 10)
ORDER BY score DESC
LIMIT 10;The table and column arguments are identifiers. The search text is a string, and the fourth argument bounds the candidates returned by the function. The query aliases Spice's _score column to score.
For application search, choose refresh behavior according to the source connector and required freshness. Available data integrations have different change-capture capabilities. Indexing a dataset does not automatically make every source update visible immediately.
BM25 Full-Text Search FAQ
What does BM25 stand for?
BM25 stands for Best Match 25, a ranking function associated with the Okapi information retrieval system. It scores lexical matches using term frequency, term rarity, and document length. It does not require an embedding model.
What are the default BM25 parameters?
Lucene uses k1 = 1.2 and b = 0.75 by default. The k1 parameter controls term frequency saturation, and b controls document length normalization. Defaults and tuning options differ between engines.
How does BM25 differ from TF-IDF?
BM25 uses explicit term frequency saturation and a tunable adjustment against average document length. TF-IDF describes several weighting methods, some of which also normalize document vectors. Compare specific implementations against labeled relevance examples.
When should I use BM25 vs. vector search?
BM25 fits queries where exact terminology, names, or identifiers matter. Vector search can retrieve conceptual matches that use different wording. Evaluate hybrid retrieval when the collection needs both query styles.
Can BM25 handle multi-language search?
BM25 can rank text in multiple languages, but retrieval quality depends on the analyzer. Each language needs appropriate tokenization and normalization. Verify the capabilities of the selected engine before assuming that it stems or segments every language correctly.
Are BM25 scores comparable across different queries?
Raw BM25 scores are not calibrated across different queries or collections. Term statistics, query length, and engine details affect their scale. Use them to rank candidates within a search, and validate any cross-query threshold separately.
Learn more about full-text search
Technical guides on building full-text and hybrid search with BM25, vector similarity, and SQL in a single runtime.
Search Docs
Learn how Spice provides full-text, semantic, and hybrid search capabilities in a single SQL-native runtime.
True Hybrid Search: Vector, Full-Text, and SQL in One Runtime
Build hybrid search without managing multiple systems. Query vectors, run full-text search, and execute SQL in one unified runtime.
Real-Time Hybrid Search Using RRF: A Hands-On Guide with Spice
Learn how to build hybrid search with RRF directly in SQL using Spice, combining text, vector, and time-based relevance in one query.
See Spice in action
Walk through your use case with an engineer and see how Spice handles federation, acceleration, and AI integration for production workloads.
Talk to an engineer


