What is BM25 Full-Text Search?

BM25 ranks full-text search results by combining term frequency, term rarity, and document length into a relevance score.

BM25 (Best Match 25) ranks documents by how well their indexed terms match a search query. It combines term frequency, term rarity, and document length into a relevance score. Search engines use that score to order matching results.

A search for kubernetes deployment error can match a troubleshooting guide, a release note, and a long discussion thread. Returning every match does not help the reader choose. BM25 estimates which documents focus most closely on the query terms.

BM25 is a ranking function, not a complete search engine. Tokenizers define searchable terms, inverted indexes find candidates, and query rules determine which candidates qualify. This distinction matters when diagnosing poor results: changing the scoring formula cannot recover a term that indexing discarded.

When Is BM25 a Good Starting Point?

BM25 is useful when the wording itself carries meaning. Examples include an API method, a legal clause name, a product model, or an incident identifier. A relevant document often repeats those terms because it describes the requested subject directly.

It also creates an interpretable baseline for a retrieval system. Engineers can inspect matched terms, their frequencies, and the effect of document length. This makes it easier to explain why a release note outranks a troubleshooting guide and which change could improve the ordering.

The ranker does not require an embedding model or a training pipeline. Index construction still consumes compute and storage, and maintaining relevance still requires evaluation. That combination makes BM25 a useful first measurement before adding more retrieval stages.

For example, establish whether keyword retrieval finds known support articles before introducing a semantic model. If exact identifiers already work well, preserve those results during later changes. If conceptual questions fail, the evaluation identifies where semantic retrieval could add value.

How Full-Text Search Finds Documents

Analyze text before indexing

An analyzer converts text into searchable tokens. Depending on its configuration, it can lowercase text, normalize accents, remove common words, or apply stemming. Indexing and query analysis must produce compatible tokens.

For example, an English stemmer can reduce running and runs to run. Stemming applies language-specific rules; it does not reliably resolve every irregular form or synonym. Elastic's stemmer documentation describes available language algorithms.

Technical identifiers need particular care. Splitting ERR_CONNECTION_RESET or customer-id changes what an exact match means. A product catalog often needs separate fields for analyzed descriptions and unchanged product codes. Inspect the actual tokens before adjusting relevance parameters.

Look up candidates in an inverted index

An inverted index maps each token to the documents containing it. Each posting list records document identifiers and term frequencies. Indexes can also store token positions for phrase queries.

kubernetes -> document_3, document_17, document_42
deployment -> document_3, document_8, document_17

An AND query requires both terms. An OR query accepts either term. BM25 scores the resulting candidates; the query parser decides these matching rules.

Query text Analyze terms Read posting lists Select matching documents Compute BM25 scores Return ranked results

The index avoids reading every document's raw text for each request. Query cost still depends on term frequency across the collection, filters, candidate counts, and index layout. A common term can produce a much larger candidate set than a rare identifier.

How BM25 Scores Documents

Term frequency and saturation

Term frequency counts a query term's occurrences within a document. Repeated mentions can indicate that the document focuses on that subject. However, repeating the same word twenty times should not make a document twenty times more useful.

BM25 applies saturation: each additional occurrence contributes less than the previous occurrence. The k1 parameter controls how quickly that contribution approaches its limit. Smaller values reduce the importance of repeated mentions. At k1 = 0, matching terms contribute independently of their frequency.

Inverse document frequency

Inverse document frequency, or IDF, measures how rare a term is across the indexed collection. A term found in ten documents carries more distinguishing information than one found in nearly every document.

Rarity depends on the collection. In a Kubernetes knowledge base, kubernetes can be common and contribute little. A specific error identifier can carry much more weight. Adding documents changes these statistics, so the same query can receive different scores after an index refresh.

Document length normalization

Long documents have more opportunities to mention a query term. BM25 adjusts term frequency using document length relative to the collection average. Length usually means analyzed token count, not bytes or characters.

The b parameter controls this adjustment. A value of zero disables length normalization; one applies the full length adjustment in the formula. This does not mean that short documents always win. A longer document can still rank higher through stronger matches across the query.

The formula and its parameters

The following common form sums a contribution for each query term. The Stanford information retrieval text explains BM25's frequency and length adjustments.

score(D, Q) = sum over query terms t of:

             IDF(t) * tf(t, D) * (k1 + 1)
             -----------------------------------------------
             tf(t, D) + k1 * (1 - b + b * length(D) / avgdl)

IDF(t) = ln(1 + (N - df(t) + 0.5) / (df(t) + 0.5))
SymbolMeaning
tf(t, D)Occurrences of term t in document D
NNumber of documents in the collection
df(t)Number of documents containing t
length(D)Analyzed length of the document
avgdlAverage analyzed document length
k1Term frequency saturation parameter
bDocument length normalization parameter

This example uses the positive IDF variant documented by Lucene's BM25 implementation. Lucene defaults to k1 = 1.2 and b = 0.75. Other implementations can differ in constants, field statistics, and length encoding. Their raw scores need not match this calculation exactly.

A Worked BM25 Example

Consider a one-term query for kubernetes over 1,000 documents. Ten documents contain the term, and the average document length is 100 tokens. Use k1 = 1.2 and b = 0.75.

The formula gives an IDF of approximately 4.56. Three matching documents receive the following illustrative scores.

DocumentLength in tokensTerm occurrencesBM25 score
Short troubleshooting note5038.02
Longer overview20035.90
Detailed deployment guide200108.29

The short note outranks the longer overview because both contain three occurrences, but the note concentrates them in less text. The detailed guide ranks slightly higher than the note because it contains more matching evidence. Saturation prevents its ten occurrences from producing a proportional score increase.

For a query containing several terms, repeat the calculation for each term and add the contributions. A rare error code can change the ordering substantially. The example isolates one term so the effects of frequency and length remain visible.

These scores are relative ranking signals, not confidence percentages. A score of 8.29 does not mean an 82.9% chance of relevance. Avoid applying one score threshold across unrelated queries or collections without testing it.

TF-IDF combines term frequency with inverse document frequency. It describes a family of weighting methods, not one fixed scoring formula. Implementations can use logarithmic frequency scaling and normalize document vectors.

BM25 adds an explicit saturation curve and a tunable adjustment against average document length. Saying that TF-IDF never normalizes length is therefore misleading. The useful comparison is between particular implementations on the same retrieval task.

ApproachMain signalUseful starting pointCommon limitation
TF-IDFWeighted lexical overlapText features and simple retrieval baselinesBehavior depends on weighting and normalization
BM25Lexical overlap with saturation and length adjustmentDocumentation, names, identifiers, and technical queriesVocabulary mismatch
Vector searchSimilarity between learned embeddingsParaphrases and conceptual questionsExact identifiers can be difficult
Hybrid searchCombined lexical and vector candidatesCollections containing both query stylesMore components to evaluate and operate

BM25 cannot infer that account termination answers a query about canceling a subscription without matching terms or additional query processing. Synonyms and query expansion can improve lexical recall. Hybrid search adds semantic candidates from vector retrieval.

Reciprocal Rank Fusion combines result positions rather than raw scores. It therefore avoids requiring BM25 scores and vector similarities to share a numerical scale. Weighted score addition requires a separate calibration strategy. Neither method guarantees better results for every collection.

How to Tune and Evaluate BM25

Start with representative queries and documents labeled for relevance. Include product names, error identifiers, natural-language questions, ambiguous terms, and queries that should return nothing. Separate tuning queries from a held-out evaluation set.

Inspect failed queries before changing k1 or b. Missing results often come from tokenization, incorrect language selection, restrictive filters, or outdated indexes. Ranking changes cannot fix missing documents or access rules that exclude them.

Measure different parts of retrieval quality. Recall at ten measures how many relevant documents appear in the first ten results. Precision at ten measures the fraction of those results that are relevant. Normalized discounted cumulative gain, or nDCG, also considers result order and graded relevance.

Change one setting at a time and keep the corpus fixed during comparisons. Test lower b values when long, useful documents receive excessive penalties. Test k1 changes when repeated mentions receive too much or too little weight. Retain a setting only when it improves the chosen evaluation criteria.

Query-level inspection matters alongside aggregate scores. A small average improvement can hide regressions for critical identifiers or a particular language. Group results by query type and document source. For an internal help center, losing the exact incident runbook can matter more than improving several broad informational queries.

Measure latency, indexing time, and storage alongside relevance. Adding phrase information, more fields, or broader query expansion changes operational cost. For an application endpoint, test concurrent searches during ingestion rather than measuring an idle index alone.

Where Full-Text Ranking Runs

Full-text search does not imply BM25. PostgreSQL's built-in ranking functions, ts_rank and ts_rank_cd, use different ranking methods. A PostgreSQL deployment needs an appropriate extension or separate retrieval component when BM25 is specifically required.

Lucene-based systems, including Elasticsearch, use BM25 by default. Elasticsearch exposes similarity settings for tuning its parameters. Embedded search libraries can also implement BM25 without requiring a separate search service.

Choose the deployment based on index size, update rate, query features, and operational ownership. A separate service creates a synchronization boundary with the source database. An embedded index moves that responsibility into the application or data runtime. Both need explicit recovery and freshness policies.

Advanced Topics

Field weighting and phrase matching

A title match often deserves more weight than a body match. Independent field boosts multiply scores from separate fields. BM25F takes a different approach: it combines weighted, length-normalized field frequencies before applying saturation.

These methods are not interchangeable. Repeating text across a title, summary, and body can affect their rankings differently. Test field weighting with realistic documents instead of assuming a threefold title boost universally improves results.

BM25 itself does not enforce word order. A phrase query uses indexed positions to require adjacent terms or a permitted distance between them. Treat phrase matching, Boolean operators, and typo tolerance as query features surrounding the ranker.

Chunking changes the statistics

Retrieval systems often split long documents into passages. Each passage then becomes a searchable unit with its own length and term counts. Splitting also changes average length and document frequency across the index.

A parameter choice that works for complete manuals can behave differently for short passages. Reevaluate relevance after changing chunk size, overlap, or boundary rules. Deduplicate overlapping passages before returning context to an answer generator.

Preserve parent identifiers and section titles during chunking. Otherwise, several nearly identical passages can occupy the first results without adding useful information. Evaluate both passage relevance and coverage of distinct source documents.

Shards, filters, and candidate limits

Distributed indexes can calculate term statistics per shard. Uneven term distributions can then affect ranking across shards. Engines differ in whether they collect broader statistics before scoring, with extra coordination potentially increasing latency.

Filtering order also matters. Selecting ten global matches and then applying a tenant filter can return fewer than ten authorized results. Filtering during candidate retrieval searches the correct population before selecting the final results.

Access control must apply regardless of ranking or retrieval method. A high relevance score does not authorize a document. Test restrictive filters, empty result sets, and ranking during concurrent updates before exposing search to multiple tenants.

BM25 Full-Text Search with Spice

Spice hybrid SQL search combines lexical retrieval with SQL queries and vector search. Its built-in full-text engine uses Tantivy. The full-text search documentation defines index configuration and the text_search() SQL function.

This example requires a local Parquet file with unique id values and a text body column. Add the dataset to a Spicepod.

datasets:
  - from: file:./knowledge_base.parquet
    name: knowledge_base
    acceleration:
      enabled: true
    columns:
      - name: body
        full_text_search:
          enabled: true
          row_id:
            - id

After the dataset and index load, query the indexed column.

SELECT id, body, _score AS score
FROM text_search(knowledge_base, 'kubernetes deployment error', body, 10)
ORDER BY score DESC
LIMIT 10;

The table and column arguments are identifiers. The search text is a string, and the fourth argument bounds the candidates returned by the function. The query aliases Spice's _score column to score.

For application search, choose refresh behavior according to the source connector and required freshness. Available data integrations have different change-capture capabilities. Indexing a dataset does not automatically make every source update visible immediately.

BM25 Full-Text Search FAQ

What does BM25 stand for?

BM25 stands for Best Match 25, a ranking function associated with the Okapi information retrieval system. It scores lexical matches using term frequency, term rarity, and document length. It does not require an embedding model.

What are the default BM25 parameters?

Lucene uses k1 = 1.2 and b = 0.75 by default. The k1 parameter controls term frequency saturation, and b controls document length normalization. Defaults and tuning options differ between engines.

How does BM25 differ from TF-IDF?

BM25 uses explicit term frequency saturation and a tunable adjustment against average document length. TF-IDF describes several weighting methods, some of which also normalize document vectors. Compare specific implementations against labeled relevance examples.

When should I use BM25 vs. vector search?

BM25 fits queries where exact terminology, names, or identifiers matter. Vector search can retrieve conceptual matches that use different wording. Evaluate hybrid retrieval when the collection needs both query styles.

Can BM25 handle multi-language search?

BM25 can rank text in multiple languages, but retrieval quality depends on the analyzer. Each language needs appropriate tokenization and normalization. Verify the capabilities of the selected engine before assuming that it stems or segments every language correctly.

Are BM25 scores comparable across different queries?

Raw BM25 scores are not calibrated across different queries or collections. Term statistics, query length, and engine details affect their scale. Use them to rank candidates within a search, and validate any cross-query threshold separately.

See Spice in action

Walk through your use case with an engineer and see how Spice handles federation, acceleration, and AI integration for production workloads.

Talk to an engineer