How Do You Keep Embeddings in Sync When the Source Data Changes?
Keeping embeddings in sync means detecting every insert, update, and delete in the source data, re-embedding only the chunks that changed, and removing vectors for content that no longer exists.
An app edits a document, so the vector index holds stale vectors and a search returns the old text. A sync pipeline detects the change, re-embeds only the chunk whose hash changed, and replaces its vector, so the same search returns the current text.
An embedding records the meaning of a piece of text at the moment a model processed it. The source text keeps changing after that moment. Writers edit documents, applications update rows, and retention jobs delete records. The vector index does not see any of these changes unless a process tells it.
When the index and the source disagree, retrieval fails without an error. A support assistant quotes a refund policy that changed last week. A search returns a document that its owner deleted. Each response looks normal, so users often find the problem before monitoring does.
This guide covers how embeddings go stale, how a sync pipeline detects and applies changes, and how six detection methods compare.
Why Embeddings Go Stale
An embedding is derived data, like a cache entry or a materialized view. Derived data stays correct only until its input changes. After that, a process must recompute it, or readers get an old answer.
Source data changes in four ways, and each one leaves a different defect in the index:
| Change in the source | What the index holds if the change is missed | Required action |
|---|---|---|
| Insert | No vectors for the new content | Chunk, embed, and write the new record |
| Update | Vectors for the previous text | Re-embed the changed chunks and replace the old ones |
| Delete | Vectors for content that no longer exists | Remove every vector for the deleted record |
| Model or chunker change | Vectors built with a model or chunk layout that no longer applies | Re-embed the full corpus into a new index |
Deletes cause the most serious failures. A missed insert makes an answer incomplete. A missed delete makes the system return content that its owner removed, sometimes for legal or privacy reasons.
The Scheduled Re-Embedding Job
Many teams start with a scheduled job that re-embeds the whole corpus every night. The approach is simple, and it fails in four predictable ways:
- Staleness of up to one interval. An edit made just after the job runs waits almost a full day to reach the index.
- Cost that grows with the corpus. The job pays to embed every document, even when only a few documents changed.
- Missed deletes. A job that only upserts never removes the vectors of records that left the source.
- Partial rebuilds. A long job can fail halfway and leave a mix of old and new vectors in the index.
A full rebuild is still the correct choice in two cases: a small corpus that changes rarely, and any change to the embedding model.
How an Embedding Sync Pipeline Works
A sync pipeline replaces the full rebuild with three stages. It detects what changed, re-embeds only what changed, and applies the result to the index.
Stage 1: Detect the Change
The detector emits a stream of record keys, each with an operation: insert, update, or delete. The detection method sets how fast changes arrive and which changes the pipeline can see.
Stage 2: Re-Embed Only What Changed
Embedding calls cost time and money, so the pipeline skips text that did not change. The common method stores a content hash next to each chunk vector. When a record changes, the pipeline chunks it again, hashes each chunk, and compares the new hashes with the stored ones. Only chunks with a new hash go to the embedding model.
Hash the normalized text, not the raw bytes. Normalization removes whitespace and markup changes that do not change the meaning.
Keep fields that change often out of the embedded text. Prices, stock levels, and status flags belong in filterable metadata columns. A price change then updates one column and needs no new vector.
Stage 3: Apply the Result to the Index
The apply step writes the new vectors and removes the vectors that the change made obsolete. Key every vector by the source record ID and the chunk position, so each write replaces the correct entry. Make each write idempotent, because most change streams can repeat an event after a restart.
Order also matters. Apply events in source commit order, or compare a version number before each write. Otherwise, a delayed event for an old version can overwrite the vector for a newer one.
Change Detection Methods Compared
Six methods cover most systems. They differ in latency, in cost, and in which writes they can see.
Scheduled Full Re-Embedding
A job reads every record on a fixed schedule and embeds it again. It needs no change tracking. Its staleness equals the schedule interval, and its cost grows with the size of the corpus.
Timestamp Polling
A job queries for records with an updated_at value later than the last stored watermark. It re-embeds those records and moves the watermark forward. The method is cheap, but it has two gaps.
First, it cannot see hard deletes, because a deleted row no longer exists to match the query. Second, it depends on every writer setting updated_at, including bulk scripts and manual fixes. A long transaction can also commit a row with a timestamp older than the watermark. Re-read a short overlap window on each poll to catch those rows.
Content Hashing
A job scans the source, hashes each record or file, and compares each hash with the stored value. A complete scan also reveals deletes: an index key that the scan did not return belongs to a deleted record. A failed or partial scan proves nothing about deletes, so the job must skip deletion after one. The scan still reads the full source on every run, so hashing cuts embedding cost but not read cost.
Application Hooks
The application re-embeds a record in the same code path that saves it. Latency is low, and the method needs no extra infrastructure. The method misses every write that bypasses the application, such as migrations, bulk imports, admin scripts, and other services. It also creates a dual write: if the save succeeds and the embedding call fails, the source and the index disagree.
Database Triggers
A trigger runs inside the database on every insert, update, and delete, including writes from outside the application. One design sets the embedding column to NULL when the text changes, and a worker re-embeds every row with a NULL embedding. Another design writes the record key to an outbox table that a worker reads.
PostgreSQL LISTEN and NOTIFY can wake the worker without polling. Notifications are not durable, so a worker that is offline misses them. Keep the outbox table as the record of pending work, and use the notification only as a wake-up signal. Triggers also add work to every write transaction, which matters on tables with high write rates.
Change Data Capture
Change data capture (CDC) reads the database transaction log, such as the PostgreSQL write-ahead log or the MySQL binary log. It emits every committed insert, update, and delete in commit order. It sees writes from every client, and the application needs no changes.
CDC has its own operational cost. The database must run with logical replication enabled, and each consumer tracks a position in the log. In PostgreSQL, a stalled consumer makes the database retain log files, which can fill the disk. Monitor consumer lag and replication slot size.
Comparing the Methods
| Method | Sees hard deletes | Sees writes from outside the app | Typical staleness | Cost grows with |
|---|---|---|---|---|
| Scheduled full re-embedding | Only if it builds a new index | Yes | One schedule interval | Corpus size |
| Timestamp polling | No | Only if every writer sets the timestamp | One poll interval | Changed rows plus the overlap |
| Content hashing | Yes, by comparing key sets | Yes | One scan interval | Corpus size for reads |
| Application hooks | Only if the delete path calls it | No | Seconds | Changed rows |
| Database triggers | Yes | Yes | Seconds to one worker cycle | Changed rows plus write overhead |
| Change data capture | Yes | Yes | Seconds | Changed rows |
Many production systems combine two methods. CDC or triggers carry the normal flow of changes, and a scheduled reconciliation job finds anything the stream missed.
Chunk-Level Synchronization
Most indexes store several chunks per document. Two problems appear at the chunk level.
Chunk Boundaries Shift After an Edit
Fixed-size chunking splits text every N tokens. An insertion near the top of a document moves every later boundary. Every later chunk then holds different text and gets a new hash, although most of the document did not change. A one-sentence edit can re-embed the whole document.
Structure-aware chunking limits this effect. Split the document at headings or sections first, then split long sections into fixed-size chunks. An edit then changes only the chunks of one section. Content-defined chunking goes further: it places boundaries based on the text itself, so boundaries away from the edit stay where they were.
For short documents, re-embedding the whole document on every change is often simpler. The extra embedding cost is small, and the pipeline has no partial state to track.
Replace a Document's Chunks as a Group
Upserting chunks by position leaves a defect when a document gets shorter. A document that had six chunks and now has four still has chunks five and six in the index. Those chunks hold text that no longer exists, and retrieval can still return them.
Two patterns prevent these orphaned chunks. The first deletes all chunks for the document and inserts the new set in one transaction. The second writes the new chunks under a new version number, then switches the document's current version in one write. Queries filter retrieved chunks to the current version, so a reader never mixes old and new text. A cleanup step deletes the older versions later.
Handling Deletes
Deletes need their own design, because several detection methods cannot see them.
A soft delete marks the row with a deleted_at value and keeps it in the table. Timestamp polling sees this change like any other update, and the pipeline removes the vectors. Filter soft-deleted rows at query time until the pipeline removes them. Purge the rows later, after the index confirms the removal.
A hard delete removes the row. Only CDC, triggers, a full key comparison, or an application hook on the delete path can detect it. A reconciliation job lists the record keys in the index and in the source. It then deletes every index entry whose key the source no longer has. Run it on a schedule even when CDC is in place, because it catches deletes that occurred while a consumer was down.
A reconciliation job deletes data, so it needs guards. Read the source keys from one consistent snapshot, and skip index entries written after that snapshot started. Abort the run if the read fails or stops early, or if it would delete an unusual share of the index.
Treat a missed delete as a data retention defect, not only a relevance defect. A record deleted for privacy or legal reasons must leave every derived copy, and the vector index is a derived copy.
Keeping Vector and Keyword Indexes Consistent
Hybrid search queries two indexes: a vector index and a full-text index. Both indexes need the same inserts, updates, and deletes. If one index applies a delete and the other does not, the deleted document returns through the second index.
Three practices keep the two indexes aligned:
- One change stream for both. Drive both indexes from the same ordered stream, keyed by the same primary key. Separate pipelines with separate checkpoints can apply one change at different times.
- One transaction where possible. When both indexes live in the same database, update both in the same transaction.
- A final check against the source. Join the retrieved keys to the current source rows, and drop any key with no current row.
Choosing a Strategy
The right design depends on the source, the change rate, and how much staleness the application accepts:
- Database tables that users edit. Use CDC from the transaction log. It captures every update and delete without changes to the application.
- Append-only records such as tickets, logs, or messages. Use timestamp polling with an overlap window. Add a retention rule if old records leave the source.
- Files in object storage or a document system. Use modification times or object events to find changed files. Compare content hashes to skip files that did not change.
- Small corpora that change rarely. A scheduled full rebuild into a new index is simple and correct.
- Small candidate sets after a selective filter. Embed the text at query time. The system stores no vectors, so nothing goes stale, but every query pays the embedding cost.
Advanced Topics
Changing the Embedding Model or the Chunker
Vectors from different models live in different vector spaces, so a model change makes every stored vector incompatible with new query vectors. A chunker change also calls for a rebuild, because it changes the boundaries and keys of every chunk. Treat either change as a full rebuild into a new index.
Build the new index next to the old one. Record the current position in the change stream, then backfill the new index from a fresh snapshot. Replay changes from the recorded position into the new index, so each newer change replaces the snapshot copy of its record. Switch queries after the new index catches up with the stream, then drop the old index. Store the model name and chunker version with each vector, so a mixed index is easy to detect.
Measuring Freshness
Freshness is the time from a source commit to the moment a query can retrieve the new content. A canary measures it end to end: it writes a record with a unique token, then queries search until the token appears. Alert on the 95th percentile of that delay.
An acknowledged write is not always a visible write. Many search engines expose new documents to readers only after a refresh step. Elasticsearch, for example, refreshes active indexes every second by default. Also track two counts: rows that wait for embedding, and the gap between source rows and distinct source IDs in the index. Compare distinct IDs, because one row can own several chunk entries. In-flight changes make the counts differ briefly, so alert only on a gap that persists across checks.
Debouncing and Batching
An editor with autosave can produce dozens of updates for one document in a minute. Wait for a short quiet period after each change, then embed only the latest version. Batch embedding calls to cut per-request overhead and to stay under provider rate limits. Cache embeddings by content hash, so a reverted edit or repeated boilerplate text needs no new model call. Include the model name, model version, and embedding parameters in the cache key, so a model change cannot reuse old vectors.
Guarding Against Stale Reads
Even a fast pipeline has a delay between a commit and the index update. For data where an old answer causes harm, verify retrieved chunks at query time. Store the source version or content hash with each chunk. After retrieval, compare it with the current source row, and drop chunks that no longer match. The answer is then incomplete for a short time, which is safer than an answer built on deleted text.
Keeping Embeddings in Sync with Spice
Spice treats embeddings as columns of a dataset. A Spicepod declares which column to embed, which model to use, and how to chunk the text. When the runtime loads or refreshes an accelerated dataset, it computes the embedding column. It stores the vectors in the same row as the source text.
The refresh mode of each dataset selects one of the change detection methods in this guide.
fullreads the source again on each refresh and swaps in the new copy. It re-embeds every row, so it suits small corpora.appendwith atime_columnloads only rows newer than the stored maximum. With aprimary_keyandon_conflict: upsert, an updated row replaces its earlier version and its vectors.changesapplies real-time change data capture from PostgreSQL logical replication, MySQL binlog replication, DynamoDB Streams, MongoDB change streams, or Debezium. Spice embeds each change batch before it writes the batch.
In the accelerated table, chunk vectors and chunk offsets are list columns of the source row. An update rewrites the row with a new chunk list, so chunks from a longer version cannot remain. A delete removes the row and its vectors in the same write. A full-text index on the same column updates from the same writes, keyed by the same primary key.
This Spicepod embeds a PostgreSQL documents table and keeps it current from the write-ahead log. It uses the Spice Cayenne accelerator, which the Spice documentation recommends for high-throughput CDC.
version: v1
kind: Spicepod
name: document-search
embeddings:
- from: huggingface:huggingface.co/sentence-transformers/all-MiniLM-L6-v2
name: minilm
datasets:
- from: postgres:public.documents
name: documents
params:
pg_host: db.internal
pg_port: '5432'
pg_db: app
pg_user: spice
pg_pass: ${secrets:PG_PASS}
pg_sslmode: verify-full
acceleration:
enabled: true
engine: cayenne
mode: file
refresh_mode: changes
primary_key: id
on_conflict:
id: upsert
columns:
- name: body
embeddings:
- from: minilm
row_id: id
chunking:
enabled: true
target_chunk_size: 512
overlap_size: 64
full_text_search:
enabled: true
row_id:
- idHybrid SQL search then queries both indexes over the same synced rows and fuses the results with reciprocal rank fusion.
SELECT id, title, _fused_score
FROM rrf(
vector_search(documents, 'refunds for annual plans'),
text_search(documents, 'refund annual plan', body),
join_key => 'id'
)
ORDER BY _fused_score DESC
LIMIT 5;Embedding columns work the same way on datasets from any of the supported data connectors, with a refresh mode chosen per source. Retrieval-augmented generation pipelines then read current vectors from the same runtime that runs the SQL, with no separate embedding job to schedule.
Embedding Sync FAQ
How often should embeddings be updated when source data changes?
Update embeddings when the source text changes, not on a fixed calendar. Pipelines based on change data capture or database triggers update the index within seconds of a commit. A scheduled job leaves results stale for up to one full interval. Choose the method from the amount of staleness that the application can accept.
Does editing one paragraph require re-embedding the whole document?
Not always. With structure-aware chunking, an edit changes only the chunks of its own section, and a content hash identifies those chunks. With fixed-size chunking, an early edit shifts every later boundary, so most chunks change anyway. For short documents, re-embedding the whole document is often the simpler choice.
How do you remove deleted documents from a vector index?
Capture the delete event and remove every vector keyed by the document ID. Change data capture, database triggers, and application hooks on the delete path see hard deletes, but timestamp polling does not. Also run a scheduled reconciliation that compares index keys with a consistent snapshot of the source keys. It removes entries for deletes that the change stream missed.
Should an embedding pipeline use timestamp polling or change data capture?
Timestamp polling queries for rows with an updated_at value later than a stored watermark. It is simple, but it misses hard deletes and rows whose writer did not set the timestamp. Change data capture reads the transaction log, so it sees every committed insert, update, and delete in order. Use change data capture for tables that users edit or delete.
How can you tell whether a vector index contains stale embeddings?
Store a content hash with each chunk and compare it with a hash of the current source text. A mismatch marks a stale chunk. To measure freshness end to end, write a canary record with a unique token and time how long search takes to return it. Also watch for a lasting gap between the number of source rows and the number of distinct source IDs in the index.
Can a vector index and a full-text index fall out of sync?
Yes. Hybrid search reads both indexes, and each index must apply the same inserts, updates, and deletes. If only one index applies a delete, the deleted document can still appear through the other one. Drive both indexes from one change stream keyed by the same primary key.
Learn more about keeping embeddings in sync
Documentation and engineering posts on embedding columns, real-time change data capture, and hybrid search over live data.
Embedding Datasets Docs
Configure embedding columns, chunking, and row identifiers on Spice datasets, and choose between accelerated and just-in-time embeddings.
Spice 2.0: Real-Time Analytical Query on Operational Data, Without ETL
Native CDC replication from the PostgreSQL WAL, MySQL binlog, and MongoDB oplog keeps accelerated datasets current without Debezium or an external streaming layer.
Real-Time Hybrid Search Using RRF: A Hands-On Guide with Spice
Learn how to build hybrid search with RRF directly in SQL using Spice, combining text, vector, and time-based relevance in one query.
See Spice in action
Get a guided walkthrough of how development teams use Spice to query, accelerate, and integrate AI for mission-critical workloads.
Get a demo


