Back to News & insightsEngineering

Changing embedding models without breaking your search

Follow an embedding migration from relevance judgments and index design to backfills, permission checks, shadow queries, cutover, and rollback.

Editorial guide · Updated September 20, 2026 · 8 min read
Floating metal and glass layers connected to a central core, illustrating application architecture.

An embedding model turns text into numerical representations used by a search system. Replacing it can change which documents appear relevant even when every document and query stays the same. That makes an embedding migration a change to product behavior, not merely a library update or an exchange of one model identifier for another.

The index, query encoding, text preparation, and relevance thresholds form a coordinated system. A new model can require a new index and different evaluation assumptions. Successfully generating vectors proves only that the pipeline ran. It does not prove that users can still find the answer they came for.

Consider a fictional equipment documentation service. Engineers search manuals, service bulletins, and installation notes. Some queries contain exact part numbers, while others describe symptoms. The service also separates public documents from customer-specific instructions. This guide follows a migration for that scenario; its design examples are recommendations to test, not measured results from a real company.

State the reason for the migration

Write down the problem you intend to solve. Perhaps symptom descriptions retrieve weak matches, the supported languages are expanding, or the current deployment no longer fits operational requirements. Those are different reasons, and each requires its own evidence. A strong score on an unrelated public dataset is only a reason to investigate a candidate.

Collect examples of current failures before implementing the replacement. For the documentation service, save cases where the correct manual exists but does not appear near the top. Include cases that currently work well. Otherwise, the evaluation may demonstrate improvement on a small problem set while overlooking regressions in ordinary searches.

Also define the acceptable operating envelope: index storage, query latency, update delay, and infrastructure effort. A candidate that improves relevance but cannot keep newly published bulletins searchable within the required time may be unsuitable for the overall service.

Match the model to the retrieval task

Sentence Transformers distinguishes symmetric search, where queries and candidates resemble each other, from asymmetric search, where a short query retrieves a longer passage. Its documentation also describes dedicated query and document encoding methods. Follow the selected model's intended encoding conventions instead of assuming every text should be prepared identically.

For the equipment service, a short symptom description and a manual section are different kinds of input. Check whether the candidate expects a task prefix, instruction, or particular pooling and normalization behavior. Record that configuration as part of the model version used by the index.

Test exact identifiers separately from semantic paraphrases. A model that connects “motor becomes hot” with “thermal overload” may still confuse two nearly identical part numbers. Keep an exact-match retrieval path when the product requires it, and evaluate how its results are combined with semantic candidates.

Treat vectors as belonging to a specific space

Do not assume that old document vectors can be compared with queries encoded by a new model. Matching dimensions are not proof of compatible coordinates. Unless compatibility is explicitly established for your setup, encode documents and queries with the corresponding model and configuration.

Create a separate index generation for the candidate. Store a manifest containing the encoder identifier, revision, dimension, similarity function, preprocessing version, and source snapshot. The manifest should make an accidental mismatch detectable before the query reaches the vector store.

Put a configuration check in the query path. If the selected query encoder and index generation do not belong together, fail visibly or return to a verified pair. Silently searching incompatible vectors can produce plausible-looking results, which makes the defect harder to notice than an explicit error.

Freeze enough of the corpus to compare fairly

Choose a document snapshot for the offline comparison and record which revisions it contains. Keep the old and candidate systems on the same source material during that experiment. If one index contains a newly corrected bulletin and the other does not, a relevance difference may come from the document update rather than the encoder.

Inspect extraction and chunking before the backfill. A broken parser that discards table headings can damage both systems. For the first migration experiment, hold chunk boundaries constant where possible. If you also want to change chunking, evaluate that as a separate condition so you can identify what caused an improvement.

Preserve document identity across index generations. A stable source identifier and revision are more useful than relying only on the vector store's internal row identifier. They allow reviewers to compare results, enforce deletions, and identify duplicates even when the underlying vector records are replaced.

Build relevance judgments around real questions

For each diagnostic query, identify passages that answer the question and passages that merely mention related vocabulary. More than one result may be relevant. A symptom query might require both a troubleshooting paragraph and a warning that limits the recommended procedure.

Have a domain reviewer inspect disagreements and ambiguous labels. Keep queries grouped by purpose: exact lookup, symptom search, comparison, and questions with no supported answer. A single average makes it too easy to miss that the migration helps one group and harms another.

Measure whether relevant material enters the candidate set and where it appears after ranking. Also inspect what the answer generator receives if search feeds a RAG workflow. Better retrieval metrics are useful, but the final check is whether the user obtains supported, applicable information with manageable effort.

Backfill with checkpoints and bounded load

Process source records in batches and record the last completed checkpoint. Use stable identifiers so repeating a batch does not create duplicate active records. Account for failed encodings explicitly; a completed job with a silent missing tail is not a complete index.

Limit concurrency according to your infrastructure and provider allowances. Keep foreground traffic protected while the backfill runs. Measure the rate at which documents are processed and project the remaining work from observed throughput, rather than promising a completion time from a small warm-cache trial.

Maintain an error queue containing the source identifier and a safe diagnostic reason. Oversized documents, invalid text extraction, and transient service errors deserve different recovery paths. Avoid copying confidential document contents into unrestricted job logs just to make a failure easier to inspect.

Keep updates and deletions consistent

A long backfill runs while the source corpus continues to change. Capture updates that occur after the initial snapshot and apply them to the candidate generation before cutover. The exact mechanism depends on your database and architecture, but the migration needs an explicit way to close that gap.

Treat deletions as first-class events. A document removed during the backfill must not reappear when an older batch finishes later. Use revision checks or tombstones as appropriate, and verify that a delayed job cannot overwrite the deletion with a stale copy.

Keep a reconciliation report: source record count, indexed count, missing identifiers, duplicate active revisions, and the age of the latest applied update. Counts alone do not establish correctness, but they help locate gaps. Sample the content behind matching identifiers to check that the right revisions were actually encoded.

Preserve permissions through every branch

Apply the same authorization rules to the candidate index that protect the current system. Include account scope and document visibility in the indexing contract, and enforce current permissions during retrieval. A vector result should never gain access merely because it came from a new implementation.

Run paired searches with accounts that have different rights to similar manuals. Inspect returned identifiers, snippets, and any cached answer. Permission filtering after a generated answer is too late to prevent restricted evidence from influencing that answer.

Test a visibility change during migration. Make a customer document private, revoke an account's access, and search again through both index generations. The old and new branches need to agree on the access boundary even if their relevance ordering differs.

Shadow queries before changing the user experience

Send an appropriately sampled, permitted set of queries to the candidate system while continuing to serve results from the current system. Ensure the shadow path is read-only and does not create extra user-visible actions. Include its compute and logging costs in the rollout budget.

Compare result overlap, relevance on reviewed cases, empty-result frequency, and latency. Low overlap does not automatically mean the new model is worse; it may be finding better passages. High overlap does not prove quality either. Use disagreement sampling to direct human review toward the cases that actually changed.

Prevent the experiment from creating an unexpected data transfer. If the candidate encoder runs through another provider, confirm that the selected traffic is allowed on that route before shadowing it. The query path's data boundary should be documented just as carefully as its ranking configuration.

Recalibrate thresholds and downstream assumptions

A similarity cutoff tuned for the old encoder may not have the same meaning with the new one. Inspect score distributions and labeled outcomes before reusing it. Avoid converting a raw similarity value into a probability of correctness without evaluating that interpretation.

For the equipment service, test the “no relevant documentation found” decision explicitly. Lowering a cutoff may reduce empty results while introducing misleading procedures. Decide whether the user should see a weak match, be asked for a part number, or receive a clear absence-of-evidence message.

Review downstream rerankers, caches, and analytics. Cache keys should distinguish index generations where results can change. Monitoring should identify the active encoder and index so a later regression can be traced. An embedding migration is incomplete if only the vector-writing job knows which version is running.

Cut over with a coherent rollback plan

Switch the query encoder and index generation together through a controlled configuration change. Start with a bounded share of traffic and observe both relevance and operational behavior. Maintain the previous pair for a defined rollback window, keeping updates and deletions consistent for as long as it remains a valid fallback.

Set rollback triggers before the release: authorization failures, material regressions on critical queries, or unacceptable response times. If you return to the old index, verify that it still contains current permitted data. A fallback that has been abandoned during the migration may no longer be safe or useful.

Finally, document what justified the change and what remains uncertain. Keep the evaluation queries, index manifest, and reconciliation results with the release record. A successful migration leaves users with better search and leaves the team able to explain, reproduce, and reverse the change when necessary.

Further reading

Sentence Transformers: Semantic search and query encoding

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.