Reranking adds a second relevance pass after initial retrieval. Instead of relying only on vector distance, keyword score, or reciprocal rank fusion, a reranker evaluates the query and each candidate document together and produces a more precise ordering. Databricks AI Search now includes a generally available Databricks reranker that can be enabled on ANN, hybrid, or full-text queries, and Databricks recommends trying reranking for RAG workloads where answer quality matters more than minimum retrieval latency.
Within Generative AI on Databricks, reranking is the bridge between broad candidate recall and the small evidence set ultimately given to the LLM. It is most valuable when the first retrieval pass finds roughly the right documents but does not order them reliably enough.
The current AI Search API reranks the top candidate set before returning the requested final results, so candidate count, selected columns, and latency budget are all part of the design.
Initial retrieval should optimize recall
The first-stage search—ANN, hybrid, or full-text—should retrieve a broad enough candidate set that the relevant document is likely present.
Reranking cannot recover a document that was never retrieved.
Use query rewriting, filters, hybrid search, better embeddings, and larger candidate pools when first-stage recall is the actual weakness.
The Databricks reranker re-scores the top candidates
Current AI Search passes the top candidate results into the Databricks reranker when a reranker is configured.
The reranker considers the query and selected text/metadata columns jointly rather than only the precomputed embedding distance.
This cross-encoder-style second pass is slower than vector lookup but can separate subtle relevance differences better.
columns_to_rerank determines what evidence the reranker sees
You can choose columns such as chunk text, title, parent summary, section, or description for reranking.
The selected columns are concatenated for the reranker.
Include metadata that actually helps relevance; dumping dozens of noisy columns increases input without necessarily improving ranking and can increase latency/cost.
Reranking works with ANN, hybrid, and full-text retrieval
Because reranking is a post-retrieval step, the same reranker can be applied to different first-stage algorithms.
Evaluate combinations such as ANN+rerank, hybrid+rerank, and full-text+rerank on the same query set.
One corpus may benefit most from hybrid recall plus reranking, while exact-code/documentation search may get little value beyond full-text.
Quality gains must be measured against latency
Databricks notes that reranking introduces a small additional latency and can improve retrieval quality materially, with current guidance citing roughly 10% improvement as a typical reference point rather than a guarantee.
For interactive chat, measure p95 retrieval latency and total time to first answer alongside DCG/relevance.
High-throughput low-latency search endpoints may prefer simpler ranking for some query classes.
Use retrieval-quality evaluation instead of anecdotal examples
Databricks AI Search includes a retrieval evaluation workflow that compares query types with and without reranking using metrics such as DCG@10 and average relevance.
Inspect per-query winners and failed searches; averages can hide a subset where reranking consistently hurts exact-match queries.
Use a representative query set before enabling reranking as a blanket default.
Reranking can reduce context-window waste
When better ranking moves truly relevant chunks to the top, the application can often send fewer documents to the LLM without losing answer quality.
This reduces model input tokens, noise, and the chance that irrelevant passages distract the generator.
The added retrieval latency may therefore be offset partly by shorter model context and faster/cheaper generation.
Do not rerank already tiny result sets blindly
If first-stage retrieval returns only three highly precise documents, a second model pass may provide little benefit.
Reranking is most useful when there is meaningful ordering uncertainty in a larger candidate pool.
Use query-class routing or thresholds so trivial exact-ID lookups can bypass the reranker while ambiguous conceptual questions receive it.
Fine-tuned rerankers can target specialized domains
Current AI Search APIs support a fine-tuned reranker hosted through a Model Serving endpoint in addition to the base Databricks reranker.
This can help legal, biomedical, code, or proprietary-domain retrieval where general semantic ranking misses organization-specific relevance.
Fine-tuning adds dataset, serving, versioning, and monitoring burden, so establish a strong base-reranker evaluation before training a custom model.
Reranking and query rewriting should be evaluated together
A rewritten query changes the candidates and the semantic wording the reranker receives.
RAG Query Rewriting in Databricks should therefore be optimized jointly with reranking rather than in isolated A/B tests.
Track original query, rewritten query, first-stage scores, reranker scores, selected chunks, and final answer so regressions can be traced to the correct stage.
Reranking succeeds when the second pass improves evidence order enough to justify its cost
The mature system retrieves broadly, selects useful reranker columns, compares algorithms with and without reranking, measures DCG/relevance and latency, bypasses reranking for easy exact queries, and keeps candidate/reranker traces for debugging.
Reranking is valuable not because it is another model call, but because it turns a noisy candidate set into a smaller, better evidence set for generation.
Candidate-set size is a major reranking knob. Too small and the relevant document may never reach the second pass; too large and latency grows because the reranker scores many weak candidates. Tune top-N candidates against recall and p95 latency instead of copying a fixed number from examples.
Reranker input should include concise fields with complementary signal. Chunk text plus title or parent summary is often stronger than chunk text alone, while repeating long body fields increases latency. Test column combinations using AI Search retrieval-quality evaluation rather than assuming more context always helps.
Metadata filters should run before reranking when they enforce hard constraints such as tenant, language, date, or document type. The reranker is a relevance model, not an authorization engine. Never rely on a low reranker score to suppress content the caller is not allowed to see.
Score interpretation should remain internal. First-stage similarity and reranker relevance scores are produced by different models/scales and should not be compared as if they were calibrated probabilities. Use them for ranking/thresholding only after corpus-specific evaluation.
Reranker outages need a degradation policy. For many applications, returning the first-stage AI Search ranking is better than failing the whole answer; for high-stakes systems, quality requirements may justify abstaining. Make fallback explicit and log it so degraded responses can be analyzed separately.
Fine-tuned rerankers should be evaluated against the base Databricks reranker with a held-out query/document relevance set. Domain tuning is useful only when it improves ranking enough to justify model-serving cost, maintenance, and version management.
Context selection after reranking should consider redundancy. The top five results can all be nearly identical chunks from one document. Apply document diversity, parent grouping, or maximal-marginal-relevance-style selection if answer quality suffers from duplicate evidence, while preserving the highest-ranked truly relevant passages.
Production monitoring should sample retrieval traces and compute online/offline relevance metrics over time. Corpus growth, new document types, embedding migrations, or query changes can reduce reranker effectiveness even when the reranker model is unchanged. Treat reranking quality as a living search metric, not a one-time benchmark.
Reranking should happen after security filtering and before expensive generation. This order prevents unauthorized candidates from reaching the reranker and lets the application shrink context before paying LLM tokens. Keep filters deterministic and based on caller identity, then use reranking only for relevance among allowed candidates.
Reranker latency should be traced separately from AI Search first-stage latency. If total retrieval slows, operators need to know whether HNSW/full-text retrieval, reranker service, or metadata filters are responsible. Enable debug-level timing during benchmarking and preserve aggregate stage metrics in production.
DCG@10 is useful for ranked retrieval, but final RAG quality also depends on answer-support sufficiency. A reranker can improve the top ordering while still leaving the answer without one necessary document. Pair ranking metrics with retrieval sufficiency and final grounded-answer evaluation.
Document-level deduplication should be evaluated with reranking. If many chunks from one relevant PDF occupy the top 10, a second relevant document may be pushed out. Group or diversify results based on parent IDs when multi-source coverage matters, while keeping the strongest chunk from each parent.
Reranking policy should be versioned. Changing reranker model, fine-tuned endpoint, columns_to_rerank, candidate count, or query type can all change results. Record those settings with retrieval evaluation runs so quality regressions can be reproduced and rolled back.
Reranking should be disabled for queries where deterministic sort order is the business requirement, such as newest document, highest rating, or lowest price. Current AI Search also supports sort columns in newer APIs; relevance reranking should not override an explicit business sort merely because an LLM thinks another item is semantically closer.
Cache or reuse reranker outcomes only when query, candidate set, reranker version, and relevant metadata are unchanged. Corpus updates can make an old ranking stale even if the user asks the same text again, so aggressive caching should include index/version freshness.
Keep reranking policy tied to retrieval metrics and the application’s latency budget.
Reranking decisions should remain transparent in traces so support engineers can compare first-stage order with final reranked order and explain why one source was promoted above another when users challenge a citation or answer.