Microsoft AI-103: Azure AI Search Vectorizers

Azure AI Search vectorizers perform query-time conversion from text or image input into a vector using an embedding model configured in the search index. This is different from the embedding skill used during indexing: the skill turns document chunks into vectors, while the vectorizer turns an incoming query into a compatible vector automatically at search time.

Within Microsoft AI Agents, vectorizers simplify RAG clients because the application can send the user’s text rather than managing embedding calls and vector dimensions itself. The index schema still needs a vector field, vector profile, and a vectorizer whose model matches the embedding model used during indexing.

Current Azure AI Search integrated vectorization supports Azure OpenAI vectorizers, custom Web API vectorizers, Azure Vision multimodal vectorization in preview, and Microsoft Foundry model-catalog vectorization in preview for supported models.

The query vectorizer must match the index embedding space

Similarity search only works when query and document vectors come from compatible embedding models and dimensions.

Changing the query vectorizer to a newer embedding deployment without rebuilding document vectors can make similarity scores meaningless even if the API accepts the request.

Treat the embedding model/deployment/config as one versioned pair for indexing and querying.

Vectorizers are attached through vector profiles

The index schema defines vectorizers and vector-search profiles, and vector fields reference a profile.

This separates algorithm/profile settings from the model endpoint used to convert text into a vector.

Document the profile/vectorizer/field relationship because large indexes can contain several vector fields with different models or modalities.

Azure OpenAI vectorizers reduce application plumbing

With an Azure OpenAI vectorizer, Azure AI Search calls the configured embedding deployment during query execution.

The application sends text through a vectorizable query rather than making a separate embedding call, handling retry, and injecting the float array.

This can simplify authentication and code, but query latency now includes the vectorizer call and depends on the embedding deployment’s quota/availability.

Managed identity is preferable to embedded model keys

Current vectorizer guidance supports keyless authentication with managed identity and appropriate roles for the Search service to call Azure OpenAI or other supported resources.

This avoids storing long-lived API keys in index definitions or deployment scripts.

Test role scope and private-network connectivity during deployment; a search query can fail even when the search service itself is healthy if it cannot reach the embedding endpoint.

Custom Web API vectorizers open the model boundary

A custom vectorizer can call an arbitrary compatible endpoint when the organization uses an embedding model not provided directly through built-in integrations.

This adds flexibility but also makes latency, retry, schema, scaling, authentication, and model-version compatibility the customer’s responsibility.

Keep the custom vectorizer contract narrow and benchmark under concurrent query load, not only single requests.

Foundry model-catalog vectorizers are still preview

Microsoft currently documents a Foundry model catalog vectorizer in preview that works with supported deployed models from the catalog.

Preview status should be reflected in production risk, API-version selection, and upgrade testing.

Do not assume every catalog embedding model is supported identically for portal, SDK, indexing skill, and query vectorizer paths; check the current supported-model lists.

Vectorizer latency contributes directly to search latency

Query-time embedding happens before vector retrieval, so embedding-model response time is part of end-to-end RAG latency.

Track vectorizer latency separately from Search execution where telemetry allows, and test concurrent load against embedding quota.

For repetitive queries, application-level caching may help, but cache keys must include model/version and normalized query text.

Hybrid search still benefits from query-time vectorization

A vectorizable text query can participate alongside lexical search text in a hybrid query.

The same user phrase can feed BM25/semantic ranking and the vectorizer, reducing duplicate client logic.

Candidate counts and filtered vector search settings still determine the retrieval side; vectorization only creates the query vector.

Indexing and query-time failures should be operated separately

An embedding skill can fail during indexer runs because of throttling or malformed documents, while a vectorizer can fail live user queries because the embedding endpoint is unavailable or unauthorized.

These are different SLOs and alerts.

Azure AI Search Indexer Failures covers the ingestion path; query vectorizer health belongs in the online retrieval path.

Model upgrades require coordinated re-embedding

If an embedding model is retired or the team wants better quality, create a migration plan that can rebuild document vectors and switch the query vectorizer coherently.

For large indexes, consider a parallel new field/index or blue-green rebuild so old and new embeddings are never mixed accidentally.

Measure retrieval quality on an evaluation set before switching production traffic.

Vectorizers are successful when query embedding becomes an invisible, versioned dependency

The mature design records model/deployment, dimensions, profile, authentication, latency, quota, failure behavior, and migration plan alongside the index schema.

Integrated vectorization should reduce application complexity without making the embedding model an undocumented hidden dependency inside Search.

Vectorizer configuration should be treated as part of the index contract. Changes to endpoint, deployment name, authentication, dimensions, or model version can affect every live vector query immediately. Manage it through infrastructure-as-code and review diffs just like changes to searchable/filterable fields.

Embedding-model quota can become the online search bottleneck even when Search has ample replicas and partitions. Monitor failed vectorization calls and query latency during traffic spikes. A retrieval service may look healthy from Search metrics while users receive errors because the embedding deployment throttled first.

Private-network deployments add one more dependency: the Search service must reach the embedding endpoint through a supported path. Shared private links, DNS, managed identity, and regional placement should be validated together. Test actual query-time vectorization after deployment; successful creation of the index schema does not prove the vectorizer can call the model.

Multi-language retrieval should be evaluated against the embedding model’s real behavior. A single multilingual model may serve several languages well, while another design uses language-specific fields or models. If one index contains multiple vector spaces, assign profiles/fields explicitly and route queries to the correct vectorizer rather than mixing incomparable vectors.

Image or multimodal vectorizers introduce additional input normalization concerns. Query images, text prompts, and document-side embeddings need to use the same compatible model/configuration. Preview multimodal capabilities should be tested for model-version and API changes before they become a hard production dependency.

Client-side embedding can still be appropriate when the application already maintains a central embedding service, needs custom caching, or must use a model not supported by an integrated vectorizer. Integrated vectorization reduces plumbing, but it should not force duplication of an existing well-governed embedding platform.

Query diagnostics should log which vectorizer/profile ran and the model deployment behind it. This is especially important during migrations where two index versions coexist. If retrieval quality changes, the team should be able to tell whether the document vectors, query vectorizer, HNSW profile, or filter behavior changed.

Use an evaluation set to catch accidental mixed embedding spaces. Rebuild a sample index with the candidate model, run the same queries, and compare recall/ranking before cutover. If results collapse suddenly after a configuration change, suspect model/profile mismatch before assuming the content itself became less relevant.

Vectorizer migrations should be blue-green when the index is large or business-critical. Build a parallel index or parallel vector field with the new embedding model, populate it completely, run retrieval evaluations, and only then switch query traffic. This prevents a mixed state where some documents use one embedding space and others another.

Embedding dimensions and index storage cost should be part of model selection. Larger vectors can improve quality for some workloads but consume more index storage and memory and can affect search performance. Evaluate quality gain against service size, indexing time, query latency, and cost rather than selecting the largest embedding model automatically.

Vectorizers also make dependency ownership explicit: the Search team owns index/query configuration, while the embedding deployment may be owned by an AI platform team. Define who monitors quota, rotates credentials/identity permissions, handles model retirement, and approves version changes so a query-time dependency does not fall between teams.

Finally, keep a client-side fallback strategy if the product needs graceful degradation. For some applications, lexical/hybrid text search can still return useful results when query vectorization is temporarily unavailable. If such fallback is allowed, label the response/telemetry so quality differences are visible rather than silently treating degraded search as normal.

Vectorizer configuration should be part of disaster-recovery testing. A restored or secondary Search service needs the same vectorizer definition, network path, managed identity permissions, model deployment, and compatible index schema. A copied index without a working query-time embedding dependency can pass basic document lookups while every vectorizable query fails.

Keep a deterministic lexical fallback only when product semantics allow it, and monitor how often it triggers. A frequent fallback means the vectorization dependency is unhealthy or underprovisioned and should be fixed rather than normalized as ordinary retrieval behavior.

Treat query-time embedding as a production dependency with its own SLO, ownership, migration plan, and observable failure mode.

Keep the embedding contract versioned and tested.

Keep it observable.

Embedding changes should be handled like schema changes. Record the vectorizer, model version, dimensions, chunking assumptions, and reindex strategy so queries do not silently compare incompatible representations after an upgrade or partial rebuild.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!