Azure AI Search indexer failures range from obvious document errors to quieter problems such as timeouts, blocked network access, stale change-tracking state, missing documents, skill throttling, or a scheduled run that never reaches the end of a large data source. Microsoft describes indexers as best-effort scheduled ingestion: they retry transient problems over future runs, but each run has execution limits and strict timing control may require using the push indexing APIs instead.
Within Microsoft AI Agents, indexer reliability directly affects RAG quality. An agent can produce a plausible answer from an index that silently missed the newest documents, so ingestion health must be observable as part of the application, not just a search-admin concern.
The existing RAG on Azure article covers the end-to-end pipeline; this page focuses on operating the ingestion stage.
Execution history is the first source of truth
Azure portal indexer execution history and the Get Indexer Status API show recent runs, start/end time, items processed, errors, warnings, and document-specific keys when available.
Even a run marked Success can contain warnings worth investigating.
Capture the run timestamp and exact error before resetting anything, because resets can make the original failure harder to reproduce.
maxFailedItems controls when indexing stops
Indexers stop when the number of failures exceeds configured limits such as maxFailedItems and maxFailedItemsPerBatch.
Increasing those values can keep a large ingestion moving past a few bad documents, but it should not become a blanket way to ignore systematic parsing or schema errors.
Track the skipped document keys and remediate them later or quarantine them intentionally.
Scheduled indexers provide retry behavior across runs
Microsoft currently recommends schedules for reliable indexing because transient network, throttling, or per-document failures can be retried on later runs.
For large sources, scheduled execution also lets progress continue after a run reaches its maximum processing window.
Do not rely only on manual Run calls if the pipeline needs to catch up automatically after temporary failures.
Runtime limits can leave documents unprocessed without a permanent failure
Indexers have managed execution time limits that vary by execution environment. Skillset runs in the multitenant environment commonly have a shorter maximum than private execution through shared private links.
Large files, slow custom skills, complex enrichment, or huge sources can cause the run to stop before every document is processed.
Partition data sources or simplify skill work if the indexer consistently cannot finish useful increments.
Network and identity failures often look generic
Data sources behind firewalls, private endpoints, or virtual networks need supported connectivity such as shared private links and correct managed identity roles.
A 403 or generic network error can come from public access being disabled, an unapproved private-link connection, an NSG rule, wrong DNS, or an expired key.
Test the source independently and verify the indexer execution environment before changing search mappings.
Skill failures need both Search and downstream-service diagnostics
An embedding, document extraction, or custom Web API skill can fail because of throttle, timeout, invalid input, malformed output, authentication, or network reachability.
Indexer errors often identify the skill and enrichment path, but the downstream service logs/metrics may contain the actual cause.
Azure AI Document Intelligence Custom Extraction is relevant when a custom extraction service participates in the pipeline.
Missing documents can be change-tracking problems
If a source item exists but never appears in the index, check whether it was modified after the last successful high-water mark, whether the indexer timed out first, and whether source change-tracking prerequisites are correct.
A future-dated high-water-mark field or missing database index can cause data to be skipped or queries to time out.
Indexer status exposes initial/final tracking state for supported sources, which can help diagnose this class.
Reset and Run have distinct roles
Reset clears the indexer’s internal high-water mark; it does not perform indexing by itself. Follow it with Run or wait for the next scheduled execution.
A full reset causes all source documents to be reprocessed and can be expensive for skillsets and large corpora.
Use it when schema/enrichment changes require rebuilding everything, not as the first reaction to every transient failure.
Document or skill-level reset can be more precise
Current Azure AI Search provides preview reset APIs for specific documents or skills and a resync mode for supported permission-field scenarios.
These can avoid a full-corpus rebuild when the affected scope is known.
Keep API version and preview status explicit in automation because these capabilities can evolve faster than stable Run/Reset behavior.
Reset does not automatically delete every orphaned index document
Microsoft notes that reset/run overwrites documents found in the source but does not necessarily remove index documents whose source record disappeared in unsupported deletion-tracking scenarios.
This matters after source cleanup or chunk/projection changes.
Azure AI Search Index Projections should be validated for stale child chunks after structural changes.
Indexer reliability is successful when ingestion freshness is an application SLO
The mature service monitors last successful run, items processed, errors/warnings, skipped documents, runtime duration, source lag, embedding/skill failures, and index freshness.
Agents should not quietly answer from stale corpora. Ingestion health belongs beside retrieval and model health in the production dashboard.
Indexer monitoring should distinguish freshness from last-run status. A job can report Success after processing only a small incremental set while the source is still hours behind because previous runs timed out. Track source watermark age, documents pending, and last indexed modification time for important corpora so a green status cannot hide growing lag.
Skill throttling often needs capacity coordination across services. An indexer can run many enrichment operations in parallel while the downstream Azure OpenAI, Document Intelligence, or custom API has its own quota. If the skill service is repeatedly throttled, lowering concurrency or scaling the dependency may be more effective than resetting the indexer endlessly.
Schema mismatches deserve fast failure. A changed source type or skill output can fail deserialization or field mapping when the target index expects another type. Validate representative documents in a staging index before modifying production schema, and keep field mappings/skill outputs versioned with the application that depends on them.
Large blob/document sources should be partitioned in a way the team can reason about. Multiple indexers over distinct containers, prefixes, or logical partitions can reduce one run’s duration and make recovery more targeted. Avoid overlapping partitions unless duplicate processing is intentional, because the same source can be written multiple times and make counts confusing.
Warnings should be categorized rather than ignored globally. “No text found in image” may be expected for one file type, while an embedding warning on every PDF can indicate a broken enrichment path. Build alert thresholds by warning class and volume so normal noise does not suppress meaningful drift.
Indexer credentials should have a rotation runbook. Managed identity reduces key-rotation risk, but role assignments, tenant policy, private links, DNS, and firewalls can still break. For key-based data sources, test key rotation in nonproduction and update both secret source and data-source definition through infrastructure-as-code.
Document counts should be reconciled carefully. One source document can produce several projected chunks, while parent-only rows or skipped documents can alter index count further. Compare logical parent counts and chunk counts separately rather than declaring success because the index contains “about the same number” as the source.
Finally, define when to abandon indexers for a push pipeline. If the application needs strict transaction ordering, custom retry semantics, immediate deletion guarantees, or per-document control that managed indexers cannot provide, the Documents Index API may be a better fit. Indexers are valuable for managed ingestion, not a universal replacement for application-owned ETL.
Alerting should distinguish one bad document from systemic failure. A single corrupt PDF can be quarantined; a sudden spike in thousands of field-mapping errors after a schema change needs immediate rollback. Error volume, percentage of documents affected, and repetition across runs are better escalation signals than a binary “indexer failed.”
Operational dashboards should also show last completed full/backfill run and current incremental cadence. If a team only watches incremental Success runs, it may never notice that a structural index rebuild was requested but never finished. Separate steady-state freshness from migration/backfill progress.
For agent applications, stale-index behavior should be explicit. The UI or orchestration layer can continue serving from the previous index, disable answers for affected sources, or warn about freshness depending on business risk. Search ingestion should fail observably instead of letting the model answer confidently from a corpus the platform knows is incomplete.
Keep a small failure corpus in nonproduction: malformed PDF, oversized document, missing permission, throttled skill, schema mismatch, and temporarily unavailable source. Re-run these cases after Search/API-version or skill changes to confirm alerting and retry behavior still work. Reliability improves when indexer failure modes are tested deliberately instead of learned only from production incidents.
Indexer incidents should end with a prevention change where possible: schema validation, better source partitioning, quota monitoring, network health checks, or a quarantined-document path. Repeated manual resets are a symptom that the ingestion system still lacks a reliable recovery design.
Freshness should be measurable from source change to searchable document. Track extraction, transformation, enrichment, indexing, and query visibility separately so an SLA breach points to the failing stage instead of appearing as a generic ‘search is stale’ complaint.