Understanding Images and Video with Microsoft Foundry
Understanding an image and generating a new one require different contracts. An image-generation service is asked to invent a visual asset; an analysis service is asked to make an evidence-based statement about media that already exists. If an AI assistant describes a damaged product, extracts information from a chart, or searches a training video for a relevant moment, the primary engineering requirement is fidelity to the source. A plausible but unsupported answer can be worse than an explicit admission that the available pixels, frames, or transcript do not show the needed fact.
Microsoft AI-103’s computer vision objectives include multimodal visual reasoning, captions and accessible descriptions, question answering grounded in images, video segment analysis, Content Understanding, and safety. They also require recognizing the limitations of the relevant model or service. A strong implementation treats uploaded media as evidence with permissions, provenance, quality constraints, and a verification plan, not as an instruction to the model to invent missing details.
Choose the analysis task and evidence standard
Begin by asking what the reader or downstream system actually needs. A photo library may need a short description to improve search. A manufacturing workflow may need a known object or component identified. A user with a visual accessibility requirement may need alternative text that conveys the image’s functional role. A training-video library may need accurate timestamps for a spoken explanation or a segment showing a particular process. These outputs differ in structure and risk, even if the same multimodal model could produce fluent text for several of them.
Define the intended answer and the acceptable source. For an image question, the source might be one photograph, several related images, or a chart with a caption. For video analysis, it might be selected frames plus a transcript, not every moment of the full recording. An application should know when its evidence is insufficient: a blurred label, an occluded component, or a video frame sampled too sparsely does not support a precise answer. Mark uncertainty and request better input rather than producing a guess disguised as a fact.
For instance, consider a fictional equipment photo in which a warning label is partially hidden. The assistant can describe the visible equipment and point out that the full text of the label is unreadable. It should not infer an exact safety instruction merely because similar products usually carry one. If the workflow must identify a serial number, run a suitable OCR or specialized parser and verify the source region instead of depending on the model’s narrative description.
The computer vision troubleshooting process is relevant here because an incorrect output may originate in capture quality, resize settings, model suitability, thresholds, or downstream interpretation. The model is only one part of the evidence pipeline. A good test distinguishes a failure to see the evidence from a failure to reason about evidence that was captured correctly.
Select supported Foundry analyzers and understand API lifecycles
Azure Content Understanding in Foundry Tools supports structured analysis of documents, images, audio, and video. Its current documentation identifies 2025-11-01 as the generally available API version for production-oriented workflows and 2026-06-01-preview as the separate preview for newer capabilities. The preview may offer agentic document understanding and enhanced analysis options, but those are not silently available in the GA contract. Region support, resource configuration, compatible model deployments, quotas, output shape and access requirements should all be verified for the exact analyzer chosen.
Prebuilt analyzers have different purposes. An image-search analyzer can describe image content in a representation intended for retrieval. A video-search analyzer can produce transcript information, keyframes and segments that support finding moments in long media. Custom analyzers can apply a business-defined field schema when an application needs named results rather than an unstructured description. Do not pretend that image retrieval, OCR, object recognition, and semantic question answering always share the same request and response fields.
Microsoft’s newer Content Understanding model also changes how teams should think about legacy vision APIs. Azure Vision Image Analysis 4.0 is documented as deprecated, with an announced September 25, 2028 retirement. That future retirement date is not a claim that Image Analysis has stopped working today, but new long-lived designs should review recommended replacements and migration guidance. A technical study article should teach supported current choices rather than accidentally encouraging a dependency that the vendor has already scheduled for retirement.
Use an appropriate resource identity and avoid placing subscription keys in a public browser. The application must also control which user can submit a media asset and who can later view the analysis or original. Source media may contain faces, private documents, location clues, or other sensitive information that does not become public simply because the system creates a searchable description.
The AI-103 study guide still mentions single-task and pro-mode Content Understanding pipelines, but the service lifecycle makes that wording historical. Microsoft’s older 2025-05-01-preview offered standard single-file analysis and a pro mode for advanced multi-document reasoning; that preview and its pro-mode configuration are retired. The newer 2026-06-01-preview provides an agentic document-analysis workflow, configured through config.workflow, but it is not a drop-in replacement for old pro mode and initially supports one input file per request. The production-recommended 2025-11-01 GA analyzer contract does not include this preview-only agentic capability. Engineers should recognize the exam terminology, select a current supported workflow deliberately, and test how migration affects field schemas, evidence, costs and latency.
Submit an image analysis operation and track its output
A useful implementation starts with a known, permitted source image and a clear question. Microsoft’s Content Understanding REST quickstart documents a prebuilt-imageSearch analysis request using the GA API version. The following Python sample illustrates submitting that operation from an authorized backend. Configure FOUNDRY_ENDPOINT, FOUNDRY_API_KEY, and APPROVED_IMAGE_URL with actual test-only values from a permitted Foundry resource and publicly or otherwise service-accessible licensed image. It intentionally stops after submission and returns the operation URL rather than claiming analysis has completed.
import os
import requests
endpoint = os.environ["FOUNDRY_ENDPOINT"].rstrip("/")
key = os.environ["FOUNDRY_API_KEY"]
image_url = os.environ["APPROVED_IMAGE_URL"]
response = requests.post(
f"{endpoint}/contentunderstanding/analyzers/"
"prebuilt-imageSearch:analyze",
params={"api-version": "2025-11-01"},
headers={
"Ocp-Apim-Subscription-Key": key,
"Content-Type": "application/json",
},
json={"inputs": [{"url": image_url}]},
timeout=30,
)
response.raise_for_status()
operation_url = response.headers.get("Operation-Location")
if not operation_url:
raise RuntimeError("Expected analysis operation URL")
print("Analysis submitted:", operation_url)
This code is syntax-checked offline against the documented request pattern; it has not contacted Azure. Real applications must handle the service’s operation lifecycle, poll the returned status using the correct authorization and current API contract, and distinguish a pending job from a completed analysis. A submitted request is not a validated caption. Once the operation finishes, check the actual response schema for image descriptions, source references, fields and any confidence or grounding information provided by the chosen analyzer. Store only what the business task requires.
Use a stable media identifier and preserve enough input metadata to reproduce the comparison. If a later model or analyzer version produces a different answer, reviewers need to know whether the source image changed, whether preprocessing differed, and whether the application interpreted the response incorrectly. That is more valuable than storing a single generated paragraph without its source context.
Generate captions and accessibility descriptions that remain truthful
A caption for search is not always good accessibility text. A retrieval caption may mention many visible objects so an image can be found with varied queries. Alt text should communicate the image’s role in its surrounding content, emphasizing information a reader needs. A decorative background may need little or no description, while a chart conveying a trend may need an extended explanation or an accessible data table. The model cannot know every editorial purpose from pixels alone.
Consider a diagram of a cloud access workflow. A generic caption such as “boxes connected with arrows” misses the relationship that matters. A more useful accessible description identifies the stages and the direction of access checks, provided the diagram actually shows them. For a chart, verify that numbers, labels and trends come from the visible data and not from plausible assumptions. If the chart’s text is too small, use its source data or an approved high-resolution export instead of relying on a model to hallucinate a readable axis.
Before publishing generated descriptions, inspect them for unsupported demographic, health, identity or behavioral claims. Visual inference can be uncertain, and a confidently worded description may introduce information irrelevant to the user’s need. Define a review and correction path. If the description will be relied on for accessibility, include users with relevant accessibility needs in the evaluation process rather than assuming a high model confidence score proves usefulness.
Existing articles about multimodal input quality explain why image resolution, glare, occlusion, layout and preprocessing can alter results. Descriptions should be evaluated against the actual displayed version of the image, not a different crop or earlier unedited file. An image that changes after publication may require its accessibility text to be reviewed again.
Answer questions using visual evidence instead of guesses
Visual question answering has a grounding requirement. If the user asks how many warnings appear in a screenshot, the application should count visible warning elements or state that the image is insufficient. If the user asks whether a person followed a safety procedure, a single still image may not contain the temporal evidence needed to answer. The task should be scoped to what the media actually shows, with a path for clarifying questions or escalation.
A robust pipeline can separate perception and reasoning. First identify regions, text, diagrams or visual characteristics using a suitable modality-specific tool. Then let a model reason over a controlled representation that includes source IDs and limitations. For some tasks, a general multimodal model can process the image directly; for others, explicit OCR, layout extraction, structured measurement, or object detection offers more reliable intermediate evidence. There is no universal rule that the largest model is the best perception tool.
Prompts should constrain output to observed facts and ask the system to distinguish uncertain conclusions. That instruction alone is not sufficient, so maintain tests where the correct response is abstention. Examples include a partially hidden serial number, an unreadable handwritten note, a chart without a legend, and an image whose timestamp has been cropped away. Evaluate whether the application states the uncertainty correctly and avoids inventing a missing value.
When the system uses a structured analyzer, keep the values and their evidence references together. A downstream process should not be allowed to treat an extracted object label as an authorization decision. For instance, identifying a visible scratch does not prove warranty eligibility, which depends on policy and perhaps other evidence. This distinction between observation and business action protects users from confident but unjustified conclusions.
Interpret videos through scenes, keyframes, and transcripts
Video analysis is not the same as generating or remixing a clip. In analysis, the input is evidence that already exists; the objective is to find relevant moments, describe events, or extract structured metadata. Content Understanding’s video-search workflow can provide transcripts and keyframe or chapter information for supported inputs. The transcript represents spoken language, while frames describe visual content. Combining them can improve retrieval, but neither alone captures every meaningful event.
Sampling imposes important limits. Microsoft’s video documentation describes sampled frames at roughly one frame per second, with frames resized for analysis. Very brief visual events or small text can therefore be missed. Music, ambient sound, or other non-speech audio may not appear in a speech transcript. An application should not assert that an event never occurred simply because a sampled representation fails to show it. For high-consequence review, preserve the original video and support human examination of the relevant interval.
Design a query such as “Find the moment when the instructor demonstrates a secure sign-in step.” Determine whether the evidence should come from spoken instructions, a visible screen, or both. Index segments with timestamps and keep a path back to the original clip. A retrieved chapter label is an aid to navigation, not proof that the specific operation was performed correctly. The application should be able to show the footage or transcript that supports its answer.
Segmentation adds business choices. A long recording might be divided by scene changes, speaker turns, or topic boundaries. Overly short segments lose context; overly long segments can bury the answer. Custom segment schemas and agentic analysis may have different release status, costs and constraints. Test representative recordings and do not assume current preview behavior applies to every GA analyzer. When users rely on a result, explain the time range and source media that support it.
Protect media from prompt injection, misuse, and over-interpretation
An image can contain text that tries to instruct the AI application to ignore its task. A video frame may show a malicious prompt, an unrelated password, or sensitive material. OCR and multimodal understanding should treat this content as untrusted data. It cannot override the application’s trusted instructions, authorize new tools, or grant access to a private data source. A system that reads an instruction on a poster and treats it as an operator command has crossed a dangerous trust boundary.
Visual moderation also requires context. A classifier may produce a signal that media is unsafe or disallowed, but downstream policy still needs to decide how to handle uncertainty, human review, appeals, and data retention. Use only supported provider controls and avoid creating instructions to circumvent filtering. A production application should record the moderation outcome without needlessly exposing sensitive media to operators who do not need access.
Visual policy enforcement involves distinct checks that should not be collapsed into one image-safety score. A content-moderation service can classify supported harmful-image categories; prohibited symbols, brand marks, or unauthorized logo usage may require a separate approved reference set, purpose-built detector, or trained human reviewer. For generated media, the application can add a visible disclosure or controlled watermark at export where policy requires it, then check that the approved mark survives the delivery pipeline. Brand and provenance decisions should be tested against the actual media and rights documentation, not inferred from the quality of an AI-generated caption.
Provenance detection is also narrower than an authenticity guarantee. Microsoft documents tools that can inspect C2PA Content Credentials and invisible watermarking signals for supported content generated by Microsoft AI systems; absence of such a marker does not establish that arbitrary third-party media is human-made. In a brand-review test, use licensed sample images that include an approved mark, an altered mark, and a prohibited symbol. Record which classifier or rule produced each flag, route ambiguous results to review, and verify that false positives do not silently discard legitimate content. This turns watermarking, brand policy, and disallowed-symbol checks into explainable application controls rather than unsupported claims about a universal vision model.
Rights and consent are separate from technical support. The fact that a service accepts a URL does not prove the application is permitted to process its content. Use owned, licensed, or otherwise authorized test media. Keep source and derivative analysis accessible only to appropriate users, establish retention and deletion rules, and restrict inference about private people or confidential documents to legitimate user tasks.
Evaluate the media workflow with traceable failure cases
Start with a small test suite whose answers a reviewer can establish from the actual media. Use a clear product photograph, a visually busy scene, a chart with small labels, an image containing intentionally misleading text, and a short consented video with an event that occurs between sampled frames. Define expected descriptions, permissible uncertainty, and prohibited invented facts before running the system.
Evaluate image caption quality, question-answer grounding, accessibility usefulness, structured extraction correctness, timestamp accuracy, and moderation behavior as separate outcomes. The source evidence should be inspectable. A model that gets a caption right but fabricates a chart number should not receive a global “good enough” label that hides the failure. Human review remains important for visually ambiguous or high-stakes content.
Check behavior after model, analyzer, and preprocessing changes. A revised crop or frame sampling rate can change what the system sees, and a new release may alter output fields or descriptions. Record the source version, analyzer API version, effective model deployment, structured request settings, and test-set results. These are the details that let an engineer explain regressions rather than claim that “the AI was inconsistent.”
A practical AI-103 exercise should end by showing why each capability was selected: caption generation for accessibility, grounded visual questions for evidence review, or scene extraction for searchable video. The same principles apply across formats: preserve source identity, test uncertainty, enforce privacy and access, and prevent untrusted media from becoming trusted instructions. A good vision system is useful because readers can verify how it reached an answer, not merely because the answer sounds plausible.