Video generation is not a longer version of image generation. An image can look convincing in one frame while a video becomes unusable because an object changes shape, the camera jumps unexpectedly, or motion loses continuity. Video also adds duration, audio, editing, asynchronous processing, and much larger files to the application design. For the computer vision portion of Microsoft AI-103, engineers need to decide how a video is requested, produced, revised, reviewed, and delivered—not merely recognize that an AI model can produce clips.
Microsoft Foundry documents Sora 2 video generation as a preview capability, including text-to-video, image-to-video, and remixing supported generated videos. Preview documentation and model availability must be checked before a project commits to this route: regions, entitlement, service terms, supported inputs, model deployment names, and content restrictions can change. A practice article can demonstrate the architectural workflow without claiming that an account has access or that a sample clip was actually generated.
Identify the video’s real job before selecting a model
A business may want a short animated product concept, a visual explanation of a process, or a stylized introductory scene. Those tasks have different tolerance for invention. A conceptual animation can benefit from creative motion, whereas a security training diagram that shows a physical lock or machine procedure needs exact steps and labels. A generative video that changes the equipment geometry between frames could teach the wrong procedure. For factual demonstrations or regulatory training, conventional footage, verified animation, or human-authored compositing may be more suitable.
Define the intended audience, distribution channel, approved subject matter, target aspect ratio and duration, language or audio requirements, and which elements must remain consistent. Clarify whether the request starts from a text brief, a licensed still image, or an existing video originally produced by the same system. This distinction matters: a still reference asks the model to invent motion from visual evidence, while a remix starts from a generated clip and asks for a controlled change. The API support and rights to use the input media must both be verified.
For example, an internal learning team might request an eight-second stylized scene of a paper boat crossing an illustrated sea. Its purpose is to test motion consistency and visual composition, not claim an accurate depiction of a real person or product. The requirements might say that the boat remains recognizably the same shape, the camera does not jump, and the movement is calm. Such a neutral synthetic task is useful for a lab because it does not need private customer media or a real person’s likeness.
The broader multimodal AI design questions still apply: why is the modality useful, what evidence does the input carry, how much source media must be retained, and who reviews the output? Video adds temporal coherence as an additional engineering requirement. Two frames can each be plausible while the transition between them is wrong.
Check preview conditions, permissions, and deployment boundaries
Do not build a production release on a preview capability without accepting the preview’s support and service-level boundaries. Microsoft describes Sora 2 within Azure OpenAI and Foundry documentation, but deployment name, available model version, region, and authorization can differ by subscription. Some content categories or uses are restricted. Confirm those conditions on the date the service is configured and record which endpoint and deployment have actually been provisioned. An exam learner should not infer that every Foundry resource supports every video operation.
Separate the identity that invokes the model from the user who submitted the prompt. The application can authenticate to Azure, yet still need its own policy deciding whose media may be uploaded, who may request a costly job, and who can retrieve a completed clip. Give each request a stable job record with a requester identifier, permitted purpose, model deployment, content policy classification, source-media provenance, and retention rule. The raw input itself should not be copied into broadly accessible logs.
Before an invocation, validate the request against current model limits. Duration, dimensions, input reference size, accepted media format, and account throughput are provider-specific. A platform may support a given option only in certain deployments or versions, so the validation layer should be based on a verified capability list rather than an assumed universal default. If a video request violates a policy or cannot fit the supported dimensions, explain the limitation without repeatedly sending the same invalid request.
Another important boundary is the difference between a locally stored file and a hosted generation job. A model request may return a job identifier long before any video exists. That identifier must be protected like any other resource reference tied to a user’s work. The application should not allow one customer’s job ID to fetch another customer’s video merely because an API exposes a download operation.
Build the asynchronous create, poll, and download lifecycle
Microsoft’s documented Sora 2 workflow starts a video job, receives its identifier, checks progress, and downloads media only after successful completion. A frontend should not hold an ordinary HTTP request open indefinitely while the model creates a clip. A backend job controller can persist state and present a separate progress view. The essential states include queued or processing, completed, failed, and cancelled, although exact provider fields must be checked against the current interface. A successful create response is an accepted job request, not evidence that the resulting video passed quality checks.
With an authorized Azure OpenAI resource that has a supported Sora 2 deployment, a representative Python example follows the documented OpenAI-compatible Azure v1 video interface. Configure AZURE_OPENAI_BASE_URL with the resource’s /openai/v1/ endpoint, store an API key in AZURE_OPENAI_API_KEY, and use the actual video deployment name in SORA_DEPLOYMENT. Install a compatible openai Python package first. The sample uses an invented, non-sensitive scene; no clip has been generated for this draft.
import os
import time
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AZURE_OPENAI_API_KEY"],
base_url=os.environ["AZURE_OPENAI_BASE_URL"],
)
video = client.videos.create(
model=os.environ["SORA_DEPLOYMENT"],
prompt="A stylized paper boat sailing across a calm illustrated sea",
size="1280x720",
seconds=8,
)
deadline = time.monotonic() + 900
while video.status not in ("completed", "failed", "cancelled"):
if time.monotonic() > deadline:
raise TimeoutError(f"Video job still unfinished: {video.id}")
time.sleep(20)
video = client.videos.retrieve(video.id)
if video.status != "completed":
raise RuntimeError(f"Video job ended in state: {video.status}")
client.videos.download_content(
video.id, variant="video"
).write_to_file("clip.mp4")
The timeout here is an application safeguard, not a guarantee about service completion time. A job can still finish on the provider after a local poller stops. In a full application, save the job ID first so the task can resume safely without creating duplicate paid generations. Persist terminal failure details without dumping raw sensitive prompts into logs. A cancelled state should not be mistaken for a successfully produced asset, and an application restart should not erase the ability to retrieve a completed job.
Download only the completed output, inspect the actual file type, and store it with the right ownership permissions. A download URL or returned object may have a limited lifetime. The system should not assume a video can be fetched indefinitely from the generation endpoint. If a user-facing preview requires a thumbnail or adaptive streaming derivative, create it in a controlled post-processing step after the original artifact has been safely stored.
Use reference images and generated-video remix carefully
An image reference can help specify composition or subject appearance, but it cannot promise perfect identity preservation through motion. Before submitting a reference image, verify licensing, privacy, supported dimensions, and input constraints. Start with an innocuous reference that has clear shapes and controlled lighting. The model must infer movement, occluded details, and frames that were absent in the still. The resulting clip may therefore introduce unexpected design choices even when the first frame resembles the source closely.
Editing an existing generated video introduces a different relationship. In the documented Sora 2 remix operation, the caller references the identifier of an earlier generated video and requests a targeted change through a new prompt. The underlying clip supplies continuity that a fresh text-only request lacks. This is not the same as importing any arbitrary public video file for frame-level editing, and it should not be described as a universal post-production editor.
One effective editing experiment changes only the color palette of the paper-boat clip while preserving the camera path and boat motion. The prompt could request muted teal and warm yellow while retaining the same scene. Compare the resulting clip with the original across several time points; inspect whether the boat shape, horizon, framing, and apparent speed remain sufficiently consistent. A remix that changes all these aspects may fail the assignment even when its colors match the request.
Keep version history explicit: parent video ID, remix job ID, input prompt, requested change, generated output ID, review outcome, and whether the new version became approved. This lineage matters for troubleshooting and rights control. An asset that was approved under a particular campaign policy should not silently be replaced when a new generated variation arrives. Approval should follow the version of the actual media, not the text prompt alone.
Do not assume an image mask and a video remix use the same semantics. Image edits can describe spatial regions within one source frame. Video editing must preserve or intentionally change temporal relationships across many frames. The request shapes, limitations, and acceptance tests are different; copying an image editing parameter into a video request is not a valid implementation strategy.
Design safety and intellectual-property review into the workflow
Generated video has risk beyond single-frame content. It can suggest that an event occurred, depict an identifiable individual, or make a branded product appear to perform an action that it never performed. The application should avoid unauthorized likenesses, deceptive representations, unlicensed source materials, and restricted content. Microsoft documents content safeguards for Sora 2 preview and notes content categories that the service blocks. Treat those documented restrictions as enforceable conditions for a preview service rather than inventing techniques to bypass them.
At intake, check whether the prompt or reference media can be processed under the organization’s policy. For example, a team should not submit confidential customer footage just to test a new model’s animation quality. Use synthetic or properly licensed references wherever possible. Do not rely on the model’s own safety filter as the only rights or consent check. A request may be technically accepted while still being inappropriate for the organization’s intended use.
At output, decide whether content needs a visible synthetic-media disclosure, a watermark, provenance metadata, or manual signoff before publication. Those requirements depend on the use case, jurisdiction, platform, and organizational policy. If a training clip depicts a procedure, have a knowledgeable reviewer confirm that the steps and equipment remain correct through the entire sequence. If a marketing clip depicts a product, confirm it does not invent unsupported features or a real customer endorsement.
Review audio separately if the selected model generates it. Speech or background sound can create unintended statements even when the visual scene is harmless. Captions and accessibility descriptions should be based on what the final clip actually contains, not what the prompt requested. A helpful output may still need synchronized captions, non-audio alternatives, or a human-created descriptive summary.
Evaluate motion, temporal consistency, and user value
Video quality must be judged over time. Inspect object continuity, camera motion, transitions, lighting consistency, physical plausibility, timing, and the presence of visual artifacts. In the paper-boat example, watch whether the boat retains its outline across frames, whether the water behaves coherently, and whether a movement requested as calm becomes unexpectedly abrupt. A representative set of clips should include difficult motions, occlusion, camera movement, and changing light—not just short demonstrations that happen to succeed.
Use separate tests for task adherence, safety, technical quality, and audience usefulness. Correct output dimensions and playable encoding can be checked automatically. Temporal stability might be screened through computer-vision comparisons, but aesthetic coherence and truthfulness often require reviewers. An automatic score should not conceal that a generated demonstration invents a critical step or transforms a product beyond recognition.
For each candidate model or prompt revision, use the same evaluation set and record the observed differences. An improvement in one scene may worsen another. Monitor how often the clip requires remix, how many attempts succeed, how long users wait, and how many outputs are rejected for content policy or visual failure. These measures are closer to product usefulness than counting generated seconds alone.
Failure analysis should be specific. A bad video can result from poor prompt requirements, unsupported reference media, an unstable model response, restrictive content policy, or a post-processing error. Correcting each demands a different action. A prompt improvement cannot repair a failed authentication request; increasing retry volume does not solve a policy violation; and a successfully generated clip is not enough if the player cannot decode it.
Manage operational cost and delivery reliability
Video creation is usually more expensive and slower than text inference, and Microsoft describes Sora 2 billing on a per-second basis under applicable service terms. The actual bill depends on current pricing, selected model, output configuration, number of attempts, and repeated editing jobs. Require the application to show users the expected workload and implement per-user quotas where appropriate. A batch of failed or repeatedly remixed videos can consume resources without producing a useful deliverable.
A queue should enforce concurrency and a maximum cost or attempt policy. Each generation needs a lifecycle record that distinguishes original request, retry, resumption of an existing job, and intentional new variation. When a polling worker fails, resume its existing job reference instead of automatically issuing another video request. Distinguish service throttling from temporary network failure and from a job that has already moved to a terminal state.
The same end-to-end telemetry architecture used for agent workflows applies to video. A correlation ID can connect authorization, the submitted request, provider job ID, progress checks, moderation outcome, saved output, and review. Avoid storing the complete media in general logging systems. Track file storage, transfer, playback and deletion separately from model execution; user-visible success requires an accessible and approved clip, not merely a completed backend job.
Plan for preview lifecycle changes. A future model update might alter available sizes, accepted inputs, remix behavior, or content rules. Keep those capabilities out of hard-coded client assumptions. Revalidate the documentation, test a representative video suite, compare costs and moderation behavior, and retain a way to stop or roll back the application’s use of a preview option when requirements change.
Practice the AI-103 decision process without inventing results
A useful AI-103 exercise starts with a clear fictional scene and written requirements. Choose an allowed model and deployment, identify its supported request and output formats, then sketch the application components: intake validation, job submission, durable status tracking, terminal-state retrieval, artifact storage, and review. Explain how an image reference would change the input contract and how a remix would depend on an already completed generated video.
Use source-approved test media and a subscription you are actually authorized to operate. If that environment is unavailable, work through request shapes and failure states offline; do not claim a service call occurred. Evaluate the planned result across visual continuity, policy, cost, accessibility, and retention. The most valuable lesson is that a generative video is an asynchronous, versioned, potentially sensitive software artifact. Treating it as a one-line prompt misses the design decisions the current computer vision domain is intended to assess.