Microsoft Foundry now provides first-party CI/CD patterns for hosted agents using Azure Developer CLI (azd), GitHub Actions or Azure DevOps. Current guidance supports azd pipeline config, GitHub OpenID Connect authentication, azd provision/azd deploy, environment-specific Foundry projects, agent versioning, smoke tests and azd ai agent eval run as a regression gate. Infrastructure can be managed with Bicep or Terraform while agent versions remain deployed through the Foundry/azd lifecycle.
Within Microsoft AI Agents, CI/CD should treat agent source, infrastructure, model deployment, configuration, tools, evaluations and production endpoint routing as one release system.
Keep source and deployment configuration in Git
Store agent code, azure.yaml, infrastructure templates, dependency manifests, eval configuration and pipeline files together.
Do not make the Foundry portal the only source of truth for production settings.
A deployment should be reproducible from a commit and environment configuration.
Use azd for provisioning and deployment
Current Foundry guidance uses azd up for first-time provision plus deploy, azd provision for infrastructure and azd deploy for agent code/version deployment.
This separation lets code releases avoid re-provisioning infrastructure unnecessarily.
Pin the Foundry extension/tool versions in CI where practical.
Configure CI with federated identity
Current Microsoft templates use GitHub OIDC / federated credentials rather than long-lived Azure secrets.
Assign only the roles the pipeline needs on the target Foundry project/resource group.
Keep application runtime identity separate from deployment-pipeline identity.
Separate development, test and production projects
Current promotion guidance recommends distinct environment bindings rather than relying on environment names alone to isolate resources.
Use separate Foundry projects/resources where production isolation matters.
Promote the same source revision through environments so failures are not mixed with code drift.
Deploy code changes as new agent versions
Hosted agent deployments create versions.
Previous versions can remain available for rollback or pinned endpoint routing.
Do not overwrite the only production revision without a tested return path.
Smoke-test every deployment
Microsoft’s current GitHub Actions quickstart deploys the hosted agent, reads status and invokes a safe test prompt.
Add checks for expected output shape, tool availability and environment identity.
A successful infrastructure deployment is not proof the agent can actually answer or call its dependencies.
Run evals in CI
Current azd ai agent eval run can run stored evaluation configuration in CI.
Set release thresholds for important quality/safety metrics and fail the pipeline when candidate behavior regresses.
Keep eval datasets/versioning independent from a deployment so a failing candidate cannot rewrite the test it is judged against.
Manage infrastructure as code separately from agent versions
Foundry project initialization supports Bicep or Terraform infrastructure choices.
Infrastructure state and agent version state are related but not identical.
Use IaC for resource topology and azd deploy/Foundry agent lifecycle for application versions.
Secrets should stay out of repository configuration
Store deployment variables and secrets in GitHub/Azure DevOps secure stores or Azure-managed secret systems.
Map them into azd environments deliberately.
Never commit model/provider keys, MCP credentials or production connection secrets into azure.yaml.
Production rollout should be gradual
Current Microsoft guidance supports promoting hosted agent versions across environments and gradually routing production traffic.
Canary the new version, monitor evaluations and runtime telemetry, then increase share.
Keep the previous known-good version ready for fast route rollback.
Foundry CI/CD succeeds when the pipeline tests behavior, not just deployment
The mature pipeline provisions through IaC, authenticates with federated identity, deploys versioned agents, isolates environments, smoke-tests endpoints, runs eval gates, protects secrets and supports gradual promotion/rollback.
The release artifact is the behavior of the deployed agent plus its infrastructure and tools—not merely a container or source ZIP.
Pipeline design should separate CI from CD. Pull requests should run unit tests, linting, security scans and offline/preview evals without touching production. Merges can deploy to a development or staging Foundry project, then promotion to production should require the appropriate quality/security approvals for the application’s risk.
Foundry project endpoint and environment variables should be explicit in every job. Current `azd ai` commands can resolve project context from flags, active azd environment, global config or environment variable. CI should set one deterministic source so a runner cannot accidentally target the developer’s default project.
Infrastructure and agent code should have independent drift checks. Terraform/Bicep can show whether Azure resources match desired state, while `azd ai agent show` can confirm the deployed agent version/configuration. A project can be infrastructure-clean but still serve an unexpected agent version if endpoint routing changed manually.
Production endpoints should be pinned during candidate deployment when necessary. Current hosted-agent deployment creates new versions and endpoint behavior can follow latest unless configured otherwise. Keep production pinned to the known-good version, deploy the candidate, evaluate it, then update routing deliberately.
Pipeline smoke tests should exercise dependencies, not only return non-empty text. If the agent needs a search index, MCP server, Azure Function or database, include one safe test proving the integration is reachable with the deployed identity. This catches missing RBAC and secret/config problems before users do.
Eval thresholds should be versioned beside the eval configuration. A team can accidentally make a failing release pass by lowering the threshold in the same change. Protect critical eval configs with code-owner review and report baseline-versus-candidate metrics in the pipeline artifact.
Secrets and model/deployment names should be separated. Non-sensitive IDs/endpoints can be GitHub Actions variables; credentials and API keys belong in secrets or managed identities. Prefer OIDC/federated auth to long-lived service-principal secrets wherever supported.
Environment promotion should reuse the same source revision but allow environment-specific infrastructure and model deployment names. Production may use a different region, SKU, private networking or model capacity. Keep these differences parameterized rather than forking the codebase.
Rollback should be tested periodically. Re-route to the previous agent version, verify the endpoint, and ensure old dependencies/secrets still work. A rollback plan that has never been exercised can fail because the previous version expects an index, connection or secret that was already removed.
Pipeline observability should retain deployment logs, agent version, commit SHA, Foundry project, model deployment, eval results and smoke-test evidence. This gives incident responders a complete answer to ‘what changed?’ without reconstructing state from the portal.
Foundry CI/CD should keep generated cloud state out of the repository. Local `.azure` environment state and temporary credentials should not be committed. Store only declarative `azure.yaml`, IaC, workflow and source files, while CI reconstructs environment state from secure variables and cloud resources.
Pipeline permissions should be split by stage where risk justifies it. A PR-validation identity may need no deployment rights; staging deployment may use one federated identity; production promotion may require a separate protected environment with reviewer approval. Least privilege reduces the blast radius of a compromised workflow.
Source-code and container deployment modes should be chosen intentionally. Current Foundry supports direct source ZIP deployment for supported runtimes and container-based deployment for more control. The CI pipeline should test the exact mode used in production rather than a different local shortcut.
Deployment artifacts should include dependency lockfiles and runtime version. Remote build is convenient, but reproducibility depends on resolving the same dependency set. Pin versions and review build logs so a package update does not change agent behavior without a source-code diff.
Operational monitoring should continue after the pipeline turns green. Deployment success is the start of the observation window, not the end. Track agent errors, tool failures, latency, token/cost, eval/feedback signals and version-specific incidents during canary promotion.
Use Agent Lifecycle Management to define promotion ownership beyond the pipeline mechanics. CI/CD can deploy quickly, but a governed release still needs an owner, risk classification, change record and rollback criteria.
Model deployment changes should be isolated from agent source changes where possible. If the pipeline updates both code and the underlying model version simultaneously, a behavior regression is harder to diagnose. Promote the code against a pinned model first, then evaluate a model upgrade as a separate release.
Quota and capacity checks should run before production rollout. A staging agent may pass smoke tests at low volume while the production model deployment lacks enough throughput for real traffic. Include a lightweight load or capacity validation step and monitor throttling during canary rollout.
Use Azure OpenAI Model Versioning as part of the deployment manifest so each hosted-agent version records the exact model deployment/version it expects. This prevents a mutable shared model deployment from changing agent behavior outside the agent pipeline.
CI/CD itself should be monitored. Alert on repeated failed deployments, stale production versions, drift between main and deployed commit, failed eval gates and manual portal changes after deployment. A pipeline is only a control if teams can detect when production bypasses it.
Release records should capture who approved promotion, the source commit, infrastructure revision, agent version, model deployment, eval suite version, smoke-test result and production routing change. This makes a Foundry incident traceable to one concrete release rather than a sequence of portal and pipeline events spread across several systems.
Keep pipeline ownership and production promotion authority explicit so emergency changes remain auditable and reversible.
Behavioral tests should cover prompt templates, retrieval quality, tool permissions, structured outputs, safety rules, and representative failure cases. A pipeline that only proves the infrastructure deployed successfully can still promote an application whose user-visible behavior regressed.