When a Microsoft Foundry project does not behave as expected, changing the prompt is often the wrong first move. The failure may sit in resource configuration, identity, networking, model deployment, connection setup, tool permissions, quota, or the transition between older hub-based patterns and the newer Foundry resource model. A disciplined troubleshooting path isolates the layer before changing the system.
The current AI-103 blueprint explicitly includes planning and managing Azure AI solutions, selecting Foundry services, setting up projects, deployments, integrations, and agent solutions. That makes project design more than portal setup. It is the operational foundation that determines which resources an application can reach, which identities are trusted, and which failures can be diagnosed cleanly.
Microsoft now uses the Microsoft Foundry name for the unified platform, while older documentation and deployed environments may still use Azure AI Foundry or hub-based terminology. Troubleshooting has to recognize both models without assuming they are identical. The most reliable approach is to begin with the exact resource and project type that is failing, then move outward one dependency at a time.
Confirm the resource model before diagnosing anything else
New Foundry projects use a more unified resource model than earlier hub-based projects. That difference affects endpoints, SDKs, role names, connections, and where configuration lives. If an engineer copies instructions from a classic project into a new project without checking the model, an apparently reasonable configuration can fail because the underlying assumptions are different.
Start by recording the project type, resource identifier, region, endpoint, SDK or API version, and whether any classic resources are still involved. This simple inventory prevents hours of debugging at the wrong layer. It also makes migration-related defects easier to distinguish from ordinary permission or network failures.
Separate authentication failure from authorization failure
Authentication asks whether the caller has a valid identity; authorization asks whether that identity is allowed to perform the requested operation. The symptoms can look similar in an application, especially when SDKs reduce the underlying error to a generic failure. Test the identity explicitly and identify whether the caller is a developer, workload identity, shared project identity, or agent-specific identity.
Then verify role assignment at the scope the resource actually uses. A broad subscription role is not always equivalent to the required data-plane or Foundry role, and a role assigned to a developer does not automatically transfer to a published agent. The Entra identity mental model is useful because it forces the investigation to name the principal, resource, permission, and scope instead of treating “access” as one setting.
Check region, model deployment, and quota before changing application code
A Foundry project can be healthy while the selected model is unavailable in the region, not deployed, rate-limited, or configured with insufficient capacity. Verify the model deployment name used by the application, the region that hosts it, quota and rate-limit status, and whether the request matches the supported API or model capabilities.
This is especially important in multi-environment deployments where development and production may use different model names or capacity reservations. A configuration value copied from one environment can produce failures that look like code defects. Establishing deployment health independently of the application narrows the fault domain quickly.
Test connections as independent dependencies
Projects often connect to search, storage, databases, APIs, and other Azure resources. Treat each connection as its own dependency with an endpoint, authentication method, network path, and permission set. First prove that the target service is healthy. Then prove that the project or agent identity can reach it. Only after those checks should you investigate orchestration or prompt behavior.
Private networking adds another layer because name resolution, private endpoints, firewall rules, and identity can all be correct or incorrect independently. A connection that works from a developer laptop does not prove that the hosted agent can reach the same target. Test from the runtime path that matters.
Tool failures should be debugged as contracts, not conversations
When an agent calls a tool, inspect the requested function name, arguments, schema, authentication, response, and timeout. The natural-language conversation is secondary. A malformed schema can cause the model to generate invalid arguments; a valid call can still fail because the backend rejects a field or because the tool returns an unexpected response shape.
Use a minimal known-good invocation outside the agent where possible. If the function works directly but fails through the agent, the problem is in tool definition, orchestration, or identity propagation. If it fails directly, fix the backend or connection before changing the model. This separation preserves causality.
Retrieval problems belong to the retrieval layer first
A weak grounded answer does not automatically mean the model is hallucinating. Inspect whether the right documents were indexed, whether the query retrieves the expected passages, whether filters remove needed content, and whether the chunk size preserves the information required to answer the question. Retrieval quality should be tested before generation quality.
When Azure AI Search or another retrieval service is involved, evaluate keyword, vector, hybrid, and ranking behavior against representative queries. If the expected evidence never reaches the model, prompt changes can only hide the problem. The system should be able to show what context was retrieved for the failing request.
CI/CD and environment drift deserve their own checks
A deployment pipeline may succeed while leaving an environment with the wrong variable, connection, role assignment, or model deployment. Compare the intended release manifest with the actual production configuration. Look for manual changes, stale secrets, missing permissions, and resources created outside the deployment process.
The lessons from Azure Pipelines and GitHub Actions are relevant because automation is only as reliable as the state it manages. The pipeline should surface drift rather than simply repeat deployment steps against an unknown environment.
Use end-to-end tracing only after the basic layers are known healthy
Once resource, identity, deployment, connection, and tool checks pass, end-to-end tracing becomes powerful. Follow a request through the application, Foundry endpoint, model call, retrieval step, tool invocation, and final response. Correlation IDs and structured logs make it possible to distinguish latency, retry, policy, and data-quality problems.
The adjacent AI-300 path becomes useful at this point because operational AI requires monitoring, evaluation, and production diagnostics. Troubleshooting should leave behind observability improvements so the next incident produces better evidence than the last one.
Close the incident by proving recovery at the failed layer
A fix is not complete because one prompt worked after a change. Re-run the smallest test that reproduces the original failure, then run representative end-to-end scenarios. Confirm that authorization boundaries still hold, that no other environment was unintentionally changed, and that monitoring shows the expected healthy state. If the defect involved migration from classic terminology or resources, update the runbook so the same ambiguity does not return.
The clean troubleshooting sequence is: identify project model, prove identity, prove resource health, prove connection, prove tool or retrieval contract, trace end to end, and validate recovery. That order avoids random configuration changes and makes Microsoft Foundry projects understandable as systems rather than as portal screens.
Do not ignore DNS and endpoint resolution when private networking is involved. A private endpoint can exist correctly while the runtime still resolves the public address, or a development machine can resolve a private name that the hosted runtime cannot. Record the hostname the application uses, inspect where it resolves from the failing environment, and compare that result with the intended network design. This simple check often separates a networking defect from an application defect quickly.
A clean incident record should capture the failed request, timestamp, correlation identifier, caller identity, project endpoint, model deployment, relevant connection, and the exact remediation. That record becomes more valuable than a screenshot of the portal because it can be compared across incidents. Over time, recurring failure categories reveal where the project design itself needs simplification rather than another troubleshooting note.
For production systems, define a smoke test that exercises the minimum critical path after every deployment: authenticate, reach the model, retrieve one known item, invoke one safe tool, and emit one trace. A five-minute smoke test can catch configuration drift before users encounter it and gives operators a known-good reference when later symptoms appear.
Permission troubleshooting should include inherited and indirect access. A caller may appear to have the right role through a resource group while a data-plane permission is still missing, or a managed identity may have access to the service but not the specific index, database, or secret it needs. Write down the exact resource path being accessed and test that path with the runtime identity; broad role screenshots can conceal the missing edge.
Troubleshooting documentation should distinguish symptoms from causes. A 403 response, an empty retrieval result, and a tool timeout are observations; they are not diagnoses. Runbooks should list the evidence that separates likely causes and the safest verification step for each. This keeps operators from making broad configuration changes because a familiar error message appeared and reduces the risk of fixing one incident by weakening security or changing unrelated resources.