Agent telemetry needs a trace of decisions, not only request counts
Production agents make sequences of model calls, tool calls, retrieval operations, policy checks, retries, and user-facing responses. A simple success counter cannot explain why a request failed or why latency changed. Microsoft Foundry uses OpenTelemetry-oriented tracing and Application Insights integration to expose that execution path.
For AI-103, observability is part of the application contract. Engineers need enough evidence to separate model behavior, tool failures, retrieval gaps, policy blocks, and platform capacity without logging every sensitive conversation in full.
AI observability starts with correlation. Every request should carry identifiers that connect client activity, agent execution, model calls, tool spans, and downstream services so an incident can be reconstructed across boundaries.
Design telemetry before deployment. Adding trace fields after an incident usually means the exact evidence needed to explain the failure was never captured.
Trace taxonomy should be agreed across teams before dashboards are built. Consistent span names for model calls, retrieval, tools, policy checks, and approvals make cross-agent queries possible and reduce the need for each product to invent its own observability vocabulary.
Sampling strategy should preserve rare high-severity events. Uniform sampling can discard the exact traces needed to investigate data leakage, permission failures, or destructive tool calls, so policy events and severe errors should be retained at a higher rate than routine successful turns.
Server-side and client-side tracing cover different parts of the path
Foundry can emit server-side telemetry for agent invocation, while client-side OpenTelemetry instrumentation captures application code around the managed service. Combining them shows both what the agent platform did and what the surrounding product did before and after the call.
Client traces should include orchestration logic, custom preprocessing, cache decisions, identity checks, and calls to services that Foundry cannot observe directly. Server traces can then connect the managed agent’s model and tool behavior to the broader request.
Propagate trace context across services instead of starting a new disconnected trace at every hop. When a tool triggers another service, preserving the parent-child relationship makes latency and failure analysis far more useful.
Use stable service names and environment attributes so traces from development, staging, and production do not blend together. Operational queries should be able to isolate one release and one deployment boundary quickly.
Client and server clocks can differ, so distributed traces should rely on propagated context and span timing rather than manual timestamp joins. This becomes important when agents call services in different regions or through queues where absolute times are harder to compare.
Some hosted-agent tracing and Insights features remain preview or vary by agent type. Production observability plans should verify the current support status for the chosen runtime and region, and they should retain an independent logging path for signals that the managed tracing surface does not yet guarantee.
Model spans need workload context to be useful
Model telemetry should capture model or deployment identity, latency, token usage where available, response status, retry count, and relevant configuration version. Those fields explain cost and performance changes without requiring operators to inspect raw prompts by default.
GenAI observability ties model signals to application outcomes so operators can judge whether lower latency or token use actually improves the product. A faster model call is not an improvement if it causes more retries, tool errors, or low-quality responses downstream.
Record prompt-template and agent-version identifiers rather than copying entire prompt text into every trace. Version IDs allow comparison and rollback while reducing the amount of sensitive content stored in monitoring systems.
Segment metrics by request class. Long research tasks, short routing decisions, and tool-heavy transactions have different latency and token profiles, so blended averages can hide regressions.
Token metrics should be interpreted with request success. A drop in tokens may indicate efficiency, but it can also mean context was omitted or a response failed early, so cost dashboards should join usage with quality and completion signals.
Tool telemetry should expose both intention and effect
A tool span should identify the tool, normalized operation, latency, result status, retries, and downstream dependency. For sensitive tools, log parameter classes or redacted identifiers rather than full payloads while preserving enough detail to distinguish one operation from another.
Authorization and approval outcomes belong near the tool span. A denied call is not the same as a network failure, and a user-cancelled action is not the same as a tool exception. Collapsing them into one error metric creates noisy incident response.
Agent analytics connects repeated tool patterns to user-visible symptoms. If one connector accounts for most slow turns, operators can fix the actual bottleneck instead of tuning the model.
Track side-effect completion separately from model completion. An agent response can fail after a tool succeeded, or succeed textually while an external action failed. Production telemetry needs to represent both states.
Tool spans should include a stable operation category even when backend endpoints change. That lets historical dashboards remain useful through implementation refactors and keeps business-level monitoring separate from low-level API naming.
Dependency naming should remain stable across retries. If each retry creates an unrelated operation name, dashboards can exaggerate unique failure types and hide that one downstream service is repeatedly causing the same class of incident.
Content recording should be opt-in and purpose-limited
Full prompts, responses, tool payloads, and retrieved documents can be valuable for debugging, but they can also contain personal data, secrets, contracts, or customer records. OpenTelemetry instrumentation should not become an uncontrolled duplicate data lake.
Prefer structural telemetry by default: sizes, identifiers, status, timing, model name, tool name, and version metadata. Enable content capture only for defined scenarios, protected environments, and retention periods.
Redaction should happen before export when possible. Removing secrets later from a monitoring store does not undo earlier exposure to operators, backups, or downstream analytics systems.
Document who can access trace content and why. Observability data often has broader technical access than production databases, so content capture can quietly weaken the original application’s data controls.
Privacy controls should be tested like any other feature. Verify that redaction occurs before export, that privileged users cannot casually turn on full-content capture in production, and that retention jobs actually delete trace content when policy requires it.
Metrics should describe service objectives, not dashboard abundance
Choose metrics that answer operational questions: Is the agent available? Is it getting slower? Are tools failing? Is quality degrading? Are policy blocks increasing? Is cost per successful task rising? A long list of counters without response ownership adds noise.
AI monitoring should connect symptoms to traces. An alert on high latency is most useful when operators can pivot directly to traces that show whether the delay came from a model, retrieval, tool, or client.
Use percentiles for latency and distributions for token usage instead of relying only on averages. Agent traffic is often long-tailed, and a small number of complex turns can dominate user dissatisfaction or spend.
Define service-level indicators around completed user tasks where possible. Model request success is a weak success metric when the user needed a multi-step workflow to finish.
Alerts should point to a runbook and an owner. An alert on agent error rate that nobody knows how to triage simply converts telemetry into noise, while a threshold tied to trace examples and escalation steps shortens recovery time.
Production traces can feed evaluation and optimization
Observability and evaluation should form a loop. When traces reveal a repeated failure pattern, sanitize representative cases and add them to the evaluation set so the next release is tested against the behavior that actually occurred.
Foundry Insights can analyze trace data for recurring behavior, while evaluation services can score quality and safety on selected examples. Keep human review in the loop when the inferred cause or proposed optimization affects high-impact behavior.
Tag releases and experiments so improvements can be attributed. If model, prompt, tool schema, and retrieval configuration all change at once, telemetry may show that performance moved without explaining which change caused it.
Preserve enough history to compare before and after a deployment. Fast-moving agent platforms make it easy to lose the baseline needed to prove that a “better” configuration actually improved production.
Evaluation pipelines should retain links back to production trace identifiers when a sanitized incident becomes a regression test. That lineage helps engineers understand why a case was added without keeping sensitive customer data inside the test fixture.
Telemetry design should make incidents cheaper to resolve
The practical measure of observability is how quickly a team can answer: what happened, who or what was affected, which layer failed, whether an external action occurred, and what changed recently. Trace design should serve those questions directly.
Create incident views that join agent version, model, trace ID, tool status, dependency latency, policy events, and deployment metadata. Operators should not need five portals and manual timestamp matching to reconstruct one conversation.
For Microsoft AI agents, Application Insights and OpenTelemetry provide the technical substrate for traces and metrics. Product teams still need naming conventions, retention policy, privacy controls, dashboards, alerts, and ownership that turn that raw telemetry into an operating system.
Good telemetry reduces the temptation to add speculative logging after every incident. When core identities and spans are designed well, teams can debug new failure modes by querying existing evidence rather than capturing everything forever.
Observability cost should be reviewed as traffic grows. High-cardinality attributes, full-content spans, and long retention can make monitoring expensive, so teams should preserve the fields needed for diagnosis while sampling or aggregating lower-value detail.
Runbooks should include the minimum evidence operators need before escalating to model or platform teams. Trace ID, agent version, model, failed span, tool outcome, and recent deployment history usually provide a far stronger starting point than a screenshot of the user-visible error.