Category Archives: AI & Machine Learning
Agent-to-agent communication becomes difficult when every team invents its own invocation contract. One agent exposes a REST endpoint, another expects a function call, a third returns a proprietary task object, and none of them describe capabilities in the same way. The Agent2Agent protocol is meant to reduce that friction by giving agents a standard way […]
Conversation state is what lets an agent behave as though the second turn belongs to the first. It sounds simple until the application has to decide which messages belong together, how long they should persist, what happens when two requests arrive at once, and whether the conversation is allowed to contain sensitive data. In Microsoft […]
Retries are one of the easiest ways to make an agent system look more reliable while quietly making it less safe. A transient 429 from a model deployment may deserve another attempt. A malformed tool request does not. A timeout after a payment API call is especially dangerous because the call may have completed even […]
A shared agent endpoint does not imply shared user state. In a production Microsoft Foundry deployment, session isolation is the boundary that keeps one caller’s conversations, files, runtime state, and stored data from appearing in another caller’s workspace. This matters most in enterprise and multi-tenant systems, where the agent may look like one service from […]
Engineering with Claude is less about finding one perfect prompt and more about designing a dependable system around a probabilistic model. A production application has to decide what Claude may see, which tools it can call, how state is carried across turns, what happens when an API request fails, how much a run may cost, […]
Semantic caching changes the economics of an AI gateway by allowing it to reuse an earlier model response for a new prompt that is similar in meaning, not merely identical in text. Azure API Management supports this pattern with LLM semantic cache lookup and store policies backed by an external vector-capable cache. The appeal is […]
A production agent loop is not simply a model call repeated until the answer looks finished. It is the control system that decides when Claude can inspect data, call a tool, change state, ask for input, recover from failure, and stop. The Anthropic Agent SDK matters because it packages the same general agent machinery used […]
Large language model capacity is usually measured in tokens, not only requests. Two API calls can consume radically different amounts of model capacity because one contains a short question and the other includes a long document, several tool definitions, and a large response. Azure API Management addresses this mismatch with an LLM token-limit policy that […]
Reliable Claude API integrations treat errors as part of the protocol, not as exceptional surprises. The practical question is not whether a request can fail, but whether the application can tell the difference between a malformed request, a permission problem, a rate limit, a temporary service condition, and a failure that happened after a streaming […]
AI transformation becomes difficult to defend when the organization can describe what it deployed but cannot explain what changed. Agent counts, prompt volume, token spend, and Copilot licenses are activity metrics. They do not by themselves prove business value. Microsoft’s current guidance on agent ROI starts from a more disciplined premise: define value before building, […]
Idempotency becomes important the moment a Claude-powered workflow can be retried. A network timeout, worker restart, queue redelivery, or user double-submit can cause the same logical job to reach the application more than once. If the workflow only generates text, that may produce duplicate cost or duplicate output. If the workflow can call tools that […]
When a company has one AI application, it is easy to connect that application directly to a model endpoint. The design changes when dozens of applications, agents, teams, and tenants share model deployments. They need a common place to authenticate callers, allocate capacity, observe consumption, enforce policy, route traffic, and protect backends from bursts. Azure […]
Claude API rate limits are easier to operate when they are treated as capacity signals rather than as arbitrary request failures. Anthropic separates monthly spend controls from rate limits, and the API can constrain both request frequency and token throughput. A system that watches only requests per minute can therefore look healthy while still exhausting […]
Streaming changes a Claude integration from “wait for one response” into a sequence of events that the application must interpret correctly. Anthropic’s Messages API uses server-sent events (SSE) when streaming is enabled. That lets text, tool-use information, and other deltas arrive incrementally, which can reduce perceived latency and keep long-running HTTP connections active. It also […]
Audit logging for Claude is not one feature with one data source. Anthropic now exposes several layers of operational and governance data, and they answer different questions. The Compliance API provides per-event activity records for security, legal, and compliance workflows. Analytics APIs provide aggregated usage and adoption metrics. OpenTelemetry can stream detailed runtime telemetry from […]