Audit logging for Claude is not one feature with one data source. Anthropic now exposes several layers of operational and governance data, and they answer different questions. The Compliance API provides per-event activity records for security, legal, and compliance workflows. Analytics APIs provide aggregated usage and adoption metrics. OpenTelemetry can stream detailed runtime telemetry from products such as Claude Code. Application teams also need their own logs around Claude API calls and tool execution.
The design challenge is deciding which events must be reconstructable after an incident. A security team may need to know who performed an administrative action, which organization was affected, and which session generated a change. A FinOps team may need aggregated cost by workspace. An application owner may need the request ID, latency, model, and tool sequence for one failed agent run. Treating those needs as one generic “AI log” creates either too little evidence or an expensive pile of sensitive data.
A mature Claude engineering platform defines the audit model before incidents force the organization to discover what was never recorded.
The Compliance API is the event-level governance layer
Anthropic’s Compliance API provides programmatic access to organization activity for eligible Claude Enterprise and Claude Console configurations. Its Activity Feed covers authentication, chat, file, project, administrative, and platform activity. Current documentation says activity is queryable shortly after it occurs and retained for a long period, making the feed suitable for security investigation and SIEM ingestion.
The API also distinguishes access by key type and scope. A Compliance Access Key can unlock broader Claude Enterprise content and organization data depending on scopes, while an Admin API key can access the Activity Feed in supported Claude Console scenarios. This is important because an audit integration should use the minimum scope required rather than treating a powerful compliance credential as an ordinary application key.
Anthropic also records access to the Compliance API itself. That is a valuable property: the system used to investigate user activity also leaves evidence about who queried the audit data. Security teams should ingest those events rather than excluding them as internal noise.
Analytics data is not a substitute for an audit trail
Aggregated analytics answer questions such as adoption, usage, productivity, and cost. They are valuable for operations and governance, but an aggregate cannot normally reconstruct the exact sequence of user or administrative actions that led to an incident. The data model is intentionally different.
For Claude Code, Anthropic provides analytics capabilities that can report daily productivity metrics. Claude Enterprise analytics can cover broader engagement and cost across Claude products. Those sources are useful for trends: who is adopting the tool, whether usage is growing, which teams are consuming spend, or whether a rollout is active.
Security investigations need event-level evidence. FinOps needs cost attribution. Engineering needs request and tool traces. Keeping those purposes separate prevents a common failure where one logging pipeline tries to serve every audience and ends up being too coarse for audit but too sensitive for routine dashboards.
Application logs should preserve correlation without copying every prompt
An application that calls the Claude API should log enough metadata to connect a user request to the model interaction and downstream effects. Useful fields can include the application trace ID, Anthropic request ID, timestamp, model, workspace or tenant, operation type, latency, attempt number, token usage where available, and final status. Agent systems should additionally record tool names, tool outcome, and the application’s own operation identifiers.
Raw prompt and response logging should be a deliberate choice, not a default. Prompts can contain customer data, source code, credentials accidentally pasted by users, or regulated information. If full content is required for evaluation or investigation, access controls and retention should match the sensitivity of the source system.
Often the better pattern is layered retention: keep high-level metadata broadly, retain content only for approved workflows, and use hashes or references to the system of record where direct duplication is unnecessary. This aligns with private-data governance because observability systems can become an overlooked secondary data store.
SIEM ingestion needs stable identifiers and cursor discipline
A durable compliance integration should treat the Activity Feed like any other event source. Consume it incrementally, persist the cursor or equivalent checkpoint, and make ingestion idempotent so a retry does not create duplicate alerts. Events should be normalized into the organization’s security schema while preserving the original event ID and actor identifiers.
Anthropic recommends stable user identifiers rather than email addresses as primary join keys. That is good general practice because email addresses can change. The SIEM can enrich the stable identifier with current directory information without rewriting historical evidence.
Correlation is where the audit trail becomes useful. A Claude event can be connected to identity-provider logs, repository activity, ticket changes, or application traces. A suspicious action then becomes part of a timeline rather than an isolated AI event.
Claude Code and agent telemetry require a runtime view as well
Compliance events answer who did what at the product level, but agent operators also need the internal run. Claude Code and Agent SDK workloads can expose OpenTelemetry-style data that includes operational metrics such as token use, cost, and host or session context. That data can reveal slow tools, repeated retries, expensive workflows, or unusually long sessions.
This is the domain of AI observability. A successful API status does not mean the agent behaved efficiently. An audit event that says a session occurred does not explain whether the agent called the same tool ten times. Runtime telemetry and governance events should be correlated, not conflated.
For high-value workflows, the application may also maintain an append-only decision record that captures the business operation, approvals, agent result, and external state change. That can be more useful to auditors than a raw token-by-token transcript because it expresses the business meaning of the run.
Retention should be designed from the investigation backwards
Anthropic documents specific retention behavior for Compliance API records and certain session data, while customer-controlled data can be subject to organization retention settings. An enterprise should not simply mirror everything forever “just in case.” It should decide what investigations and legal obligations require, then retain the minimum evidence that satisfies them.
Different records can have different horizons. Security events may need years. High-volume debug traces may need weeks. Raw prompts may need much shorter periods or no centralized retention at all. Aggregated cost metrics can often be kept longer because they are less sensitive than conversation content.
Deletion behavior must also be understood. If an organization relies on Anthropic as the only source of retained evidence, it should know which records can be deleted by users or policy and which remain in the activity history. If it copies data to its own SIEM, that copy becomes subject to the organization’s own privacy and access controls.
Audit alerts should be tied to meaningful behavior
A logging program is incomplete if no one knows what to alert on. Useful detections may include unusual administrative changes, new compliance-key creation, suspicious access patterns, large changes in agent usage, repeated authorization failures, unexpected tool activity, or an application suddenly operating outside its normal workspace.
The threshold should reflect context. A developer using Claude Code heavily during a migration is different from a service account that suddenly begins accessing many repositories. Behavioral baselines are more useful than one global number for every actor.
Security governance is strongest when evidence is connected to accountability. An alert should identify who can investigate, which system owns the data, and what response is appropriate.
An audit schema should make the event understandable without requiring an investigator to reconstruct the entire application. Useful fields normally include the actor or service identity, organization or workspace context, action, target resource, result, timestamp, and a stable correlation identifier. Where the API exposes a request identifier, preserving it beside the application trace creates a clean bridge between an internal incident record and provider-side support evidence.
Content and evidence should be separated deliberately. Security teams may need proof that a file was uploaded, a project setting changed, or a model request occurred without retaining every prompt and response in the SIEM. Storing unnecessary content expands privacy exposure and makes access controls harder. A better design records the event metadata needed for accountability and keeps sensitive conversation content in the system that already governs it.
Identity normalization deserves the same care. Anthropic’s compliance integration guidance recommends stable user identifiers for joins rather than relying on mutable display attributes such as email alone. That matters during account changes, mergers, and investigations that span several identity systems. If the audit pipeline cannot reliably answer who an event belonged to at the time, long retention does not create useful evidence.
Good audit logging makes Claude activity explainable
The goal is not maximum log volume. It is to make important actions reconstructable without turning the logging platform into an uncontrolled archive of sensitive AI content. Compliance events, analytics, runtime telemetry, and application logs each have a distinct role. Used together, they can answer who acted, what the platform recorded, what the agent actually did, what it cost, and what external state changed.
That structure also makes routine operations easier. Production monitoring can move from a symptom to a specific request and session. Security can correlate that session with identity activity. FinOps can attribute the cost. Governance can verify that retention and access policies were followed. Auditability becomes part of the architecture instead of a forensic afterthought.