Claude’s Messages API is stateless: the application sends the prior conversational turns in the messages parameter and Claude generates the next turn. Anthropic’s API reference explicitly describes this as stateless multi-turn conversation. By contrast, Claude Managed Agents sessions are a beta higher-level surface that maintain conversation history across multiple interactions. Production architecture should therefore decide whether conversation state lives in the application database around Messages API calls or in a Managed Agents session designed for long-running agent work.
Within Claude Engineering, state is not only chat text. It includes user identity, thread metadata, tool results, retrieved evidence, approvals, business workflow state, model/prompt version, and the subset of history that is actually safe and useful to send back to Claude.
Messages API state belongs to your application
Each Messages request includes the conversation turns Claude should consider.
Store the canonical conversation on your side and construct the request deliberately rather than assuming Anthropic maintains a hidden server-side chat session for the direct API.
This makes retry, audit, deletion, tenant isolation, and replay explicit engineering responsibilities.
Separate canonical history from model context
The durable record may contain every user/assistant/tool event, while the model context should contain only the information needed for the next decision.
Keep raw history for audit/replay as policy permits, and build a context projection that can summarize, trim, or retrieve older content.
Claude Context Editing covers strategies for keeping the active context efficient.
Store structured workflow state outside natural-language turns
Order ID, approval status, account tier, selected project, tool execution result, and retry counter should live in typed application state where possible.
Do not make Claude infer critical state from a sentence written 40 turns ago.
Inject the relevant structured facts into each turn or tool result so the state machine remains deterministic.
Tool results are state transitions
A tool call can create a ticket, change a database, send a message, or retrieve current data.
Record tool call ID, parameters, result, external object ID, idempotency key, and completion status independently from the assistant prose.
Claude API Idempotency is relevant when retries could repeat a side effect.
Summaries should be treated as derived state
Conversation summaries save tokens but can omit details or introduce interpretation.
Keep the raw source events so a summary can be regenerated after prompt/model changes.
Version the summary schema and include explicit fields for commitments, unresolved questions, user preferences, entities, and workflow status rather than one free-form paragraph for everything.
Prompt caching reduces repeated history cost
Long conversations repeatedly resend the system prompt, tool definitions and earlier turns.
Anthropic currently recommends prompt caching as a primary cost/latency lever for repeated context.
Cache stable prefixes while keeping new turns and volatile state outside the cached segment.
Token counting should gate context construction
Use the token-counting endpoint before requests when histories, documents, images or tool schemas are large.
Set a context budget for conversation history, retrieved evidence and tool definitions.
When the budget is exceeded, summarize or retrieve older state rather than blindly truncating the earliest messages, which may contain critical instructions.
Managed Agents sessions are stateful by design
Current Managed Agents beta sessions reference an agent/environment and maintain conversation history across multiple user events.
Use this higher-level model when Anthropic-managed session execution, tools and event streams match the application’s needs.
Do not mix Managed Agent session history and a separate app-side transcript without defining which is authoritative.
Thread and tenant isolation should be explicit
Every conversation should have a stable internal thread/session identifier scoped to the correct tenant/user.
Never select history from a shared cache by user-visible title or loose similarity.
Authorization should happen before state is retrieved, and audit logs should show which conversation record fed each API call.
Deletion and retention must cover derived state
If a user or policy requires deleting a conversation, remove or expire raw messages, summaries, embeddings/retrieval records, tool-output caches and any replicated analytics artifacts according to the relevant data-handling rules.
Anthropic’s API/data-retention options do not replace your application’s own storage policy.
State architecture should make deletion paths knowable before production.
Conversation state succeeds when the model sees the right context without becoming the database
The mature application owns durable thread history, stores workflow facts structurally, records tool transitions, caches repeated prefixes, budgets tokens, derives summaries reproducibly, and uses Managed Agent sessions only when their state model is intentional.
Claude should reason over state supplied to it; critical business state should remain controlled by systems that can validate, query, and audit it deterministically.
Conversation IDs should be immutable internal identifiers while user-visible titles remain editable metadata. Users can rename threads or create identical titles; business logic should never use display text to locate state. Store tenant ID, user/profile ID, application ID and conversation ID together so authorization is checked on every state lookup.
Messages should be stored in a normalized event model that can represent text, images/documents, tool calls, tool results, refusals/errors and system-level metadata. Claude content blocks can evolve beyond plain strings. Keeping the original structured blocks avoids losing information needed for replay, evaluation or migration.
Retries should not append duplicate assistant turns to canonical history. When a network timeout occurs after Anthropic processed a request, the application may not know whether a response was generated. Use request IDs/idempotency where applicable, status reconciliation and deterministic turn IDs so a retry does not create two logical assistant turns.
Branching/editing requires explicit parent relationships. If a user edits an earlier prompt or regenerates an answer, decide whether the application creates a new branch or mutates history. Immutable events plus parent pointers preserve both paths and let eval/debugging reproduce the exact branch that produced a later tool action.
Context selection should prioritize system policy, current user request, unresolved commitments and authoritative tool state over conversational pleasantries. A chronological ‘last N turns’ strategy is simple but can discard the oldest instruction that still controls the workflow. Build context by semantic/structural importance, not only recency.
Retrieved long-term memory should be treated as evidence with provenance. If the app stores preferences or past facts in a vector store, include source thread/event and timestamp, and allow stale/contradicted memory to be superseded. Never let a fuzzy memory retrieval silently override a current explicit user instruction.
Model migrations can change how much history is useful. A new model/context window may handle longer raw history or different summary formats. Keep context construction modular so changing models does not require rewriting the durable conversation database or accepting one summary scheme forever.
Tool approval state should not be inferred from conversation wording. Record an explicit approval object with requested action, scope, approver, timestamp and expiry. If Claude says ‘the user approved this earlier’, the application should still check the structured record before executing a destructive tool.
Streaming responses should become one logical assistant event. Persist partial deltas for resilience if needed, but mark the event incomplete until a terminal message/end state arrives. If a stream disconnects, the UI can show partial output while the canonical state distinguishes complete from interrupted generation.
State metrics should include average context tokens per turn, cache-hit rate, summary frequency, retrieval hit/miss, duplicate/retry events and per-thread storage growth. These reveal when conversation architecture is becoming expensive or losing important information long before users report incoherent behavior.
State compaction should be triggered by semantic milestones as well as token count. Completing an order, closing a support case or finishing a research phase is a natural point to summarize resolved details and keep only the outcome plus unresolved items. This produces cleaner context than repeatedly summarizing an arbitrary moving window.
Conversation state should distinguish user-provided facts from model inferences. Persisting Claude’s inference as if the user stated it can compound errors across turns. Mark provenance and confidence, and require explicit confirmation before promoting inferred high-impact facts into durable business state.
External data fetched by tools should have freshness metadata. Account balance, inventory, ticket status or policy can change after the turn that retrieved it. Store retrieval timestamp/version and refetch when the workflow requires current truth rather than replaying old tool output from the transcript.
Privacy-sensitive state should be minimized. If a tool returns a full record but later turns need only an eligibility flag or ID, persist the minimum derived value where policy permits. Large conversation histories can become accidental data stores containing far more personal or confidential detail than the feature needs.
Session archival is different from deletion. For Managed Agents or app-side threads, archived/inactive state may remain retrievable for audit while being excluded from active UI. Define explicit retention/deletion transitions so ‘close conversation’ does not ambiguously mean hide, archive or erase.
Conversation migration tests should replay representative historical threads when changing model, system prompt or state-construction code. Compare outputs and tool decisions at key checkpoints. This catches assumptions where the new model interprets summaries or prior assistant text differently.
State architecture should also make replay safe. A replay used for debugging or evaluation should call mock/read-only tools or a sandbox unless the original side effects are explicitly intended. Replaying a stored conversation against production tools without guardrails can resend emails, recreate tickets or repeat purchases.
Keep conversation-state migrations reversible when schemas change.