Databricks inference tables make runtime AI traffic queryable by logging model-service requests and responses into Unity Catalog Delta tables. In the current architecture, inference tables are a Unity Gateway feature used for debugging, monitoring, optimization, and compliance. The older legacy model-serving inference-table experience was retired in 2026 and should not be the basis for new designs.
Within Generative AI on Databricks, inference tables sit between serving and evaluation. They provide the raw operational record of what applications asked, what the model or gateway returned, which route was used, and how the request behaved. MLflow traces and scorers can then add richer application-quality evidence around that runtime record.
Because payloads can contain sensitive user input and model output, inference logging is also a data-governance decision, not only an observability feature.
Unity Gateway inference tables are the current path
Current Databricks documentation describes inference tables under Unity Gateway model services. The feature is billed and logs requests and responses into Unity Catalog tables chosen by the administrator.
The legacy endpoint inference-table path stopped accepting new enablement in February 2026 and ended support in April 2026. Existing guidance directs teams to migrate to AI Gateway/Unity Gateway inference tables.
New platform standards should therefore ignore the legacy model unless maintaining historical endpoints during migration.
Request IDs and invocation IDs help reconstruct routed traffic
Current Unity Gateway inference tables include identifiers such as request_id and invocation_id. A request can produce multiple invocations when guardrails, fallbacks, or multi-step serving behavior result in more than one underlying inference action.
This distinction is valuable when traffic splitting or fallback is enabled because one user request can touch more than one backend.
Operators should preserve client-side correlation IDs as well so application traces can be joined to the gateway record.
Request and response payloads enable deep debugging
Inference tables can retain the raw request and response content that moved through the service. This lets teams reproduce malformed inputs, inspect model output, identify prompt regressions, and compare problematic requests with healthy ones.
That visibility is powerful enough to require strict access. Payload tables may contain personally identifiable information, proprietary documents, credentials accidentally supplied by users, or regulated content.
Access should follow least privilege, and the organization should decide whether all payloads need full retention.
Sampling and logging cost should be designed deliberately
Inference-table logging is billed, which means a high-volume service can create meaningful observability cost if every request is stored indefinitely.
Sampling, retention, filtering, and workload criticality should determine how much data is captured. Security or high-risk workflows may justify broader capture; lower-risk, high-volume services may use sampling plus aggregate metrics.
Cost reduction should not destroy the evidence required for incident response or compliance, so the policy needs named owners.
At-least-once delivery means duplicates are possible
Databricks documents at-least-once log delivery for inference tables, which means duplicate log rows can occur occasionally.
Analytics should use request/invocation identifiers rather than assuming every table row represents one unique user request.
This matters for cost estimates, error-rate calculations, and evaluation sampling where duplicate rows could bias the result.
Inference logs should be joined with route and model metadata
Useful analysis asks which model, destination, service, or model version handled the request. Current inference-table schemas expose serving metadata and gateway context that can be combined with routing configuration.
Databricks Model Serving Routes is directly relevant because traffic splitting and fallbacks change which model actually generated a response.
Quality analysis should compare outcomes by model destination rather than only by application endpoint.
Payload logs are not the same thing as MLflow traces
Inference tables show gateway/model-service interactions. MLflow tracing can capture a much wider application path: retrieval, agent steps, tool calls, prompt assembly, business functions, and model invocations.
Use inference tables when the question is about serving traffic or gateway behavior. Use traces when the question is how the whole GenAI application reached its outcome.
The two evidence sources are complementary and can share correlation identifiers.
Retention should reflect privacy and audit requirements
Production teams should define how long payloads are retained, whether sensitive fields are redacted before serving, who can query the tables, and how deletion requests are handled when user data appears in logs.
Unity Catalog provides the governance surface, but the organization still decides the policy.
Logging every prompt forever “for debugging” can create a second uncontrolled repository of sensitive application data.
Inference tables can feed offline evaluation datasets
Production requests that reveal quality problems can be curated into MLflow evaluation datasets. Operators can identify representative failures, add expected facts or guidelines, and rerun scorers against candidate fixes.
This creates a practical quality loop: production evidence → curated example → evaluation → prompt/model/retrieval change → monitored release.
Raw traffic should be curated rather than copied blindly; evaluation sets benefit from representative, reviewed examples rather than every production request.
Inference logging is successful when it reduces mean time to explanation
A production team should be able to answer: which request failed, what model route handled it, what payload was sent, what response returned, whether a guardrail or fallback was involved, and which application trace corresponds to the request.
Inference tables provide that serving evidence. Their value comes from making runtime behavior queryable enough that debugging and quality work start from facts rather than anecdotes.
Schema evolution in the logging table should be handled like any other production data contract. Gateway features can add fields or metadata over time, and dashboards should select named columns deliberately rather than depend on brittle positional assumptions.
Request and response bodies can be large. Teams should avoid scanning the full payload columns for every dashboard query when aggregate usage or error analysis can rely on metadata fields. Curated views can expose operational signals while keeping raw payload access narrow.
Compliance use cases should distinguish audit evidence from debugging retention. A regulated application may need certain metadata retained for years while raw prompts should be removed much sooner. Separate curated audit tables can preserve required evidence without keeping every payload indefinitely.
Fallback and guardrail behavior can create multiple invocations for one request, which makes request-level and invocation-level metrics different. Count user requests when measuring product traffic; count invocations when measuring backend load, evaluator activity, or provider cost.
Inference tables should be joined with system serving/billing data when the organization wants cost per request or per model destination. The gateway payload table shows what happened; billing usage shows what it cost.
Sampling should be evaluated for bias. If only successful traffic is sampled heavily while errors are always retained, quality dashboards need to account for that policy. Sampling_fraction and request metadata should be preserved in analysis so rates are not interpreted incorrectly.
Inference logging becomes operationally powerful when curated queries are attached to runbooks: failed status codes, slow invocations, policy-triggered requests, fallback paths, route experiments, and repeated user errors. The raw Delta table is the evidence store; the operational value comes from the views and workflows built on top of it.
Payload logging should also be tested with streaming responses and multi-turn requests so the team understands exactly what is persisted. The logged representation may differ from the client’s incremental view, and large streamed outputs can create materially larger storage usage.
Curated inference views can hash or remove direct identifiers while preserving dimensions needed for analysis. This is often a better default for broad analyst access than granting raw-table access to every team investigating model quality.
Retention policies should be automated. Relying on analysts to remember to delete old payload data defeats the privacy and cost controls the organization intended. Delta table lifecycle, access reviews, and archival policy should be part of the inference-logging setup.
When multiple applications share one model service, request tags should identify product, environment, feature, and tenant class so inference-table analysis can separate them. Shared infrastructure without request metadata creates blended quality and cost signals.
Inference-table changes should be coordinated with downstream dashboards. If a routing or guardrail feature begins producing multiple invocation rows per request, request-count queries that previously counted table rows directly may become wrong.
Inference-table schemas should be abstracted behind stable curated views for downstream analytics. This gives the platform team room to adopt new Unity Gateway fields without forcing every dashboard or quality notebook to change simultaneously.
High-risk applications can enrich inference records with request tags such as release, feature, tenant class, and experiment cohort so security and quality analysis can isolate affected traffic quickly.
The operational goal is not maximum logging. It is enough trusted evidence to reconstruct important interactions while preserving privacy, controlling cost, and keeping query access governed.
Operational queries should also identify repeated client-request IDs or suspicious duplicate invocation patterns. Those can reveal retry storms, client bugs, or users resubmitting the same request because the application failed to acknowledge completion.
When raw payload access is restricted, platform teams can expose sanitized views with latency, status, model route, tags, token use, and selected quality flags so product teams still have useful observability without broad access to user content.
Keep sanitized operational views versioned with the gateway schema so dashboards remain stable as logging fields evolve.