Mosaic AI Agent Evaluation has moved into managed MLflow 3 on Databricks. Current documentation says the Agent Evaluation SDK methods are now available under mlflow.genai, with MLflow 3 providing tracing, evaluation datasets, built-in and custom scorers, LLM judges, conversation evaluation, synthetic conversation simulation, human feedback workflows, and production monitoring. The approved title preserves the older product name; production code should follow the current MLflow 3 APIs.
Within Generative AI on Databricks, evaluation is the quality control loop around agents, RAG systems, prompts, and tools. It should begin before deployment and continue on production traces after users encounter cases the development dataset never covered.
The existing MLflow for GenAI article provides the broader lifecycle context.
MLflow 3 is now the current evaluation surface
Legacy MLflow 2 Agent Evaluation code used databricks.agents.evals and related packages.
Current Databricks guidance migrates to mlflow.genai.evaluate, mlflow.genai.scorers, mlflow.genai.judges, and managed labeling APIs.
Plan migration rather than maintaining two evaluation stacks because new observability and production-monitoring capabilities are built around MLflow 3.
Evaluation datasets capture inputs and optional expectations
MLflow evaluation datasets store JSON-serializable inputs plus optional expectations such as expected facts, expected response, and guidelines.
Add source, tags, and lineage so each case explains why it belongs in the suite.
Keep a stable regression core plus an evolving set of production failures; otherwise scores become difficult to compare across agent versions.
Built-in scorers cover common GenAI quality dimensions
Current MLflow includes scorers such as correctness, relevance, safety, retrieval groundedness, retrieval relevance, and retrieval sufficiency.
Use the scorer that matches the failure mode rather than one universal “quality” number.
A RAG agent can answer fluently while retrieval is poor, and a tool agent can retrieve perfectly while calling the wrong function.
Custom scorers encode product-specific requirements
Write custom scorers for business rules such as mandatory citations, required escalation, forbidden tool use, expected JSON fields, or response-time/cost thresholds.
Prefer deterministic scorers when correctness can be computed directly.
Use LLM judges where semantic judgment is genuinely needed, and calibrate them against human labels on representative cases.
Tracing is the substrate for agent evaluation
MLflow Tracing records model calls, tool calls, retrieval, intermediate messages, latencies, errors, and other spans across supported frameworks.
Evaluation can run on these traces during development or production.
Without traces, a low score shows that the final answer is wrong but not whether the root cause was retrieval, prompt, tool selection, tool output, or model reasoning.
Conversation evaluation handles multi-turn behavior
Current MLflow evaluation supports multi-turn conversation quality with scorers for aspects such as completeness, user frustration, and dialogue coherence.
This is important for support or workflow agents where one good final answer can hide five frustrating turns.
Evaluate full conversations and state transitions, not only isolated question/answer pairs.
Conversation simulation expands scenario coverage
MLflow can simulate synthetic multi-turn conversations to test user behaviors and edge conditions at larger scale.
Use simulation to explore ambiguous users, corrections, adversarial instructions, tool failures, and escalation scenarios.
Keep human-reviewed golden conversations as anchors because synthetic users can also encode unrealistic behavior.
Review App gathers human feedback into reusable evidence
Managed MLflow includes labeling/review workflows that let experts score or annotate traces and evaluation examples.
Use domain experts for criteria automated judges cannot reliably assess, such as legal interpretation or nuanced support quality.
Convert recurring human feedback into dataset expectations or custom guidelines so the evaluation suite improves rather than repeating the same manual review forever.
Production monitoring should reuse development scorers
Databricks is building production monitoring around scheduled scorers over real traces.
Reusing the same scorers from predeployment testing makes regressions comparable across lifecycle stages.
Sample intelligently to control cost and privacy; not every production conversation needs every LLM judge, especially when deterministic safety/latency metrics already cover some properties.
Evaluation should drive release decisions, not generate dashboards only
Define promotion thresholds, acceptable regressions, must-pass cases, and rollback triggers before evaluating a candidate.
Compare model/prompt/tool versions on the same dataset and inspect low-performing examples, not just aggregate means.
Prompt orchestration and evaluation provides the broader release-control rationale.
Agent Evaluation succeeds when failures become permanent tests
The mature workflow traces every version, evaluates stable datasets with deterministic and LLM scorers, simulates conversations, captures expert feedback, monitors sampled production traces, and converts incidents into regression cases.
The product name has moved into MLflow 3; the engineering goal remains the same: agent quality should be measurable enough to improve deliberately.
Scorer calibration should be treated as a model-evaluation problem itself. Before an LLM judge becomes a release gate, compare its scores with trusted human labels across good, bad, borderline, multilingual, adversarial, and domain-specific examples. Record disagreement rate and failure modes so teams know where automated judgments require manual review.
Retrieval-specific scorers should be evaluated independently from generation. RetrievalGroundedness, RetrievalRelevance, and RetrievalSufficiency can reveal whether the RAG evidence was present and useful before the final answer was generated. This lets engineers fix chunking, indexing, or search parameters instead of endlessly rewriting the answer prompt.
Tool-using agents need trajectory scorers. Measure whether the agent selected an allowed tool, supplied valid parameters, avoided duplicate side effects, respected approval boundaries, and recovered correctly after errors. A final answer can look correct while the agent took a dangerous or expensive path to reach it.
Cost and latency should appear in the evaluation table alongside semantic scores. A candidate that improves correctness by one point but triples model calls and p95 latency might be unacceptable for production. Use custom scorers or trace-derived metrics so quality optimization remains bounded by service economics.
Production monitors should segment scores by model version, prompt alias, tool version, user cohort, and query category. A global average can hide one workflow collapsing after a tool schema change. Attach version metadata to traces so scheduled scorers can identify exactly which release introduced the regression.
Human review should focus on uncertainty, disagreement, and high-impact cases. Sending every trace to experts is expensive and slow. Use automated scorers to surface low-confidence, conflicting, safety-sensitive, or business-critical examples, then convert expert labels into improved datasets and judge calibration.
Evaluation datasets need privacy/retention controls. Traces and examples can contain customer prompts, document excerpts, tool outputs, and sensitive identifiers. Store only what is needed, restrict access by project/team, redact where possible, and apply deletion requirements consistently to derived evaluation datasets.
Release gates should have must-pass cases that no average can hide. Examples include ‘never reveal another tenant,’ ‘always escalate suicide risk,’ ‘never execute a write without approval,’ or ‘must cite the policy source.’ A candidate that achieves a high mean score while failing one critical invariant should be rejected.
Evaluation runs should record model, prompt, retrieval index, tool schema, framework/library versions, and application commit. If a score changed, the team needs to know which component changed. Treat the evaluated agent as a composite release rather than attributing every quality movement to the foundation model.
Golden datasets should include non-answer scenarios: insufficient evidence, missing permissions, tool unavailable, ambiguous request, and user correction. High-quality agents know when to ask, abstain, or escalate. A dataset containing only answerable happy-path questions trains the release process to reward overconfidence.
Judge prompts and scoring rubrics should themselves be versioned. Changing an LLM judge can move evaluation scores without any agent improvement. Store judge/scorer version with results and rerun a calibration subset when updating judge model or rubric.
Statistical uncertainty matters when score differences are small. A candidate that improves average correctness by 0.3 points on 20 examples may not represent a real improvement. Use enough diverse cases, per-category breakdowns, and confidence/significance reasoning before making large production changes from tiny deltas.
Production feedback signals such as retry, thumbs-down, human takeover, edited response, or tool cancellation can enrich evaluation datasets, but they are noisy proxies rather than labels by themselves. Sample and review them before treating them as ground truth.
Safety evaluation should exercise both direct user prompts and indirect content arriving through retrieval or tools. An agent may ignore a malicious user instruction but follow a prompt injection hidden inside a document. Build traces that reproduce both pathways and score whether system/tool policy remains intact.
Evaluation infrastructure should have cost controls. LLM judges over large trace sets can be expensive. Use deterministic checks first, sample production traces, cache stable judgments where appropriate, and reserve expensive multi-judge evaluation for release candidates or high-risk segments.
Evaluation dashboards should preserve raw examples behind aggregate scores. Reviewers need to inspect the specific traces that improved or regressed, especially when two versions have similar averages. A release decision is stronger when engineers can explain the representative wins/losses rather than pointing only to one headline metric.
Sampling strategy in production should be biased toward informative events: failures, retries, escalations, high latency, unusual tool paths, low-confidence judge scores, and representative normal traffic. Pure random sampling can miss rare but severe behaviors, while evaluating only failures gives a distorted view of overall quality.
Keep evaluation ownership explicit: one team should maintain dataset quality and scorer calibration, while application owners remain accountable for fixing regressions. Shared evaluation infrastructure works best when responsibility for interpreting and acting on scores is not lost between platform and product teams.