LLM Evaluation: Judges, Metrics, and What They Miss

LLM evaluation looks scientific because it produces scores, but every score embeds assumptions about the task, rubric, judge, data, and reference answer. The current Databricks Generative AI Engineer Associate exam guide expects engineers to use evaluation frameworks, choose judges and metrics, and understand what those results mean. The useful mental model is causal: an input and trace are presented to a scorer, the scorer applies a definition of quality, and the output is only as trustworthy as that definition and evidence.

Databricks MLflow supports built-in LLM judges for qualities such as relevance, safety, groundedness, and correctness, as well as custom judges and deterministic code-based scorers. Those options exist because not every quality dimension can or should be delegated to another language model.

The same lesson behind data-quality accountability applies to evaluation data. A carefully designed judge running against mislabeled, unrepresentative, or outdated examples will produce consistent evidence about the wrong task.

A metric begins with a decision

Before choosing a scorer, state what decision the metric supports. Is it a release gate, an experiment comparison, a monitoring signal, or a diagnostic clue?

A numeric score used only to rank two prompt variants can tolerate different calibration from a score that automatically blocks production deployment.

Every metric should have an owner and an action. If nobody knows what to do when groundedness falls from 0.91 to 0.86, the number is descriptive rather than operational.

Metrics should also be separated by task category before they become gates. A support agent may be excellent on common account questions and unsafe on refund authorization. If the release threshold averages those cases together, strong routine performance can hide a critical failure. Define category-level minimums when business consequences differ materially.

Release thresholds should also state what happens near the boundary. A candidate scoring 0.801 against a 0.80 gate is not meaningfully safer than one scoring 0.799 when evaluator variance is material. Use confidence intervals, repeated judgment, or human review around marginal cases instead of converting uncertain measurements into falsely precise pass/fail decisions.

Built-in judges are fast starting points

Predefined judges give teams standard definitions for common quality dimensions and make early evaluation easier.

They are useful for broad relevance, groundedness, safety, or correctness checks where the application matches the judge’s assumptions.

Treat them as baselines, not universal truth. Domain-specific requirements, unusual output formats, regulated wording, or business-specific behavior often need a custom judge or deterministic rule.

Built-in judges should be validated against a small human-reviewed calibration set. If the judge consistently rates concise domain answers lower than verbose generic answers, the issue is evaluator fit rather than application quality. Calibration does not require thousands of examples; it requires representative disagreement cases that expose systematic bias.

Custom judges move ambiguity into the rubric

A custom LLM judge lets teams express nuanced criteria in natural language and can inspect inputs, outputs, expectations, or traces.

That flexibility is powerful and means the judge prompt becomes evaluation code. Vague instructions produce variable judgments; hidden examples or conflicting criteria create unstable scoring.

Version the rubric, model, threshold, and examples. A changed judge can alter release results even when the evaluated application is identical.

Custom judge instructions should specify what evidence the judge may use. A groundedness judge should know whether it may rely on world knowledge or only provided context. A policy judge should know whether one missing disclaimer is a failure or merely a style issue. Clear evidence boundaries make scoring more reproducible.

Judge rubrics should include negative examples where possible. Showing what should fail helps reveal hidden interpretation better than adding more positive prose. If a style rule says ‘concise,’ include an example that is technically correct but too verbose, so evaluators and reviewers share a concrete boundary rather than one subjective adjective.

Code-based scorers are stronger for deterministic rules

Some requirements do not need an LLM opinion. JSON validity, field presence, exact identifiers, latency, token limits, citation count, or known forbidden strings can be checked deterministically.

Use code-based scorers where a precise rule exists. They are cheaper, easier to reproduce, and less sensitive to model variation.

A robust evaluation suite often combines deterministic checks with LLM judges rather than forcing one method to cover every dimension.

Deterministic scorers are especially useful for invariants. If every response must contain a source ID, valid currency code, or machine-readable status field, a code rule can provide exact failure evidence. That keeps LLM judges focused on genuinely semantic qualities instead of spending tokens deciding whether a required field exists.

Ground truth helps and can mislead

Expected answers are useful when the task has stable correct outputs. They become brittle when multiple responses are equally acceptable or the source data changes.

Ground-truth labels should include provenance, version, and review ownership. A ‘golden’ answer can become wrong after business policy changes.

Where exact reference answers are inappropriate, define expected facts, behaviors, citations, or constraints instead of comparing free-form text to one ideal sentence.

Reference sets should have review dates. Product names, policy wording, and source facts change, and an outdated expected answer can penalize a newly correct system. Tie each example to the source or business rule that justifies it so reviewers can update the evaluation dataset when the authoritative requirement moves.

Judges can inherit model bias and instability

An LLM judge can favor verbose answers, familiar writing styles, or outputs similar to its own preferences. Different judge models can disagree, and model updates can change scoring over time.

Calibrate automated judges against human expert review on representative samples. Measure agreement and inspect disagreement categories.

Do not use one judge as both generator and unquestioned authority when high-consequence decisions require independent evidence.

Judge disagreement should be surfaced instead of averaged away. Compare multiple judges or human review on high-impact cases and record where interpretations diverge. A release decision can require consensus for safety-critical examples while accepting statistical movement on subjective style. The point is to match evidence strength to consequence.

Model-based judges should be monitored for cost and latency too. Large evaluation suites can become expensive enough that teams run them less often, weakening release discipline. Use cheaper deterministic rules for simple checks and reserve stronger judges for semantic dimensions where they add real information.

Evaluation datasets must represent the failure space

Include ordinary examples, hard examples, edge cases, refusal cases, multilingual inputs, ambiguous questions, tool failures, and examples gathered from production incidents.

Average scores can hide a critical category. A system can improve common support answers and become worse at billing disputes or safety-sensitive requests.

The broad experimentation discipline of Python data science workflows applies: stratify results, inspect distributions, and preserve enough metadata to understand which population moved.

Evaluation coverage should also include tool and retrieval traces. The final answer may be correct for the wrong reason—for example, a tool returned stale data but the model guessed the current value. Trace-aware judging can detect whether the system used the expected source, called the right tool, or followed the intended reasoning path.

Production-derived evaluation cases need privacy controls. A valuable failure example may contain customer information, credentials, or sensitive business context. Redact or synthesize where possible while preserving the failure mechanism; otherwise the evaluation dataset becomes another sensitive system with broad engineering access.

Offline scores do not replace production evidence

A candidate can win on the evaluation set and fail under real user phrasing, traffic, retrieval freshness, tool latency, or adversarial behavior.

Production monitoring should reuse compatible scorers where possible and compare sampled traces with the offline baseline.

Use logging and monitoring to connect score movement with application versions, prompt changes, model changes, and dependency failures instead of treating quality as an isolated dashboard.

Production sampling should be stratified. Random sampling can underrepresent rare but high-risk flows, while only sampling failures can make quality look worse than typical user experience. Combine routine traffic, critical categories, user-reported failures, and release-specific cohorts so monitoring reflects both prevalence and consequence.

Evaluation is strongest when it exposes uncertainty

Report metric definitions, sample sizes, judge versions, confidence or disagreement where available, and the categories in which the evaluator is weak.

Use human review for disputed, high-impact, or poorly calibrated cases rather than hiding ambiguity behind a decimal score.

LLM evaluation becomes useful engineering when teams know what the judge measured, what it missed, how stable that measurement is, and which production decision the evidence is strong enough to support.

Evaluation reports should preserve uncertainty and failure examples, not just leaderboards. Show sample size, scorer version, category breakdown, human disagreement, and the actual traces behind low scores. That lets reviewers decide whether the measured difference is meaningful enough to justify a production change.

The final release review should distinguish evidence from policy. Metrics and judges describe observed behavior; product and risk owners decide what quality is acceptable. Keeping that distinction visible prevents engineers from treating an evaluator threshold as an objective law when it is actually one organization’s tolerance for error.

Evaluation pipelines should preserve the exact application version and data snapshot used for each score. If a judge result cannot be tied back to the prompt, model, retriever, tool configuration, and evaluation dataset that produced it, later comparisons become historical anecdotes rather than reproducible evidence.

Threshold changes should themselves go through review. Lowering a release gate because a promising model barely misses it can be legitimate when the rubric was too strict and dangerous when it simply rationalizes a preferred release. Record why the threshold changed and rerun stable baseline versions to understand the effect.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!