When should you use human feedback vs automated metrics? | Scorable

Updated: 2026-04-18
By: Ari Heljakka

Short answer

Human feedback and automated metrics play complementary roles in production evaluation rather than serving as substitutes for one another. Automated metrics (deterministic checks and LLM-as-judge) are what give you scale, consistency, and continuous coverage on every release, while human review remains the calibration source for the judges and the only credible signal on high-stakes slices and on dimensions that resist reduction to a deterministic check. The pattern that holds up under real load is to automate the routine 80 percent, have humans review the ambiguous 20 percent, and track agreement between the two as a first-class metric that drives ongoing recalibration.

Key facts

Key takeaways

Definition

Evaluation in an LLM system can use one of three mechanisms.

At the system level the question is which mechanism scores which dimension. Format compliance always goes to a deterministic check. Faithfulness usually goes to a judge calibrated against a small human-labeled set. Cultural appropriateness and ethical nuance go to humans, with judges only as a secondary signal.

When this matters

The trade-off becomes binding whenever the cost of evaluation starts to compete with the cost of the system itself.

How it works

A working system layers the three mechanisms by cost and coverage.

Tier 1: Deterministic checks on 100 percent of traffic

Run every output through code-based scorers first. Schema validation, format checks, length, banned phrases, latency, and cost belong here. These are cheap, fast, and reproducible. Anything that fails a deterministic check fails the run; the judge does not need to see it.

Tier 2: LLM-as-judge on 100 percent of traffic (or a sampled fraction)

Anything that passes the deterministic checks goes to one or more judges. Each judge scores a single dimension (faithfulness, relevance, tone, helpfulness) on a normalized 0 to 1 scale. The judge prompt is versioned, the rubric is versioned, and the agreement against a human-labeled calibration set is tracked over time.

Judge inference cost can be material at high volumes. A common pattern is to run the judge on 100 percent of traffic for low-volume features and on a sampled fraction (5 to 25 percent) for high-volume features, with sampling biased toward anomalies.

Tier 3: Human review on the ambiguous and high-stakes slice

Three populations belong in the human queue:

Tier 4: Agreement as a first-class metric

The most important measurement is not any single score; it is the agreement between the automated layer and the human layer on the calibration set. Track it as a metric in its own right. When it drops, the issue is rarely the humans; it is usually a judge drift, a rubric ambiguity, or a distribution shift. Treat that drop as an actionable signal and recalibrate before trusting the automated scores again.

A common heuristic: do not let an automated dimension drive a release gate until its agreement with humans on the calibration set exceeds a threshold (for instance, Matthews correlation above 0.6 for binary judgments, or rank correlation above 0.7 for graded scores).

Example

A consumer support assistant handles 12,000 conversations per day. The team operates four objectives: format compliance, factual grounding, brand-voice tone, and refusal correctness.

Deterministic checks (100 percent of traffic). JSON schema compliance, response length within bounds, no banned phrases, latency under 1.5 seconds. Cheap to run; never wrong about format.

Judges (sampled 20 percent of traffic). Three judges score factual grounding, tone, and refusal correctness on 0 to 1. Each judge has a versioned rubric and a calibration set of 80 examples; agreement against humans is recomputed weekly.

Human review.

Agreement metric. Tone judge agreement drops from 0.71 to 0.58 after an upstream model upgrade. The team pauses the deploy that depended on that judge, recalibrates against fresh labels, revises the rubric to disambiguate two edge cases, and reruns. Agreement returns to 0.74; the deploy proceeds.

The system scales because automation is the default. It stays trustworthy because humans anchor it.

Limitations

Evidence and sources

FAQ

How do I decide which dimensions go to humans?
Dimensions that resist a written rubric, that depend on context the judge cannot see, or that carry regulatory or reputational stakes belong in the human queue. Anything with a clear, testable definition can usually be scored by a calibrated judge.

How big should the calibration set be?
50 to 150 examples per judge is usually enough to compute a stable agreement metric. Add more when agreement is borderline or when the underlying distribution is heterogeneous.

What is the right agreement threshold to trust automated scores?
For binary classification, Matthews correlation above 0.6 is a defensible bar. For graded scores, rank correlation (Spearman or Kendall) above 0.7 is a useful threshold. Below those bars, treat the judge as a signal, not a gate.

How often should I recalibrate?
Whenever the underlying judge model changes, whenever the rubric changes, and on a fixed cadence (weekly or monthly depending on traffic) to catch silent drift. Treat recalibration as scheduled maintenance, not a one-time task.

Can I skip the judge layer and use only deterministic checks plus humans?
For low-volume, low-stakes systems, yes. As volume grows, the gap between what deterministic checks catch and what humans can review opens up; the judge layer fills it. Without a judge, every semantic failure either goes to the small human queue or escapes evaluation entirely.

What if my judges and humans disagree?
Treat the disagreement as the most valuable data you have. It either reveals a rubric ambiguity (humans interpret it differently), a judge drift (the judge no longer scores the rubric correctly), or a distribution shift (the inputs have moved away from what the rubric anticipated). Each cause has a different fix.