---
title: "LLM Evaluation for Production Systems | Kolena"
url: "/blog/llm-evaluation-production-systems/"
description: "Public LLM benchmarks don't tell you whether a model improved at your task. How to design evaluation that detects regression: ground truth construction, three-way scoring, scoping model-based judges, and re-testing pinned versions for drift."
categories: ["AI Quality"]
updated: 2026-10-01T21:03:04.534353+00:00
---

# LLM Evaluation for Production Systems: Designing a Benchmark That Tells You Something

Public benchmarks measure general capability, not whether a model got better at your task. A practical guide to building evaluation that detects regression — ground truth, three-way scoring, when to use a model as judge, and why pinned versions still drift.

## Why model benchmarks don't answer your question

Every frontier model release arrives with a scorecard: reasoning benchmarks, coding benchmarks, a chart showing the new version ahead of the old one. None of it tells you whether the model got better at the thing you actually use it for.

Public benchmarks measure general capability on standardized tasks. If your task is reading a scanned rent roll, abstracting a forty-page lease, or reconciling a loan tape against the underlying files, a general benchmark is a proxy at best. Models that improve on reasoning benchmarks sometimes regress on document extraction, and nothing in the release notes will tell you which happened.

That gap is why teams running LLMs in production end up building their own evaluation. Not because public benchmarks are wrong, but because they answer a different question.

## Two questions that look identical

Before designing an evaluation, it is worth being explicit about which question it answers, because two very similar-sounding questions pull a suite in opposite directions.

**"How good are we?"** wants a representative sample of real work, cases weighted by how often they occur, and a score that approximates the error rate a user would actually experience.

**"Are we getting better?"** wants close to the reverse: the hardest cases available, weighted by how much they discriminate between systems, and an absolute number that nobody should read too literally.

A suite built for the second question will report a lower score than the first, by design — because trivial cases get retired once everything passes them, and what remains is the difficult residue. Reporting that number as a quality claim would be misleading. Using it to detect regression is exactly right.

Most teams want the second. Choosing it deliberately avoids a lot of confusion later.

## Preference ranking vs. ground truth

The best-known public evaluations measure preference. Two anonymized responses are shown side by side, a human or a model picks the better one, and votes aggregate into an Elo-style ranking.

Preference is the right instrument when there is no correct answer. For open-ended generation — summarization style, tone, helpfulness — there is no key to grade against, and aggregated human judgment is the best available signal.

Document extraction is not that. A tenant's name is transcribed exactly as it appears or it is not. An escalation date is the one printed on the page. Nobody's preference is relevant, and a preference vote cannot tell you that a model read a value out of the wrong row — it only tells you which output someone liked better.

For tasks with a correct answer, score fields against ground truth. The cost is that ground truth has to be built, which is the expensive part of any serious evaluation.

## Building ground truth that means something

If an LLM generates the ground truth and it is accepted without review, the evaluation measures agreement between two models rather than correctness. That is a real failure mode and an easy one to fall into, because generating keys is fast and reviewing them is not.

For a forty-page lease with hundreds of values, signatures, and checkboxes, a usable key is built by hand and verified by a second reader. There is no way around that cost, and attempts to automate it tend to encode the same errors the system under test would make.

Two practical consequences. Suites grow slowly, so they need to be designed for longevity rather than rebuilt each quarter. And the cases that matter most — the genuinely ambiguous ones — always require human adjudication, because they are ambiguous precisely where a model would guess.

## Scoring: three outcomes, not two

Binary pass/fail loses the information that matters most. A table with 47 of 50 rows correct, a date in the wrong format, a value that is right but typed as a string instead of a boolean — these are not failures in the same sense as a hallucinated number, and collapsing them into "incorrect" hides where behaviour is actually changing.

A three-way outcome — correct, partial, incorrect — surfaces drift earlier. Changes in model behaviour tend to appear in the partial column before they appear in the failure column, which makes the partial count a leading indicator rather than a footnote.

Reporting the counts alongside the aggregate score also matters. "500 / 66 / 69" says considerably more about what is happening than "86.35%" does.

## When to use a model as judge

Some comparisons need a model. "Acme Holdings LLC" and "ACME Holdings, L.L.C." are the same company, and a string comparison will insist otherwise. Semantic equivalence, paraphrase, and formatting variation are genuine judgment calls.

Most comparisons do not. Comparing two numbers to four decimal places does not require a language model, and routing it through one introduces a failure mode where none existed.

The risk with model-based judging is that it fails quietly. A grader that degrades looks exactly like a system that degraded — the score drops, and the obvious conclusion is the wrong one. Scoping the judge to only the comparisons that genuinely need judgment limits how much damage a bad grader can do, and makes the failure easier to isolate when it happens.

## Version strings are not stability guarantees

Providers tag stable models with version identifiers, and it is reasonable to assume a pinned version behaves consistently. In practice, that assumption does not always hold — the same version string, same configuration, and temperature at zero where available can still produce measurably different results from one week to the next.

This has a direct consequence for evaluation design: testing a model once, at adoption, is not sufficient. Re-running the same suite against the same pinned version on a schedule is what turns silent drift into something visible.

It also reframes what the evaluation is for. It is not only a selection tool used when choosing a model. It is a monitoring tool for the model already in production.

## Capability suites and workflow suites

Two groupings answer different questions, and running only one leaves a gap.

**Capability suites** isolate behaviours independent of business context: numeric operations, table handling at scale, long-context retrieval, tool use, format parsing. They work like unit tests — when one regresses, it points at a specific mechanism.

**Workflow suites** run a real end-to-end process from a real domain. They work like integration tests — they catch the case where every individual capability improved but the composed workflow got worse.

That case is not hypothetical. A model update can improve table extraction and arithmetic while degrading a workflow that depends on both, because the interaction changed. Only the workflow suite sees it.

## What this costs

None of this is glamorous work. Ground truth is built by hand. Documents are real, which means they arrive rotated, skewed, scanned at low resolution, and occasionally as a photograph of a screen. Suites have to be large enough that genuine regressions clear the noise from ambiguous cases.

The output is typically one number, on a schedule, that tells you whether to investigate further. That is a modest return for the effort — and it is still the only reliable way to know whether the model behind a production system got better, got worse, or quietly changed underneath you.

Kolena runs this process continuously across its document pipeline, which is how model selection decisions get made for each stage rather than adopting one model everywhere. The [LLM Output-Quality Evaluation agent](/agent-library/llm-output-quality-eval/) applies the same approach to customer workflows: scoring generated output against a defined rubric, with each judgment cited to the evidence behind it.
