LLM Eval Tooling: Four Categories, and the Two That Matter
A held-out task set and twenty assertions beat any platform bought without one. Plus the judge biases that invalidate scores.
Jump to 5 sections
Quick answer: LLM evaluation tooling splits into four types: assertion frameworks that check outputs against rules, LLM-as-judge harnesses that score with another model, human annotation platforms, and production monitoring that evaluates live traffic. Most teams need a small assertion suite and a held-out task set long before they need any commercial platform.
Evaluation is the part of building with models that everyone agrees is essential and most teams postpone indefinitely. The tooling has improved considerably; the discipline it requires has not gotten easier.
This roundup covers the eval tooling landscape in late 2026 — the four categories, what each measures well, where LLM-as-judge breaks down, and what a minimum viable evaluation setup looks like for a team shipping today.
The four categories of eval tooling
They answer different questions and are frequently confused for each other, which is how teams end up with an expensive platform measuring something they do not care about.
| Category | Measures | Cost | Best for |
|---|---|---|---|
| Assertion frameworks | Deterministic properties | Near zero | Format, schema, safety rules |
| LLM-as-judge | Subjective quality at scale | Token costs | Ranking, regression detection |
| Human annotation | Ground truth | High, slow | Validating judges, hard cases |
| Production monitoring | Live behavior | Infrastructure | Drift, incidents, real distribution |
Start with assertions. Checking that output parses as valid JSON, contains required fields, stays under a length limit, or never emits a forbidden string costs nothing and catches a surprising share of real failures. Teams reach for sophisticated scoring before they have written twenty deterministic checks.
In our own reviews at Model Drop, we have found a simple twenty-assertion suite catches somewhere around a third to half of the regressions a team eventually notices in production — schema breaks, truncated output, a forbidden phrase creeping back in. None of that requires a judge model or a dollar of token spend, which is exactly why it belongs first, not last.
Agentic systems need a fifth thing that none of these four categories covers well: trajectory evaluation. When a model takes eight steps to reach an answer, scoring only the final output tells you nothing about whether it got there sensibly or stumbled into it. Step-level evaluation — did it call the right tool, with the right arguments, in a reasonable order — is what catches an agent that succeeds by accident and will fail on the next variation. The same measurement gap shows up when evaluating coding agents against a repository, where a passing test says little about whether the change was sound.
Human annotation is expensive and irreplaceable. It is the only category that produces ground truth, and everything else is ultimately calibrated against it.
Where LLM-as-judge breaks down
It is the default technique now because it is cheap and scales, and it carries documented biases that invalidate results when unaccounted for.
Four biases are well established in the evaluation literature indexed on arXiv's computation and language section. Position bias, where a judge favors the first or second response depending on presentation order. Length bias, where longer answers score higher independent of quality. Self-preference, where a model rates its own family's outputs more favorably. And style bias, where confident formatting outranks correctness.
Those are all mitigable. Randomize position and evaluate both orderings. Control for length explicitly. Use a judge from a different family than the model under test. Score against a rubric with specific criteria rather than asking for a general preference.
The non-negotiable step is validation. Label 100 examples by hand, run the judge on the same 100, and measure agreement. If the judge agrees with human labels 70% of the time, every score it produces carries that error rate — and a 3% difference between two model versions is noise. Model Drop's position is that an unvalidated judge produces numbers, not evidence.
A judge that agrees with human raters 85-90% of the time is a genuinely useful instrument for ranking and regression detection. One agreeing 60-70% of the time is still usable for catching a large regression, a change from mostly-good to mostly-broken, but should never be trusted to detect a small quality improvement between two similar prompts — the noise floor is simply too high at that agreement rate.
What does a minimum viable eval setup actually look like?
A minimum viable setup is a 50-100 example task set drawn from real usage, ten to twenty deterministic assertions on format and content, and a judge model used only after its agreement with human labels has been measured. It is smaller than most tooling vendors suggest, and a small team can build the first version in a day.
- Collect 50 to 100 real tasks. From actual usage, covering the distribution you serve, including the hard and unusual cases. This is the single most valuable artifact and no tool provides it.
- Write deterministic assertions. Schema validity, required fields, forbidden content, latency and length bounds. Fast, free, and they catch real regressions.
- Add a validated judge for subjective quality. Only after checking its agreement with human labels on a sample.
- Run it on every model or prompt change. The point is regression detection, not an absolute score.
Version the task set the way you version code. When someone adds examples mid-quarter, scores shift for reasons unrelated to the system, and a comparison across that boundary is meaningless. Keep the set frozen between deliberate revisions, tag each revision, and record which version produced each result.
Keep a handful of deliberately hard cases in there too. A task set drawn only from typical traffic saturates quickly — everything passes, and the suite stops discriminating between a good change and a neutral one. The examples that occasionally fail are the ones doing the work.
That last point is the one that changes how you use the results. You are not trying to establish that your system is 87% good. You are trying to notice when a change makes it worse, which requires only a consistent measurement, not a correct one.
Public benchmarks do not substitute for this. A model's score on a standardized suite says nothing about your prompt, your data, or your task — the argument our guide to reading launch announcements makes about vendor charts applies with equal force to third-party leaderboards. Standardized suites like those MLCommons publishes are valuable for comparing systems under fixed methodology and still not a substitute for your own task set.
Evaluating production traffic
Offline evaluation measures a frozen distribution. Production is where the distribution moves, and the gap between the two is where most incidents live.
Three things are worth measuring on live traffic. Output validity rates, which catch format regressions immediately and cost nothing. Sampled quality scores on a small percentage of requests, which detect gradual drift. And explicit user signals — retries, edits, abandonment — which are the least noisy quality measure available and the most underused.
A retry rate climbing from a baseline of 4% to 9% over two weeks is a clearer signal of real quality decay than any offline eval score, because it reflects actual users rejecting actual answers rather than a proxy measure agreeing with itself. We have seen teams catch a genuine regression this way days before their offline suite, built from a stale task set, showed any change at all.
Cost per request belongs on the same dashboard as quality. A prompt change that improves output slightly while tripling token consumption is a regression in most businesses, and teams that monitor quality without monitoring spend discover this at the end of the month — the tier arithmetic in our breakdown of LLM API pricing across the 2026 tiers makes the magnitude clear. Track tokens per request alongside every quality metric.
Input drift deserves separate tracking. A prompt tuned on one query distribution degrades when users start asking different things, and that failure is invisible to an offline suite built from last quarter's examples. Refresh the task set periodically from recent traffic — quarterly at minimum, monthly for a product still finding its user base, since the query distribution in the first few months after launch typically moves faster than the eval discipline keeping up with it.
Governance frameworks such as the NIST AI Risk Management Framework treat ongoing monitoring as a requirement rather than an enhancement, which is a reasonable posture for anything making consequential decisions.
The bottom line on LLM eval tooling
Four categories, and the cheapest two do most of the work. A held-out task set drawn from real usage plus deterministic assertions outperforms any platform purchased without one. LLM-as-judge is useful once validated against human labels and misleading before that.
Your next step: collect 50 real tasks from your own logs this week and write ten deterministic assertions against them. That takes an afternoon and will catch more regressions than a tooling purchase.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.