FDE Műhely

Evals

How to evaluate a task with no single right answer

A summary or a deck can be done well a thousand ways. It's still measurable: break it into criteria and build in human feedback from the start.

An invoice total can be settled unambiguously. The quality of a client summary cannot — the answer depends partly on who’s reading.

That doesn’t mean giving up on measurement. It means using a different instrument.

Score the properties, not the output

The trick is to break “is this summary good?” into sub-questions that answer yes or no.

For a client meeting summary, say:

  • Does it contain every deadline that was agreed? (yes / no)
  • Does it contain any claim that wasn’t said? (yes / no)
  • Is every owner named? (yes / no)
  • Is it under 200 words? (yes / no)
  • Does the opening sentence carry the decision rather than the greeting? (yes / no)

Five binary questions produce a score that’s measurable, comparable, and arguable — but at least arguable about something specific.

The second one matters most: does it contain anything absent from the source? That’s the hallucination measurement, and it’s what decides trust.

Where the criteria come from

From existing material. If there are five hundred past summaries the team considers acceptable, you can read out what they have in common. If there are ten that got sent back, you learn even more: the reason for rejection is the criterion.

We write that list with the domain team, and the first version is never the final one. It grows steadily over the first couple of months.

Human feedback isn’t optional

However good the criteria list, it won’t be complete for a judgement-based task. So these cases always keep a human in the loop, and the interface always carries one thing: a way to give feedback.

Not a five-star rating. One question: was this acceptable? And if not, one field: what was missing?

Those two inputs produce two things:

  1. A new criterion. Anything that comes up three times in the free-text field goes into the list.
  2. A new eval case. The rejected output, paired with the corrected version, becomes a measurement case.

So the eval set isn’t a static document. It grows with the system.

What you shouldn’t promise here

There is no 99% on a judgement task. Anyone promising it either doesn’t understand the task or hasn’t measured it.

What you can promise: that the output is good at an acceptable rate, that failures are visible and fixable, and that the rate measurably improves. That’s enough to make a process faster — provided reviewing is faster than writing.

That last condition is decisive. If reviewing and fixing a summary takes as long as writing one, the system isn’t helping, it’s reshuffling the work. That has to be measured during discovery, not assumed.


Related: building a golden dataset and human in the loop.

Questions on this topic

Can we use a model as the judge (LLM-as-a-judge)?

Yes, but only after calibrating it: run it alongside human scoring and check how closely they agree. An uncalibrated machine judge measures its own bias.

What does this look like at your company?

If this problem sounds familiar, let's start with one process. Tell us which department burns the most manual hours — we'll come back with a concrete proposal.

Related articles