WebDior · 14 August 2026 · 6 min read

An agent without an eval is an anecdote

Most agent demos are true. They are also useless as evidence. Here is what we insist on before an agent is allowed anywhere near production.

Every agent demo we are shown works. That is not a compliment. A demo is one run, on one input, chosen by the person who built it. It tells you the system can succeed. It tells you nothing about how often, on what, or what it does when it fails.

In ordinary software this distinction barely matters, because ordinary software is deterministic: if it worked on the demo input it will work on the same input tomorrow. Agents are not like that. They fail probabilistically, they fail mid-plan, and they fail silently — producing something that looks like an answer. A single successful run is an anecdote.

What an eval actually is

An evaluation set is a collection of real inputs, each paired with what a correct outcome looks like, that the agent is run against before it ships and every time it changes. The score is the contract. If it drops, the change does not merge.

The important word is real. Public benchmarks measure whether a model can do tasks in general; they say nothing about your documents, your edge cases, or the specific way your process fails. The eval has to be built from the client's own history — the cases that went wrong, the cases that were ambiguous, the cases where a human had to step in.

How we build one when there are no labels

Clients rarely have a clean labelled dataset. What they have is history: tickets that were escalated, decisions that were reversed, outputs a reviewer corrected. That history is the eval set in disguise.

We start with the reversals. Every case where a human overrode the previous process is a case with a known correct answer and a known wrong one. A few dozen of those, plus a sample of routine cases, is enough to tell whether a first version is acceptable — and, more importantly, enough to catch a regression when the prompt or the model changes.

What we refuse to do

We decline engagements where the client will not agree on how correctness is measured. It sounds severe. It is the only honest position: a system nobody can prove is working cannot be defended after it ships, and it is the client who will be asked to defend it.

Tagged: Agentic AI · Evaluation

More writing

Labs · Ella · Agentic AI

What a permission model actually needs

'The agent asks before it acts' is easy to say. Building Ella taught us the four properties a permission model has to have before that sentence is true.

Start here

Let's build
the impossible.

Tell us the system you can't get built. We come back with a short, paid discovery — a clear plan and a fixed first milestone — usually within two working days.

01

Discovery

A working session to map the problem and define what "good" is measured against.

02

The plan

Architecture, milestones and a fixed first deliverable — yours to keep, either way.

03

We build

Embedded with your team or as a dedicated pod, shipping with traces, evals and docs.

hello@webdior.comDelhi, IN — working worldwide Booking Q3 2026