What Is Agent Evaluation? Metrics, Trajectories, and LLM-as-Judge | Covasant

Agent evaluation

Agent evaluation is establishing whether an agent is doing its job, judged on the route as well as the result. Four different activities share the name, and most teams do one of them. The instruments for the most-quoted one turn out to be exploitable.

Definition

Agent evaluation means checking whether an AI agent is doing its job well and how it got there. The tricky part is that ‘agent evaluation’ is an umbrella term for four different types of checks, and most teams only run one of them. To make it worse, the most used method also turns out to be the easiest one to game, since agents can pass its tests without genuinely doing good work.

What is agent evaluation?

Agent evaluation establishes whether an agent is working, and it has to judge the route as well as the result. An agent that reaches the right answer through a dangerous or wasteful sequence of tool calls has still failed. That is why the practice looks nothing like model testing.
This is explained in more detail under AgentOps. In short, a model can be scored against a labelled dataset because there's one correct answer. An agent completing a multi-step task usually has several acceptable paths and no single right one. So evaluation shifts from checking correctness to judging quality.

There is no single agent quality score, and anyone offering one has hidden a weighting decision from you. Task completion, trajectory validity, tool-call correctness, cost, latency, and intervention rate can move in opposite directions. An agent that completes more tasks by taking twice as many steps has improved on one axis and regressed on another. Collapsing that into one number requires deciding how much a step is worth relative to an outcome, which is a business judgement rather than a measurement.

What are the four things called agent evaluation?

Capability benchmarking, pre-release testing, production measurement, and incident review. They answer different questions, use different methods, and have different owners. Confusing them is why a team can hold an impressive benchmark number and still have no idea whether its own agent works.

Activity When and what it answers Method What it cannot tell you
Benchmarks Before you build. Can a model do this class of task at all? Public suites, standardized tasks, leaderboards Whether your agent works on your data with your tools
Pre-release Before you ship. Does our agent do the job we built it for? A curated case set of your own, scored on outcome and trajectory How it behaves on traffic nobody anticipated
Production While it runs. Is it still working, and is it drifting? Sampled scoring of live runs, plus outcome telemetry Why a specific run went wrong
Incident After it fails. What happened in this particular run? Trace reconstruction across the full span tree Whether the problem is systemic or a one-off

Why do agent benchmark scores mislead?

Three reasons: benchmarks test a different workload than yours, high scores don't guarantee the agent will succeed on repeated attempts, and the scoring methods themselves can be gamed. That third point is a recent finding, and it changes how much trust a published score deserves.

Benchmarks are still useful for choosing a model. The mistake is treating them as proof that a deployed agent will work, and there are now three separate reasons that assumption breaks down.

What should you measure when evaluating an agent?

Six measures cover most cases: task completion, trajectory validity, tool-call correctness, intervention rate, cost per completed task, and reliability across repeated attempts. Accuracy is not among them, because a multi-step task rarely has a single correct output to be accurate against.

Measure What it captures Why it matters
Completion Share of runs that finished the task end to end, not per-step accuracy The only measure a business stakeholder recognizes
Trajectory Whether the path taken was a reasonable one, scored against a reference or a rubric A right answer by a dangerous route is still a failure
Tool calls Correct tool selected, correct arguments, graceful recovery on failure Where most agent errors actually originate
Intervention Proportion of runs needing a human, and whether it is falling The clearest signal of whether the agent is earning its place
Cost Spend and latency per completed task, not per model call Per-call cost hides agents that succeed expensively
Reliability Success across repeated attempts at the same task, not a single pass Production retries; a one-shot score overstates it

Why does an LLM judge need its own evaluation?

Because a model scoring another model is a measurement instrument with an unknown error rate. Most teams adopt LLM-as-judge and never check its agreement with human judgement, which means they are managing an agent by a number whose reliability nobody has established.

This is the gap that undermines more evaluation programmes than any other, and it is rarely discussed because the alternative is uncomfortable: validating a judge requires human labour, which is the thing the judge was adopted to avoid.

The problem is not that judges are bad. Used carefully they are the only tractable way to score trajectories at volume. The problem is that they have documented, systematic biases and are treated as neutral.

How do you build an agent evaluation practice?

Layer your checks by cost. Run deterministic checks on everything, use an AI judge on a sample of cases, and have humans review a smaller subset to keep that judge calibrated. Build your test cases from real failures, not hypothetical ones. And treat evaluation as an ongoing process, not a one-time check before release.

Which category of tooling supports agent evaluation

Three categories, and most teams end up with two. Evaluation platforms provide case management, judge orchestration, and trajectory scoring, usually with tracing built in or integrated. Observability platforms increasingly add evaluation on top of the traces they already collect, which suits teams whose operations already sit there. Agent management platforms cover the release gate and runtime policy that evaluation results should feed into, which is where an evaluation finding becomes a decision rather than a dashboard.