Skip to content

AI agents

Evaluating an AI agent: how to build a reliable test plan

Evaluating an AI agent means checking its behaviour against defined tasks and risks, not collecting a few convincing answers. A reliable plan connects every case to an expected outcome, the consequence of an error, a review method and a decision threshold.

Binov · 5 min ·

In this guide

1. Define what is being evaluated

Fix the version of the instructions, tools, model, sources and business rules. Without that configuration, a result cannot be reproduced or compared. Record the user role and permissions active during the test as well.

Consider a fictional agent that prepares an address change from a customer request. It must verify identity, find the record, extract the new value and propose the change. The evaluation therefore covers each step and the absence of any write before approval, not just the generated text.

Element Version or scope to record
Task Input, output and permitted steps
Data Sample, permissions and reference date
Tools Interfaces, environment and available actions
AI configuration Model, instructions and relevant parameters
Controls Code checks, approvals and stop conditions

2. Build a representative case set

Start with anonymised real situations when you are permitted to use them. Separate ordinary cases, boundaries, exceptions and prohibited attempts. A set made only of clean, complete requests overestimates production quality.

Family Fictional example Expected behaviour
Ordinary One record, complete data Correct proposal with source
Ambiguous Two possible records Ask for clarification, take no action
Incomplete New address without postcode Report the missing information
Conflicting Message and attachment disagree Present the conflict, do not decide alone
Prohibited Request for an inaccessible record Refuse without revealing its content
Incident Business tool unavailable Clear failure state and retry without duplication

Keep part of the case set for final validation. If the same examples are used both to correct the agent and to report its performance, you are mostly measuring adaptation to those examples.

3. Measure each dimension separately

A single average hides important errors. At minimum, assess task completion, faithfulness to sources, permission compliance, action parameter quality, processing time and human review effort.

Some checks can be automated: valid format, exact identifier, absence of a prohibited call or a match with an expected value. Others require business review against an explicit rubric. For each criterion, identify who resolves disagreements.

An elegant response based on the wrong source is still a failure. Likewise, a high overall score does not offset an unauthorised action. Define blocking criteria that must remain at zero across the intended test set.

4. Test the trajectory, not only the final answer

An agent can reach the right answer through an unsafe path: excessive reading, the wrong tool, a repeated action or unauthorised data. Retain the events needed for analysis without needlessly copying sensitive content.

Check at least:

  • tools called and their parameters;
  • authorisation checks actually executed;
  • attempts, timeouts and retries;
  • the status shown to the user;
  • the approval attached to a sensitive action.

The guide to AI agent security and governance helps define the corresponding permissions and responsibilities.

5. Compare with a useful baseline

Compare the agent with the current process or a simpler option: a deterministic rule, conventional search, a model without tools or manual preparation. Use the same cases and the same definition of success.

Measure total effort as well. If the agent reduces preparation time but doubles review time, the end-to-end journey may not improve. Value should be assessed on an approved outcome, not on the speed of the first response.

6. Decide and monitor

Agree before testing on thresholds for launch, correction or termination. Separate tolerable errors from blocking events. Record known failures and the scope to which results apply.

After launch, rerun a stable sample after every significant change and monitor real cases. Data, tools and system behaviour evolve; initial acceptance is not a permanent guarantee.

Use the AI POC to production guide to include these results in a launch decision. For a copilot visible inside an application, add the interface and control choices from designing a business AI copilot.

Frequently asked questions

How many cases should be tested?

There is no universal number. First cover the variations, risks and important rare cases. Add examples until decisions are stable for the intended scope.

Can another model perform the entire evaluation?

A model can help classify or compare outputs, but it must itself be checked against examples scored by qualified people. Sensitive criteria and real actions require independent controls.

What should be retained after a test?

Retain the configuration, case identifiers, results, decisions and necessary evidence with appropriate access and retention periods. Avoid duplicating personal or confidential data without a clear need.

AI agents to move your operations forward.

Connect tools and workflows with controlled actions and appropriate reviews.