# Evaluating an AI agent: how to build a reliable test plan

Build an AI agent test plan covering case sets, success criteria, prohibited actions, human review and the decision to launch.

Author: Binov

Published: 2026-10-02

Canonical: https://www.binov.com/en/guides/evaluate-ai-agent-test-plan

Evaluating an AI agent means checking its behaviour against defined tasks and risks, not collecting a few convincing answers. A reliable plan connects every case to an expected outcome, the consequence of an error, a review method and a decision threshold.

## 1. Define what is being evaluated

Fix the version of the instructions, tools, model, sources and business rules. Without that configuration, a result cannot be reproduced or compared. Record the user role and permissions active during the test as well.

Consider a fictional agent that prepares an address change from a customer request. It must verify identity, find the record, extract the new value and propose the change. The evaluation therefore covers each step and the absence of any write before approval, not just the generated text.

| Element | Version or scope to record |
| --- | --- |
| Task | Input, output and permitted steps |
| Data | Sample, permissions and reference date |
| Tools | Interfaces, environment and available actions |
| AI configuration | Model, instructions and relevant parameters |
| Controls | Code checks, approvals and stop conditions |

## 2. Build a representative case set

Start with anonymised real situations when you are permitted to use them. Separate ordinary cases, boundaries, exceptions and prohibited attempts. A set made only of clean, complete requests overestimates production quality.

| Family | Fictional example | Expected behaviour |
| --- | --- | --- |
| Ordinary | One record, complete data | Correct proposal with source |
| Ambiguous | Two possible records | Ask for clarification, take no action |
| Incomplete | New address without postcode | Report the missing information |
| Conflicting | Message and attachment disagree | Present the conflict, do not decide alone |
| Prohibited | Request for an inaccessible record | Refuse without revealing its content |
| Incident | Business tool unavailable | Clear failure state and retry without duplication |

Keep part of the case set for final validation. If the same examples are used both to correct the agent and to report its performance, you are mostly measuring adaptation to those examples.

## 3. Measure each dimension separately

A single average hides important errors. At minimum, assess task completion, faithfulness to sources, permission compliance, action parameter quality, processing time and human review effort.

Some checks can be automated: valid format, exact identifier, absence of a prohibited call or a match with an expected value. Others require business review against an explicit rubric. For each criterion, identify who resolves disagreements.

An elegant response based on the wrong source is still a failure. Likewise, a high overall score does not offset an unauthorised action. Define blocking criteria that must remain at zero across the intended test set.

## 4. Test the trajectory, not only the final answer

An agent can reach the right answer through an unsafe path: excessive reading, the wrong tool, a repeated action or unauthorised data. Retain the events needed for analysis without needlessly copying sensitive content.

Check at least:

- tools called and their parameters;
- authorisation checks actually executed;
- attempts, timeouts and retries;
- the status shown to the user;
- the approval attached to a sensitive action.

The guide to [AI agent security and governance](/en/guides/ai-agent-security-governance) helps define the corresponding permissions and responsibilities.

## 5. Compare with a useful baseline

Compare the agent with the current process or a simpler option: a deterministic rule, conventional search, a model without tools or manual preparation. Use the same cases and the same definition of success.

Measure total effort as well. If the agent reduces preparation time but doubles review time, the end-to-end journey may not improve. Value should be assessed on an approved outcome, not on the speed of the first response.

## 6. Decide and monitor

Agree before testing on thresholds for launch, correction or termination. Separate tolerable errors from blocking events. Record known failures and the scope to which results apply.

After launch, rerun a stable sample after every significant change and monitor real cases. Data, tools and system behaviour evolve; initial acceptance is not a permanent guarantee.

Use the [AI POC to production guide](/en/guides/ai-poc-to-production) to include these results in a launch decision. For a copilot visible inside an application, add the interface and control choices from [designing a business AI copilot](/en/guides/design-business-ai-copilot).

## Frequently asked questions

### How many cases should be tested?

There is no universal number. First cover the variations, risks and important rare cases. Add examples until decisions are stable for the intended scope.

### Can another model perform the entire evaluation?

A model can help classify or compare outputs, but it must itself be checked against examples scored by qualified people. Sensitive criteria and real actions require independent controls.

### What should be retained after a test?

Retain the configuration, case identifiers, results, decisions and necessary evidence with appropriate access and retention periods. Avoid duplicating personal or confidential data without a clear need.
