Eterna Creative

How to test an AI agent that takes real actions

Your agent says it's done. Check the record, not the reply.

by Published 8 min readAutomation

A reply card that says done and a refund is on its way, above a payment system record with a check mark on its corner, in translucent white boxes on a violet to pink gradient

To test an AI agent that takes actions, check the system where the action should have landed, not what the agent said. If it reports a refund, look for the refund in the payment system. Grade the record, build your test set from real failures, and run the same checks again after every change.

This guide is for founders and ops leads who are about to rely on an agent that does more than answer questions: it issues refunds, books appointments, updates records, or files documents. For how Eterna builds AI so it holds up in production, see how Eterna builds with AI. Every check below works whether or not you ever work with Eterna.

The agent saysWhere the truth livesWhat to check
"Your refund is on its way"The payment systemOne refund exists, for the right order and the right amount
"You're booked for Tuesday"The calendar or booking systemOne appointment exists, at that time, for that person
"I've updated your address"The customer recordThe field holds the new value and nothing else changed
"I've sent the documents"The outbound mail log or the portalThe message went to the right recipient with the right attachments
"I've passed this to a person"The queue that person works fromAn item is in the queue, with the context a person needs

Why isn't the agent's reply enough?

An agent can write a confident sentence without doing the work behind it. Hugo Bowne-Anderson gives this example in AI Agent Evals: Test What Matters for Your Agent (2026): a support agent says it booked an appointment, but nothing changed in the system. Sierra, which sells AI voice agents, makes the same point in its guide to voice AI (2026). The agent should confirm the system result, and an agent that says a task is complete without a verified event is a failure signal.

A model's stated confidence needs testing too. ConfBench, a 2026 arXiv preprint by Roy and colleagues, tested 7 vision-language models on 1,346 degraded document variants and found calibration, how closely a model's stated confidence matches how often it is right, ranging from near-perfect to severely overconfident. That study is about reading documents, not taking actions. It is one reason to treat a confident answer as a claim to check, not as proof.

What should you check instead?

Check 3 layers, and grade the last one.

LayerWhat it isWhat it tells you
The callWhether the agent called the right tool with the right inputsWhether it did the work, or only described it
The steps in betweenThe state after each step of a multi-step runWhere a run broke, so you can fix the cause
The recordThe state of the system that holds the truthWhether the work happened, which is what your customer experiences

The call and the steps help you find out why a run failed. Only the record tells you whether it failed. A passing test then reads as a sentence you can say out loud: for this order, one refund of the right amount exists in the payment system.

Test the cases where the right action is no action as well. A refund request outside your policy should leave the payment system untouched, and that counts as a pass.

How do you build the test set?

Start from real failures, not from a list you imagine. Hamel Husain and Shreya Shankar's guide AI Evals: Everything You Need to Know (2026) starts the same way: read real runs first, then decide what to check. 4 steps follow.

  1. Collect real cases. Use real requests from your own work, including the messy ones. A vendor's benchmark scores the vendor's cases, not yours.
  2. Write each check as pass or fail. Husain and Shankar recommend binary pass or fail checks over rating scales.
  3. Let the person who knows the work decide what counts as right. A developer can write the check, but an operations lead knows which mistakes cost money and which are only untidy.
  4. Grow the set from failures. Every time the agent gets something wrong, in testing or in use, add that case so it cannot return unnoticed.

If you have not yet chosen which process to hand to an agent, start with which process to automate first, and take a baseline before you build.

When should the agent stop and hand over?

An agent needs a defined way to say it is not sure, and your test should include cases where the right outcome is a hand-off to a person. Check that the hand-off landed: an item sits in the queue the person works from, with the context they need to act without asking the customer again.

Where a mistake is expensive, add a second opinion that does not depend on the first model's confidence. When several models read the same input and disagree, the disagreement is a flag. Zhang and colleagues (Consensus Entropy, an arXiv preprint from 2025) found that measuring agreement across several vision-language models improved quality-verification F1 by 42.1% over a single model acting as judge. That result is for checking text read from documents, and it is a preprint, so read it as a direction, not a promise for your agent.

What does this look like in our own work?

Zonik AI is a logistics product we co-founded. It reads the broker rate confirmations its first users upload with 2 models and compares the answers field by field. Dates, times, rate, weight, and miles must match exactly, the broker name is compared after normalizing it, and notes and labels are not compared. A field where the 2 models disagree is flagged for the dispatcher. The dispatcher reviews and confirms before the load is saved. In its load tracking, unresolved loads go to one human queue. These describe how the checks are built, not how well they perform.

In September 2026, before we chose the models, we ran 14 AI models on real US broker rate confirmations in TestAImodels, our own tool for comparing AI models. The sample was 3 scanned and 2 digital documents from 3 broker templates, so read it as a test, not a benchmark. It found 7 reproducible extraction errors, which we fixed in the production prompt. The strongest scanned-document model failed both digital documents. The full test is in We tested 14 AI models on freight rate confirmations, and what 98% accuracy in AI document extraction really means explains how to read such numbers.

Every build we ship with an AI feature includes AI checks: we check AI output against real examples from your work, test that guardrails hold, and, for AI agents, check each step they take and alert when quality drifts. When a model is down or unsure, the system takes a simpler path or hands the task to a person. Our design principle for law firms is that AI prepares and the lawyer decides: nothing AI-generated reaches a client or a court without a lawyer approving it. See Legal Operations.

What does it cost and how long does it take?

The checks are part of the build, inside the fixed project price, not a separate project. A one-off automation starts from €500, in 3 ranges by complexity, and automations with complex AI inside are quoted above €3,000. We size the checks to risk: anything that touches money, data, or work that runs with nobody watching gets the strictest checks. Every build starts with the free Blueprint, which sets out what we build, what it costs, when it ships, and a clear definition of done. The cost guide shows what moves a build from one range to the next.

FAQ

What is an eval for an AI agent? An eval is a repeatable test that runs the agent on a set of cases and grades each result as pass or fail. For an agent that takes actions, grade the record in the system it acted on, not the reply it wrote.

How many test cases do I need? There is no magic number. Start with the cases that have already gone wrong, add each new failure as it appears, and keep the cases where the right answer is to stop and hand over.

Can another AI model grade my agent? It can, but check it first. Husain and Shankar (2026) recommend validating a model judge against human labels before you trust it. Where a mistake costs money, a check against the record is stronger than a judge reading the agent's own words.

Do I need my own test if the model scores well on public benchmarks? Yes. In the September 2026 test we ran for Zonik AI, a logistics product we co-founded, the strongest scanned-document model out of 14 failed both digital rate confirmations, in a sample of 3 scanned and 2 digital documents. A model that leads on one kind of input can fail on another, and your inputs are not the benchmark's.

What if the action cannot be checked automatically? Put a person between the agent and the result. In Zonik AI, a logistics product we co-founded, the dispatcher reviews and confirms before a load is saved. For law firms, our design principle is that nothing AI-generated reaches a client or a court without a lawyer approving it.

How often should I re-run the tests? After any change to the prompt, the model, or the tools the agent can call. For anything that runs unattended, add an alert for the day expected work does not happen, because an agent that fails silently produces no error to catch.

The verdict

Do not accept the agent's word that it is done. Decide where the truth lives for each action, check the record there, build the test from real failures, include the cases where the right move is to stop, and run it again after every change.

If you want an agent that does real work, and a test that shows it did, book a free audit call. To see how Eterna builds AI so it holds up in production, start with how we build with AI.

Bring the bottleneck in your system.

Book a free 30-minute strategy call. We'll show you how we'd build the system and map the lightest next step, whether we work together or not.

Book a free audit call