How we know your agent is right
An eval suite is the written definition of correct for one standing agent. This page says what is in it, how it runs on every change we make, and what we do when correct cannot be written down for a piece of work.
What an eval suite is
One agent, one suite. It is written before the agent is built, and you can read it.
A standing agent does one kind of work, unattended, in your own systems. Before we build it, we write down what a correct piece of that work looks like. That document is the eval suite. It is a list of sessions the agent must handle: for each one, the message that starts it and what the answer that ends it must contain.
The definition covers more than the words in the answer. Correct, for a standing agent, is the answer and the way it got there, so it covers four things.
- The answer. What a correct answer to this message contains, and what it must not contain.
- The authority. Which systems the agent reads, which it writes to, and which it never touches.
- The gate. Where the agent stops and waits for a named person, which is before every outward action: a sent email, a published page, a posted message, a payment.
- The cap. What the agent does when a session reaches its spend cap: it stops and asks, or stops and reports when there is nobody to ask.
The suite is yours to read. Your accounts, your keys and your code are yours from the first commit, and the suite is part of that code.
How it runs on every change
A change that fails the suite does not ship.
An agent that runs unattended keeps changing after it is live. The model under it moves, the framework it is built on moves, the systems it touches change shape, and you ask for something new. Every one of those is a change to your agent, and we make it inside your agent’s own repository and deployment.
The suite runs on every change. A change the suite passes is approved and deployed. A change the suite fails stops there: it is not deployed, and the failing entry names the piece of work the agent got wrong. That is why the definition is written down before the build. A rule that exists only in someone’s head cannot stop anything.
When what you want changes what correct means, the suite changes first and the agent follows it. The monthly report says what we changed, and the suite is what you check that against.
When correct cannot be written down
Some work has no written answer. We say so before we build.
Some work has a correct answer that nobody can write down in advance. Whether a reply to an angry customer has the right tone. Whether a refund is fair. Which of two ways to do it when both fit the brief. A person can judge each of these when it comes up, and no document can settle it before it does.
An agent judged on work like that is judged on nothing. So when we cannot write down what correct means for a piece of work, we say so before we build it, and that work does not go to the agent’s own judgement.
That work goes behind an approval gate and to a named person. The agent does the part that can be defined and stops at the gate. A named person at your company decides. The approval policy we write names that person.
An outward action is behind a gate whatever the suite says. A sent email, a published page, a posted message or a payment reaches someone outside your company, and only a named person at your company approves one. The gate for unwritten work is the same gate, used for one more reason.
What we do not do is build that work anyway and call the agent autonomous. A page that promises an agent that decides everything on its own is promising a definition of correct it has not shown you.
The other four parts of how we keep it right