Skip to content
annsa
Log inTry Annsa
Direct your agents

Definition of done for AI agents

Updated Oct 10, 2026·3 min read

The short answer

A definition of done for AI agents is a short checklist every job must pass before the agent calls it finished. Each line names something you can check, such as a test, a page or a result, and the agent records how it checked each one. "Done" then means checked, not just claimed.

Why an agent's own "done" isn't enough

Agents finish confidently. They may have written the code, but not run it, or run it locally but not where your customers are. Merged isn't deployed, and deployed isn't checked. Without a checklist, you only find the gap when a customer does.

Two levels of done

  1. Your team's standard, true for every job: tests pass, nothing else broke, it's deployed where it should be. See definition of done.
  2. This job's lines, true for this job only. These are acceptance criteria written so an agent can check them.

How to write lines an agent can check

Use "Given a state, when an act, then a result you can see":

  • Given a customer with saved filters, when they export, then the file keeps every filter.
  • Given the pricing page, when it loads on a phone, then nothing scrolls sideways.

One check per line. If a line says "and", split it.

Receipts

For each line, the agent writes down how it checked: a test name, a URL, a screenshot, a thread. A line it couldn't check stays unchecked, and it says so. You can read the receipts without reading the code.

When the agent says done but isn't

  • It checked locally, not where it runs. Ask for the live URL.
  • It checked the happy path only. Add a line for the edge case.
  • The line was vague. Rewrite it so it can fail.

Copy this

Put this at the end of every job plan, and have the agent fill in the receipts.

Done When lines with a receipt for each
## Done When
- Given <state>, when <act>, then <result you can check>.
- Given <state>, when <act>, then <result you can check>.

## Receipts (one per line)
1. <how checked>: <test name / URL / thread>
2. <how checked>: <test name / URL / thread>
Unchecked: <line>, because <reason>

Where Annsa fits

Every spec in Annsa ends with Done When lines. Your agent checks each one and records a receipt, and the spec shows which lines are checked before you tell anyone it shipped. Shipping and receipts

See how it fits together in Direct your agents.

Questions people ask.

Q1

What is a definition of done for AI agents?

A short checklist each job must pass, one checkable line each, with a receipt for how each line was checked.
Q2

How do I write a definition of done for Claude Code and Codex?

End each job plan with Done When lines in "given, when, then" form, and ask the agent to check each one against the live product and say how.
Q3

How do you know when an AI coding agent is really done?

When every Done When line has a receipt you can follow. An unchecked line means it isn't done yet, or you've decided that's fine for now.
Q4

How do I write acceptance criteria an AI agent can check?

One result per line, observable from outside the code: a page, a file, a test, a message. Avoid words like "fast" or "nice" that can't fail.

Get your agents in a row.

Try Annsa