The problem with demos
Every automation vendor can show you a demo that works. The document is clean, the question is well phrased, the answer is right. Then it meets your real invoices, your real customers' spelling, your real edge cases, and nobody has agreed in advance what "working" means. Six months later there is a system that is "mostly fine" and a team that has quietly gone back to doing the job by hand.
An acceptance test fixes this by moving one conversation to the start: what does good look like, in numbers, on your data? Below is how ours is built, step by step, using a document-checking job for a finance team as the worked example. The results quoted at the end are from a production system; the quantities in the steps are typical ranges, not one client's figures.
Step 1 — Your people pick the examples
The test set is not built by us. It is built with the people who do the job today, from real work: for a document job, a few hundred invoices, contracts and onboarding forms taken from recent months. We ask for a deliberate mix: the routine ones, the ugly scans, the supplier who formats everything differently, the ones that caused problems before. A test set that only contains easy cases proves nothing.
Rule of thumb: enough examples that the rarest important case appears several times, not once. The right size depends on how varied the work is; it is agreed with the team, not fixed in advance.
Step 2 — They write the right answers
For each example, the team records what the correct output is. For an invoice: supplier, invoice number, date, purchase-order match, net, VAT, total, bank account. For a ticket: which of the categories it belongs to, and whether it should have been answered automatically or escalated. This is the slow part, and it is also the moment the client discovers that two colleagues disagree about what the right answer is. That disagreement is worth finding before the software encodes one side of it.
The labelled set is split in two: one part to build against, and one part held back for the final test, so the system is judged on examples it has not been tuned on.
Step 3 — We agree the numbers, in writing
Four things get a number, and the numbers go in the contract:
- How often it is right. Measured per field on the held-back set, so that "total" and "bank account" cannot hide behind easy fields like "date". The threshold is agreed with the client before the build.
- How much it catches. Of all the details that should have been found, how many were.
- When it asks a person instead. Every field carries a confidence score; anything below the agreed line is routed to a person, never posted. The test measures that the routing actually happens, and that the routed cases are the genuinely hard ones.
- How fast. Processing time per document, end to end, against the manual baseline.
What this produced for one finance team in production: 99 in 100 details correct, recall above 95%, and a ten-hour day of checking became about an hour. Bank details are additionally validated against official registers, because one wrong digit there is the expensive mistake.
Step 4 — Acceptance day
The system runs on the held-back examples. The results are put next to the answers the team wrote. Anyone in the room can see whether the numbers were met. If they were, it goes live and the team is trained. If they were not, it is not finished, and the contract says so. There is no "mostly".
What it costs, and why it is the first thing we do
Building the test set is the first week of every engagement, and it is priced separately — €700, credited against the build if you go ahead. The point of pricing it separately is that you can stop after it. You keep the labelled examples and the written numbers, and you can hand them to anyone, including your own team.
Questions people ask about the test
Our data is confidential. The set is built inside your environment or on EU hosting under a signed data processing agreement; we see the part you choose to show us.
Our team has no time to label hundreds of documents. It is spread over the first week and done alongside us, and it is the best hour-for-hour investment in the project, because it is where your team defines what "right" means.
What if the job changes after go-live? The test set becomes the regression test. When your data changes — a new supplier format, a new product line — the same set, extended with new examples, tells you whether the system still passes. That is what the monthly "keep it running" work is.
If you want to see what an acceptance report contains, or work out what the numbers would be for your own job, the free 30-minute call is where that happens. If you are still deciding whether a job should be automated at all, start with the three-question test.
I-Team Collage