Answer in brief
Whoever supplies the examples decides the outcome of an AI evaluation. If the vendor brings the demo data, the procurement has already been settled before the first meeting.
The central idea
Whoever supplies the examples decides the outcome of an AI evaluation. If the vendor brings the demo data, the procurement has already been settled before the first meeting.
AI selection processes are usually run as a sequence of demonstrations. Each vendor presents, each presents well, and the buying team is left comparing impressions of systems that were each shown at their best on material the vendor chose. Everyone involved knows this is weak, and the usual remedy is a paid pilot — which is expensive, slow, and still scored informally at the end by the same people who watched the demos.
What changed, and why it matters now
The tell is in how a shortlist collapses. Teams that ran demos alone tend to arrive at a preference they find hard to articulate, and the stated reasons drift between meetings: it felt more polished, the team seemed stronger, the roadmap was clearer. Teams that brought their own examples describe their choice in numbers and disagree about weightings rather than about impressions. The second kind of disagreement is productive; the first is how organisations end up in three-year contracts they cannot justify eighteen months later.
Build the operating model
Assemble a held-out set of your own cases with known correct outcomes before any vendor contact, and score every demonstration against it under the same conditions.
Two hundred examples is usually enough and is achievable in a fortnight. Sample them across the distribution you actually see rather than the one you describe: include the routine majority, the seasonal edge cases, the malformed inputs, and the handful that experienced staff argue about. That last group is the most valuable and the most often omitted, because it is the group where the organisation does not agree with itself. Establish the correct answers internally first, and record the disagreements — a case your own experts split on cannot be counted against a vendor, but it tells you something more useful about where the process is ambiguous.
Measure what the decision produced
Score accuracy on the held-out set, the error profile by category, the cost per thousand cases at your volume, and the latency at your peak rather than the average.
The error profile matters more than the headline accuracy, because errors are not interchangeable. A system at ninety-two per cent that fails randomly is usually better than one at ninety-four that fails systematically on a category representing a fifth of your revenue. Report per-category performance and refuse a single aggregate number. On cost, model your real distribution, including retries and the long inputs, since per-unit pricing quoted against a short average input understates the bill by a wide margin at production volume.
Where execution breaks
The main risk is building the set from clean historical records, which produces a benchmark that no longer resembles the messy inputs the system will actually receive.
The second risk is leaking the set. Once examples are shared with a vendor for a pilot, they cease to be a held-out measure for any subsequent round, and a renewal evaluated on them is measuring memorisation. Keep a reserve partition that is never shared with anyone, use it only at renewal, and treat its existence as confidential. This is unglamorous administrative hygiene, and it is the difference between a renewal decision based on evidence and one based on the incumbent's familiarity.
What this looks like in practice
In practice the exercise takes one analyst two weeks and changes the procurement completely. Vendors are sent the same hundred cases, return outputs in a fixed format, and are scored by someone who did not attend the demos. The remaining hundred stay in reserve. The conversations that follow are noticeably different: vendors ask about the error categories rather than the timeline, and the ones who engage seriously with a category they scored poorly on are usually the ones worth shortlisting. It also changes the internal conversation. A procurement committee that has argued for weeks about which system feels more capable can resolve the argument in an afternoon once the same hundred cases have been run through each, and the residual disagreement is about how much a particular error category matters to the business — which is a decision the committee is actually qualified to make, and the one it should have been spending its time on. One practical caution: fix the output format before sending anything. Vendors asked for free-form responses will each return a different shape, and the scoring then becomes an exercise in interpretation that reintroduces exactly the subjectivity the set was built to remove. A short schema and one worked example is enough, and it takes an hour.
The strongest argument against this
The fair objection is that a held-out set rewards the capability you can already measure and penalises the one you cannot. Systems differ in ways an accuracy score does not capture — how they fail, how they explain themselves, how they behave on inputs unlike anything in your history — and a purely quantitative process can select a benchmark-optimised product over a genuinely better one.
So the set should decide the shortlist, not the winner. Use it to eliminate systems that cannot do the work, then choose between the survivors on the qualitative criteria that scores cannot reach: support model, data terms, exit cost, and whether the team answers hard questions straight. What the set removes is the possibility of choosing on impressions alone, which is a smaller claim than it sounds and a much larger improvement than it sounds.
A 30-day implementation sequence
Sample two hundred real cases with known outcomes this fortnight, split them in half, and send the first hundred to every vendor on the list.
Days one to three, define the categories and the sampling frame. Days four to eight, pull the cases and have two experienced people label them independently — the disagreement rate is your ceiling for any vendor. Days nine to ten, resolve or exclude disagreements and freeze the set. Day eleven, split and lock the reserve half somewhere the procurement team cannot casually access. Day twelve onwards, run every vendor through the same hundred in the same format, scored blind.
Keep the reserve half genuinely reserved
Record who has access to the reserve partition and when it was last used. The failure is rarely deliberate: someone needs a quick sanity check before a renewal, pulls the nearest available set, and the reserve is spent without anyone deciding to spend it. Treat it as a controlled asset with a named owner and a log, and refresh perhaps a fifth of it annually so it tracks the drift in what you actually receive. A set assembled three years ago and never refreshed is measuring a workload that no longer exists.
Re-score the incumbent against the reserve half at every renewal and against a fresh sample annually. Read the two together: strong performance on the original set with weaker performance on the fresh sample is the signature of drift, either in the vendor's system or in your own inputs, and the distinction is worth establishing before the contract discussion rather than during it.
Editorial conclusion
Procurement processes for AI are elaborate in every dimension except the one that decides quality. Two weeks of an analyst's time, spent before the first vendor call, converts a decision made on impressions into one made on evidence — and leaves the organisation with an instrument it can still use at renewal, which is the part nobody expects and everybody ends up needing.
Practical checklist
- First move — Sample two hundred real cases with known outcomes this fortnight, split them in half, and send the first hundred to every vendor on the list.
- What to measure — Score accuracy on the held-out set, the error profile by category, the cost per thousand cases at your volume, and the latency at your peak rather than the average.
- Failure mode to watch — The main risk is building the set from clean historical records, which produces a benchmark that no longer resembles the messy inputs the system will actually receive.
- Assign a visible owner and a review date.
- Separate evidence from interpretation.
- Capture a baseline before changing the process.
Questions and answers
How do you evaluate AI vendors objectively?
Build a held-out set of your own cases with known outcomes before any vendor contact, send every vendor the same subset in the same format, and have the results scored by someone who did not attend the demos.
How many examples does an AI evaluation set need?
Around two hundred is usually enough and takes one analyst about two weeks. Sample across the distribution you actually receive, including malformed inputs and the cases your own experts disagree about.
Why keep half of the evaluation set in reserve?
Once examples are shared with a vendor they stop being a held-out measure, and any renewal scored on them is measuring memorisation. A reserve partition that is never shared is what makes the renewal decision evidence-based.
Is accuracy the right measure for an AI vendor?
Not on its own. The error profile matters more, because errors are not interchangeable: a system at ninety-two per cent failing randomly is usually better than one at ninety-four failing systematically on a category that carries a fifth of revenue.
Should the evaluation set decide the winner?
It should decide the shortlist. Use it to eliminate systems that cannot do the work, then choose between survivors on support model, data terms, exit cost, and how straight the team answers hard questions.
