TOOKLI
AI

Evaluating AI features before you ship them

A demo is not a launch criterion. How to build an evaluation set, choose a baseline, and decide whether a model is good enough to put in front of users.

TOOKLI DataData & AI practice2 min read

The failure mode

An AI feature demos beautifully. Someone tries fifteen prompts, all fifteen look impressive, and the feature ships. Three weeks later support is fielding complaints about answers that are confidently wrong, and nobody can say whether the system got worse or was always like this.

The missing piece is not a better model. It is a measurement.

Build the evaluation set first

Before choosing a model, write down 50 to 200 realistic inputs with the answer you would accept. Draw them from real user behaviour if you have it, and from the people who will use the feature if you do not.

Include the awkward cases deliberately: the ambiguous question, the one where the correct answer is "I don't know", the one containing a typo, the one in the second language your users speak.

An evaluation set that only contains easy cases will tell you everything is fine.

Choose a baseline you could actually ship

Compare the model against the simplest thing that could work — a keyword search, a rules engine, the current manual process. Surprisingly often the baseline wins on accuracy, cost, latency and explainability at once.

That is not a wasted exercise. Finding out in week one is the cheapest possible outcome.

Decide the threshold before you see the result

Agree what "good enough" means before running the evaluation, because afterwards the number will look like whatever you need it to look like. Write it down: "we ship at 85% exact match on the evaluation set, with under 2% confidently-wrong answers."

Confidently wrong deserves its own metric. A system that says "I'm not sure" is recoverable. A system that invents a citation is not.

Keep the evaluation running

Prompts change, models get deprecated, providers silently update weights. Run the evaluation in CI on every change to a prompt or a model version, and treat a regression exactly as you would a failing test.

Where humans belong

Where an error carries real cost — money moved, a medical decision, a legal filing — design the human review step into the system from the beginning. Bolting it on later means redesigning the interface, the data model and the audit trail at once.

Related reading