Evaluating new tools
How to measure a tool on your own cases rather than on a demo, and which numbers carry a decision.
The evaluation set
- 01
Collect real cases
50 to 300 transactions from operations, drawn at random rather than chosen.
- 02
Record the expected result
For closed tasks the correct answer, for open ones a criteria list.
- 03
Include hard cases
Edge cases, ambiguous cases, and ones where the right answer is "no information".
- 04
Freeze and version it
From now on it is the yardstick across model changes.
What gets measured
| Quantity | For what |
|---|---|
| Accuracy per field or category | The actual quality figure |
| Share of unchanged acceptances | Practical usefulness |
| Refusal rate | Too low means guessing |
| Time per task | Against today's baseline |
| Cost per task | Including review effort |
| Errors nobody would have noticed | The most important category |
Comparing two tools
- On the same cases, not on two samples.
- Evaluate paired: count the cases where exactly one was right.
- With a confidence interval. Two percentage points on 200 cases is noise.
- Report the number of configurations tried. Try twenty prompts and a winner appears by chance.
- Cost per completed task, not per request.
Review effort belongs in the calculation
A tool saving 20 percent of the time while creating 30 percent review effort is more expensive than the previous process. The full calculation reads:
| Item | Before | After |
|---|---|---|
| Handling time | measured | measured |
| Review time | 0 | measured |
| Rework on errors | measured | measured |
| Cost per request | 0 | measured |
| Introduction, spread over the useful life | 0 | estimated |
When a model is used as judge
For open tasks automatic evaluation is the only way to scale. The known biases and their remedies:
- Longer answers preferred: exclude length from the rubric explicitly.
- The first option wins: swap the order and average both runs.
- Own style preferred: use a different model as judge.
- Before relying on it: measure agreement with human ratings on 50 to 100 cases. Below 70 percent, automatic evaluation is unusable.
See Reading benchmarks.
The vendor questionnaire
FREE ACCOUNT
Vendor questionnaire with a scoring sheet
The questions that must be answered in writing before a purchase, with a sheet that makes the answers comparable.
Worksheet3 items