AI Compass
Compass

Evaluating new tools

How to measure a tool on your own cases rather than on a demo, and which numbers carry a decision.

·1 min read·By Redaktion KI-Kompass
DETAIL
3 sections

The evaluation set

  1. 01

    Collect real cases

    50 to 300 transactions from operations, drawn at random rather than chosen.

  2. 02

    Record the expected result

    For closed tasks the correct answer, for open ones a criteria list.

  3. 03

    Include hard cases

    Edge cases, ambiguous cases, and ones where the right answer is "no information".

  4. 04

    Freeze and version it

    From now on it is the yardstick across model changes.

What gets measured

QuantityFor what
Accuracy per field or categoryThe actual quality figure
Share of unchanged acceptancesPractical usefulness
Refusal rateToo low means guessing
Time per taskAgainst today's baseline
Cost per taskIncluding review effort
Errors nobody would have noticedThe most important category

Comparing two tools

  • On the same cases, not on two samples.
  • Evaluate paired: count the cases where exactly one was right.
  • With a confidence interval. Two percentage points on 200 cases is noise.
  • Report the number of configurations tried. Try twenty prompts and a winner appears by chance.
  • Cost per completed task, not per request.

Review effort belongs in the calculation

A tool saving 20 percent of the time while creating 30 percent review effort is more expensive than the previous process. The full calculation reads:

ItemBeforeAfter
Handling timemeasuredmeasured
Review time0measured
Rework on errorsmeasuredmeasured
Cost per request0measured
Introduction, spread over the useful life0estimated

When a model is used as judge

For open tasks automatic evaluation is the only way to scale. The known biases and their remedies:

  • Longer answers preferred: exclude length from the rubric explicitly.
  • The first option wins: swap the order and average both runs.
  • Own style preferred: use a different model as judge.
  • Before relying on it: measure agreement with human ratings on 50 to 100 cases. Below 70 percent, automatic evaluation is unusable.

See Reading benchmarks.

The vendor questionnaire

FREE ACCOUNT

Vendor questionnaire with a scoring sheet

The questions that must be answered in writing before a purchase, with a sheet that makes the answers comparable.

Worksheet3 items

No password needed. We send you a sign-in link. An account does not subscribe you to anything. The newsletter is separate.

Was this page helpful?
Evaluating new tools