AI Compass
Compass

What AI can do today

A sober stocktake: which tasks work reliably, which work with caveats, and which do not work despite the announcements.

·2 min read·By Redaktion KI-Kompass
DETAIL
4 sections

Three tiers of reliability

TierTasksWhat to do
ReliableSummarising, rephrasing, translating, structuring, drafting, sortingCheck a sample
With caveatsAnswering from your own documents, drafting code, recognising images, transcribing speechCheck every output, demand a source
UnreliableFacts from memory, arithmetic, citations, deadlines, current eventsDo not use without a tool or a source

The first row has a common property: the material is present, and the result can be held against it. That is where the reliability comes from.

What works well, concretely

  • Bringing a ten-page note down to half a page.
  • Rendering a text in plain language or translating it.
  • Drafting an email in three registers.
  • Sorting free text into categories.
  • Extracting actions from a meeting note.
  • Reading a form or a table out of a PDF.

The question that helps

Not "can the model do this?", but:

  1. 01

    How often is it wrong?

    Measured on fifty real cases from your own operation, not on a demo.

  2. 02

    What does an error cost?

    A misrouted ticket costs minutes. A missed deadline costs a case.

  3. 03

    Will the error be noticed?

    A conspicuous error is harmless. A plausible error in a figure nobody recalculates is dangerous.

  4. 04

    What is the baseline?

    How often is the current process wrong? Without that number no progress can be claimed.

What has actually changed in recent years

AreaThenNow
Understanding and producing languageusable in narrow domainsbroadly reliable
Recognising imagesgood with many of your own examplesgood with a few hundred
Multi-step reasoningbarelyusable, expensive
Using toolsabsentwidespread, error-prone
Reliability of factspoorpoor, but solvable with retrieval

The last row matters most: what changed is not factual memory but the ability to bind the model to sources.

Limits that follow from the design

  • Hallucination. The training objective minimises distance to the text distribution, not to truth. Better models lower the rate; they do not remove it.
  • No determinism. Floating-point addition is not associative; identical inputs can produce different outputs on the same card.
  • Arithmetic. Numbers are split into tokens. Reliable arithmetic arises only by delegating to a tool.
  • The middle problem. In long inputs, the middle is measurably used less well than the start and end.
  • Cut-off. A model knows the world up to its training cut-off. Currency arises only through retrieval.

How to read announcements

A new model is typically presented with benchmark figures. Three questions make those figures interpretable:

QuestionWhy
Was the test data in training?Contamination makes figures meaningless
How many variants were tried?Multiple testing creates accidental winners
How wide is the confidence interval?Below three points of difference, usually noise

For a decision in your own organisation none of those figures counts, only the measurement on fifty of your own cases. See Reading benchmarks and Evaluating new tools.

What is likely to be different in two years

Little is defensible. Two observations hold: price per unit of capability has fallen by roughly an order of magnitude per year, and grounding in sources is improving rather than memory. Anyone planning is therefore well advised to plan on falling cost and on architectures that bind sources, not on a model that will eventually know everything.

The suitability sheet

FREE ACCOUNT

Suitability sheet: is AI right for this task?

Six questions to answer before any attempt, with a scoring that gives an answer instead of a feeling.

Worksheet2 items

No password needed. We send you a sign-in link. An account does not subscribe you to anything. The newsletter is separate.

Related courses and sources

CoursePartly free360 minEN

AI for Everyone

Andrew Ng's non-technical introductory course. Around six hours, well suited to managers.

For leaders without a technical background who have to decide rather than build.

DeepLearning.AIGo to offer
ArticleFreeEN

AI Index Report

An annual report with sourced figures on models, cost, adoption and regulation. Useful when a statement needs a source rather than an impression.

When a statement needs a source rather than an impression.

Stanford HAIGo to offer
Was this page helpful?
What AI can do today