AI Compass
Compass

Choosing a model

Large or small, hosted or local, open or closed: which question forces which answer, and what a switch costs.

·1 min read·By Redaktion KI-Kompass
DETAIL
2 sections

By task

TaskSize
Classifying, assigning categoriessmall
Extracting fields from documentssmall to medium
Summarisingmedium
Drafting customer lettersmedium
Multi-step reasoninglarge or a reasoning model
Drafting and reviewing codemedium to large
Domain questions from your own sourcesmedium plus good retrieval

The last row matters most: for questions from your own sources, retrieval quality decides more than model size.

Hosted or local

HostedLocal
Time to starthoursweeks
Dataleaves the buildingstays
Costper requestper card and operating hour
Qualitythe best availablea step below
Dependenceprovider, prices, availabilityyour own hardware and competence
Pays fromany volumehigh, steady volume

The in-house calculation

A 7-billion-parameter model in int8 on a 300-watt card:

ItemAssumptionResult
Throughput142 tokens per second, 60 % utilisation7.4 m tokens per day
Electricity with cooling0.25 EUR per kWh, PUE 1.31.40 EUR per day
Depreciation8,000 EUR over three years7.30 EUR per day
Costabout 1.18 EUR per million tokens

That figure is the benchmark against a metered service. Not included: operations effort, on-call, and the competence it takes. See Why GPUs, Quantisation and Training versus inference: they explain why the same card carries a completely different cost profile depending on the mode it runs in.

What a switch actually costs

  • Every reviewed prompt holds only for the model it was measured on.
  • Output formats are kept differently; structured outputs need re-checking.
  • Thresholds for automatic acceptance shift.
  • Context length and pricing change how retrieval and chunking should be designed.
  • The AI Act classification can be touched where accuracy or limits change.

Which puts the evaluation set before the switch, not after.

The routing arrangement

If 70 percent of requests are answered equally well by a model at a tenth of the cost, total cost falls to 0.7 × 0.1 + 0.3 × 1.0 = 0.37, a 63 percent reduction. The difficulty is in the classification: it must be cheap and must not misroute hard cases. A conservative rule that escalates when in doubt works well.

Related courses and sources

PaperFreeEN

Distilling the Knowledge in a Neural Network

A large model teaches a small one what it knows. The basis of the small models running in production today.

For anyone deploying small models in production who wants the basis for it.

CourseFree900 minEN

Hugging Face LLM Course

A free technical course on language models. Assumes Python knowledge.

For the step from using models to running open ones yourself.

Hugging FaceGo to offer
ToolFreeEN

Hugging Face model hub

Hundreds of thousands of open models with licence, model card and weights. The first place to look when checking whether a local model is enough for a task.

The licence is on the model card, and not every open model allows commercial use.

Hugging FaceGo to offer
ToolFreeEN

Ollama

Run open models locally, one command per model. The simplest way to try running without a provider at all.

For a first try with a local model, with no provider and no account.

Was this page helpful?
Choosing a model