Choosing a model
Large or small, hosted or local, open or closed: which question forces which answer, and what a switch costs.
By task
| Task | Size |
|---|---|
| Classifying, assigning categories | small |
| Extracting fields from documents | small to medium |
| Summarising | medium |
| Drafting customer letters | medium |
| Multi-step reasoning | large or a reasoning model |
| Drafting and reviewing code | medium to large |
| Domain questions from your own sources | medium plus good retrieval |
The last row matters most: for questions from your own sources, retrieval quality decides more than model size.
Hosted or local
| Hosted | Local | |
|---|---|---|
| Time to start | hours | weeks |
| Data | leaves the building | stays |
| Cost | per request | per card and operating hour |
| Quality | the best available | a step below |
| Dependence | provider, prices, availability | your own hardware and competence |
| Pays from | any volume | high, steady volume |
The in-house calculation
A 7-billion-parameter model in int8 on a 300-watt card:
| Item | Assumption | Result |
|---|---|---|
| Throughput | 142 tokens per second, 60 % utilisation | 7.4 m tokens per day |
| Electricity with cooling | 0.25 EUR per kWh, PUE 1.3 | 1.40 EUR per day |
| Depreciation | 8,000 EUR over three years | 7.30 EUR per day |
| Cost | about 1.18 EUR per million tokens |
That figure is the benchmark against a metered service. Not included: operations effort, on-call, and the competence it takes. See Why GPUs, Quantisation and Training versus inference: they explain why the same card carries a completely different cost profile depending on the mode it runs in.
What a switch actually costs
- Every reviewed prompt holds only for the model it was measured on.
- Output formats are kept differently; structured outputs need re-checking.
- Thresholds for automatic acceptance shift.
- Context length and pricing change how retrieval and chunking should be designed.
- The AI Act classification can be touched where accuracy or limits change.
Which puts the evaluation set before the switch, not after.
The routing arrangement
If 70 percent of requests are answered equally well by a model at a tenth of the
cost, total cost falls to 0.7 × 0.1 + 0.3 × 1.0 = 0.37, a 63 percent reduction.
The difficulty is in the classification: it must be cheap and must not misroute
hard cases. A conservative rule that escalates when in doubt works well.
Related courses and sources
Distilling the Knowledge in a Neural Network
A large model teaches a small one what it knows. The basis of the small models running in production today.
For anyone deploying small models in production who wants the basis for it.
Hugging Face LLM Course
A free technical course on language models. Assumes Python knowledge.
For the step from using models to running open ones yourself.
Hugging Face model hub
Hundreds of thousands of open models with licence, model card and weights. The first place to look when checking whether a local model is enough for a task.
The licence is on the model card, and not every open model allows commercial use.
Ollama
Run open models locally, one command per model. The simplest way to try running without a provider at all.
For a first try with a local model, with no provider and no account.