Local and open models
What running your own requires, which hardware suffices for what, and when it pays.
What fits on which card
| Card | Memory | Realistic |
|---|---|---|
| Consumer card | 8 GB | 7 bn in int4, short context |
| Consumer card, large | 16 to 24 GB | 7 bn in float16, 13 bn in int8 |
| Workstation card | 48 GB | 30 bn in int8, 70 bn in int4 |
| Data centre card | 80 GB | 70 bn in int8, several streams |
| Two to four cards | 160 to 320 GB | 70 bn in float16, high concurrency |
Add 10 to 20 percent for the runtime plus context memory, which grows with context length and concurrency. See Memory and bandwidth.
What operating it actually requires
- A serving runtime with continuous batching, quotas and health checks.
- A warmup after start, because the first request is orders of magnitude slower.
- Monitoring of utilisation, latency and answer length, not only availability.
- A model registry with version and checksum.
- A rehearsed fallback for when a card fails.
When it pays
- 01
Determine volume
Tokens per day, measured over two weeks, not estimated.
- 02
Compare cost per million tokens
Around 1.20 EUR in-house at good utilisation against the service price.
- 03
Set utilisation honestly
A card idling at night still costs. Below about 40 percent utilisation, self-hosting rarely pays.
- 04
Include operations effort
On-call, updates, competence. That is the item that gets underestimated.
The licence question
- Does the licence permit commercial use, and above what organisation size?
- Are there attribution or share-alike obligations?
- Are the model's outputs restricted, for instance for training other models?
- Who is liable for infringements through training data? Routinely nobody.
The last point separates open weights from hosted services: several providers of hosted models assume a copyright risk contractually; with open weights nobody does.
Where local models are clearly superior
- Processing where data may not leave the building.
- Steady high volume with simple tasks, such as classification.
- Embedding models for search, which are small anyway and run constantly.
- Processing on edge devices, see edge AI.
- Independence as a value in itself, for critical processes.
Related courses and sources
BSI on the security of AI systems
Technical guidance on operating AI systems securely. Free, and unusually concrete.
For security and operations in German-speaking countries; unusually concrete for an official source.
Hugging Face LLM Course
A free technical course on language models. Assumes Python knowledge.
For the step from using models to running open ones yourself.
Hugging Face model hub
Hundreds of thousands of open models with licence, model card and weights. The first place to look when checking whether a local model is enough for a task.
The licence is on the model card, and not every open model allows commercial use.
llama.cpp
Running quantised models on ordinary hardware, down to single cards and small boards. The reference implementation for operating without a data centre.
For running without a data centre, down to single cards and small boards.
Ollama
Run open models locally, one command per model. The simplest way to try running without a provider at all.
For a first try with a local model, with no provider and no account.