AI Compass
Compass

Local and open models

What running your own requires, which hardware suffices for what, and when it pays.

·1 min read·By Redaktion KI-Kompass
DETAIL
1 section

What fits on which card

CardMemoryRealistic
Consumer card8 GB7 bn in int4, short context
Consumer card, large16 to 24 GB7 bn in float16, 13 bn in int8
Workstation card48 GB30 bn in int8, 70 bn in int4
Data centre card80 GB70 bn in int8, several streams
Two to four cards160 to 320 GB70 bn in float16, high concurrency

Add 10 to 20 percent for the runtime plus context memory, which grows with context length and concurrency. See Memory and bandwidth.

What operating it actually requires

  • A serving runtime with continuous batching, quotas and health checks.
  • A warmup after start, because the first request is orders of magnitude slower.
  • Monitoring of utilisation, latency and answer length, not only availability.
  • A model registry with version and checksum.
  • A rehearsed fallback for when a card fails.

When it pays

  1. 01

    Determine volume

    Tokens per day, measured over two weeks, not estimated.

  2. 02

    Compare cost per million tokens

    Around 1.20 EUR in-house at good utilisation against the service price.

  3. 03

    Set utilisation honestly

    A card idling at night still costs. Below about 40 percent utilisation, self-hosting rarely pays.

  4. 04

    Include operations effort

    On-call, updates, competence. That is the item that gets underestimated.

The licence question

  • Does the licence permit commercial use, and above what organisation size?
  • Are there attribution or share-alike obligations?
  • Are the model's outputs restricted, for instance for training other models?
  • Who is liable for infringements through training data? Routinely nobody.

The last point separates open weights from hosted services: several providers of hosted models assume a copyright risk contractually; with open weights nobody does.

Where local models are clearly superior

  • Processing where data may not leave the building.
  • Steady high volume with simple tasks, such as classification.
  • Embedding models for search, which are small anyway and run constantly.
  • Processing on edge devices, see edge AI.
  • Independence as a value in itself, for critical processes.

Related courses and sources

PaperFreeDE

BSI on the security of AI systems

Technical guidance on operating AI systems securely. Free, and unusually concrete.

For security and operations in German-speaking countries; unusually concrete for an official source.

Bundesamt für Sicherheit in der InformationstechnikGo to offer
CourseFree900 minEN

Hugging Face LLM Course

A free technical course on language models. Assumes Python knowledge.

For the step from using models to running open ones yourself.

Hugging FaceGo to offer
ToolFreeEN

Hugging Face model hub

Hundreds of thousands of open models with licence, model card and weights. The first place to look when checking whether a local model is enough for a task.

The licence is on the model card, and not every open model allows commercial use.

Hugging FaceGo to offer
ToolFreeEN

llama.cpp

Running quantised models on ordinary hardware, down to single cards and small boards. The reference implementation for operating without a data centre.

For running without a data centre, down to single cards and small boards.

ToolFreeEN

Ollama

Run open models locally, one command per model. The simplest way to try running without a provider at all.

For a first try with a local model, with no provider and no account.

Was this page helpful?
Local and open models