AI Compass
Compass

What AI costs

A costing you can recompute: tokens, requests, review effort, and the items missing from most calculations.

·2 min read·By Redaktion KI-Kompass
DETAIL
3 sections

One request, computed

A typical RAG request:

ItemTokens
System prompt400
Eight passages at 600 tokens4,800
Question and history300
Answer400
Total5,900

At 3 EUR per million input tokens and 12 EUR per million output tokens: 5,500 × 3 / 1,000,000 + 400 × 12 / 1,000,000 = 0.0165 + 0.0048 = 0.021 EUR.

At a thousand requests a day that is 21 EUR a day, or roughly 640 EUR a month. With prompt caching for the system prompt and frequent passages the input share drops by 30 to 60 percent.

The full calculation

ItemOften forgotten
Request costno
Review time per taskyes
Rework on errors, times the error rateyes
Embedding and index upkeep with RAGyes
Introduction, spread over the useful lifeyes
Ongoing upkeep of templates and evaluation setsyes
Training and refreshersyes

The levers, by effect

LeverOrder of magnitudeQuality loss
Prompt caching30 to 60 % of input costnone
Shorter context, more targeted retrievallinearoften better rather than worse
Smaller model for simple requestsfactor 5 to 20 on that sharemeasurable, must be checked
Routing between models40 to 70 % overallsmall with a conservative rule
Batching when self-hostingfactor 5 to 20 throughputnone
Quantisation when self-hostingfactor 2 to 4small to noticeable

The second row is underrated: ten fitting passages produce better answers than two hundred arbitrary ones and cost a twentieth.

Self-hosting against a service

For a model with seven billion parameters in int8 on a 300 watt card at 60 percent utilisation:

ItemValue
Throughput7.4 million tokens a day
Power including cooling1.40 EUR a day
Depreciation7.30 EUR a day
Cost per million tokensaround 1.18 EUR

Not included: operating effort, on-call, expertise. The comparison with a pay-per-use service only turns in favour of self-hosting at high, steady volume. See Why graphics cards.

Avoiding budget surprises

  • Quotas per user and per application, from day one.
  • Cap maximum prompt and answer length technically, not by policy.
  • Warning thresholds at 50, 80 and 100 percent of the monthly budget.
  • Track cost per completed task, not per request.
  • With agents: a hard step limit, or an endless loop becomes an invoice.

The last point has burned through budgets in days. A step limit is not a detail.

The cost sheet

FREE ACCOUNT

Cost sheet with the calculation

A sheet that turns request cost, review effort and the forgotten items into one number per completed task.

Worksheet2 items

No password needed. We send you a sign-in link. An account does not subscribe you to anything. The newsletter is separate.

Related courses and sources

ArticleFreeEN

AI Index Report

An annual report with sourced figures on models, cost, adoption and regulation. Useful when a statement needs a source rather than an impression.

When a statement needs a source rather than an impression.

Stanford HAIGo to offer
Was this page helpful?
What AI costs