What AI costs
A costing you can recompute: tokens, requests, review effort, and the items missing from most calculations.
One request, computed
A typical RAG request:
| Item | Tokens |
|---|---|
| System prompt | 400 |
| Eight passages at 600 tokens | 4,800 |
| Question and history | 300 |
| Answer | 400 |
| Total | 5,900 |
At 3 EUR per million input tokens and 12 EUR per million output tokens:
5,500 × 3 / 1,000,000 + 400 × 12 / 1,000,000 = 0.0165 + 0.0048 = 0.021 EUR.
At a thousand requests a day that is 21 EUR a day, or roughly 640 EUR a month. With prompt caching for the system prompt and frequent passages the input share drops by 30 to 60 percent.
The full calculation
| Item | Often forgotten |
|---|---|
| Request cost | no |
| Review time per task | yes |
| Rework on errors, times the error rate | yes |
| Embedding and index upkeep with RAG | yes |
| Introduction, spread over the useful life | yes |
| Ongoing upkeep of templates and evaluation sets | yes |
| Training and refreshers | yes |
The levers, by effect
| Lever | Order of magnitude | Quality loss |
|---|---|---|
| Prompt caching | 30 to 60 % of input cost | none |
| Shorter context, more targeted retrieval | linear | often better rather than worse |
| Smaller model for simple requests | factor 5 to 20 on that share | measurable, must be checked |
| Routing between models | 40 to 70 % overall | small with a conservative rule |
| Batching when self-hosting | factor 5 to 20 throughput | none |
| Quantisation when self-hosting | factor 2 to 4 | small to noticeable |
The second row is underrated: ten fitting passages produce better answers than two hundred arbitrary ones and cost a twentieth.
Self-hosting against a service
For a model with seven billion parameters in int8 on a 300 watt card at 60 percent utilisation:
| Item | Value |
|---|---|
| Throughput | 7.4 million tokens a day |
| Power including cooling | 1.40 EUR a day |
| Depreciation | 7.30 EUR a day |
| Cost per million tokens | around 1.18 EUR |
Not included: operating effort, on-call, expertise. The comparison with a pay-per-use service only turns in favour of self-hosting at high, steady volume. See Why graphics cards.
Avoiding budget surprises
- Quotas per user and per application, from day one.
- Cap maximum prompt and answer length technically, not by policy.
- Warning thresholds at 50, 80 and 100 percent of the monthly budget.
- Track cost per completed task, not per request.
- With agents: a hard step limit, or an endless loop becomes an invoice.
The last point has burned through budgets in days. A step limit is not a detail.
The cost sheet
FREE ACCOUNT
Cost sheet with the calculation
A sheet that turns request cost, review effort and the forgotten items into one number per completed task.
Worksheet2 items
Related courses and sources
AI Index Report
An annual report with sourced figures on models, cost, adoption and regulation. Useful when a statement needs a source rather than an impression.
When a statement needs a source rather than an impression.