Why GPUs
Why AI runs on graphics cards and not processors: the difference between a few fast and many slow arithmetic units, with numbers.
The idea
A processor has perhaps 16 cores, each very fast and very flexible. A graphics card has several thousand arithmetic units, each simpler and slower, but all performing the same calculation on different numbers at the same time.
When a task consists of a million independent, identical multiplications, the second design wins by two orders of magnitude. A matrix multiplication is exactly such a task, and that is what a neural network consists of.
The comparison in numbers
| Processor (server CPU) | Graphics card (data centre) | |
|---|---|---|
| Cores | 32 to 128 | several thousand |
float16 compute | 2 to 5 TFLOP/s | 300 to 2,000 TFLOP/s |
| Memory bandwidth | 200 to 500 GB/s | 2,000 to 8,000 GB/s |
| Memory | 128 GB to 2 TB | 24 to 192 GB |
| Strength | Branching, sequence | Same operation on many data |
The calculation that determines your needs
For a language model, to a good approximation every weight is read from memory once per generated token. That makes bandwidth, not compute, the limit.
Worked through for a 7-billion-parameter model in float16, that is 14 GB, on a
card with 1,000 GB/s: 1,000/14 ≈ 71 tokens per second. In int8 it is 7 GB and
therefore around 142. That doubling from quantisation costs no extra compute at
all.
Which card for what
| Plan | Memory needed | Note |
|---|---|---|
| 7-billion model, int4 | ~5 GB | Runs on a consumer card |
| 7-billion, float16 | ~15 GB | 16 GB just about, 24 GB comfortable |
| 70-billion, int4 | ~40 GB | One large card or two smaller |
| 70-billion, float16 | ~145 GB | Multiple cards mandatory |
| LoRA fine-tune, 7 billion | ~20 GB | See LoRA and PEFT |
| Full training, 7 billion | ~120 GB | See optimisation |
Why compute is rarely the limit
A card with 1,000 TFLOP/s and 3,000 GB/s has a ratio of about 333 operations per byte. Generating a single token has an arithmetic intensity near 2, more than two orders of magnitude below that. The arithmetic units sit over 99 percent idle.
The remedy is batching: serve 64 requests together and the weights are read once and used 64 times, raising intensity to around 128. Which is why throughput per card with many concurrent requests is ten to fifty times that of a single one. See Speeding up inference.
Cost, worked through
An example for running a 7-billion-parameter model:
| Item | Assumption | Result |
|---|---|---|
| Card | 300 W under load | |
| Utilisation | 60 % across the day | 4.3 kWh per day |
| Electricity | 0.25 EUR per kWh | 1.08 EUR per day |
| Cooling, PUE 1.3 | 30 % surcharge | 1.40 EUR per day |
| Card depreciation | 8,000 EUR over 3 years | 7.30 EUR per day |
| Total | about 8.70 EUR per day |
At 142 tokens per second and 60 percent utilisation that is roughly 7.4 million tokens per day, so about 1.18 EUR per million tokens. That number is the benchmark against a metered API and answers the in-house question factually rather than ideologically. See What AI costs.
Alternatives to GPUs
| Design | Strength | Weakness |
|---|---|---|
| GPU | Universal, mature software | Price, availability |
| TPU and similar | Very efficient on large matrices | Tied to one platform |
| NPU in end devices | Very frugal | Small models, limited memory |
| CPU with extensions | Already present, no purchase | 10 to 50 times slower |
| FPGA | Adaptable, low latency | High development effort |