AI Compass
Compass

LOOK UP

Glossary

Short explanations of every term used in this wiki. Each entry links to the article that covers it in full.

Adam

MATHEMATICS

Adam keeps two extra states per parameter. At seven billion parameters that is 56 gigabytes for the optimiser alone.

See also: Momentum, LoRA

Read it in the article

Alignment

FUNDAMENTALS

Alignment teaches preferences, not knowledge and not truth. The guideline given to raters is the real specification of the behaviour.

See also: RLHF, DPO

Read it in the article

Artificial intelligence

FUNDAMENTALS

An umbrella term, not a single method. In current usage it almost always means systems that learn from data rather than following written rules.

Read it in the article

Base model

FUNDAMENTALS

A base model is not an assistant. Instruction fine-tuning and alignment are what turn it into a system that answers rather than continues.

See also: Fine-tuning, Alignment

Read it in the article

Bayes' theorem

MATHEMATICS

How often a test fires when someone is affected has little to do with how likely being affected is when the test fires.

See also: Base rate

Read it in the article

Colour space

COMPUTER VISION

The colour space is a decision, not a property of the image. For colour filters HSV is nearly always right.

See also: Pixel

Read it in the article

Confidence interval

MATHEMATICS

It says how precise the measurement is, not how good the model is. An accuracy figure without an interval is a snapshot.

See also: Standard error

Read it in the article

Context window

FUNDAMENTALS

The window covers the system instruction, history, attached documents and the answer so far. A larger window does not mean everything in it is used equally well.

See also: Token, KV cache

Read it in the article

Disparity

COMPUTER VISION

Depth error grows with the square of distance: a system accurate to a centimetre at two metres is 25 centimetres out at ten.

See also: Calibration

Read it in the article

Distillation

FUNDAMENTALS

Distillation cuts cost and latency substantially and costs measurable quality. How much depends on the task and must be measured against your own evaluation set.

See also: Inference

Read it in the article

DPO

FUNDAMENTALS

Because the optimal policy can be written in closed form through the reward model, that model drops out of the calculation. What remains is stable supervised training.

See also: RLHF, Alignment

Read it in the article

Dropout

MATHEMATICS

It usually does not help in large-scale pre-training, because the data is seen only once. In fine-tuning it very much does.

See also: Regularisation

Read it in the article

Encoder-decoder

ARCHITECTURE

Translation and speech recognition use it because the whole input is available before output begins. Language models mostly use the decoder alone, because they keep writing on.

See also: Transformer, Attention, Embedding

Read it in the article

Fairness

DATA

Several common fairness criteria are provably not simultaneously satisfiable. The choice is normative and requires justification.

See also: Bias

Read it in the article

Fine-tuning

ARCHITECTURE

Fine-tuning teaches form, not knowledge. For knowledge, retrieval is cheaper, more current and verifiable.

See also: LoRA, RAG

Read it in the article

FlashAttention

ARCHITECTURE

Memory falls from quadratic to linear in sequence length. Without it, contexts above 8,000 tokens would be unaffordable.

See also: Attention, KV cache

Read it in the article

Grounding

FUNDAMENTALS

Grounding is the most effective lever against invented statements. It consists of retrieval, a citation duty, and a check that the claims are actually supported.

See also: RAG, Hallucination

Read it in the article

Grouped query attention

ARCHITECTURE

With 32 query heads and 8 key-value heads the KV cache is a quarter of the size, at barely measurable quality cost.

See also: KV cache, Attention

Read it in the article

Hallucination

FUNDAMENTALS

A language model produces the most probable continuation. Where no information exists, a plausible one still emerges, and plausibility is not truth.

See also: Grounding, Temperature

Read it in the article

Hybrid search

ARCHITECTURE

Pure vector search fails on case numbers and proper nouns. The combination regularly lifts recall@10 by 5 to 15 points.

See also: RAG, BM25

Read it in the article

Inference

FUNDAMENTALS

During inference nothing about the model changes. Per operation it is far cheaper than training, but over a lifetime it often costs more in total.

See also: Latency, Throughput

Read it in the article

Model

FUNDAMENTALS

A model is not a program in the usual sense. It consists of an architecture and the parameters stored in it, which came out of training.

See also: Parameter, Training

Read it in the article

Momentum

MATHEMATICS

Momentum helps leave flat plateaus and saddle points, which are far more common in high dimensions than local minima.

See also: Gradient, Adam

Read it in the article

Multimodal

FUNDAMENTALS

Images are translated into tokens and placed in the context like text. One image can cost more tokens than a page of text.

See also: Token, Embedding

Read it in the article

Non-maximum suppression

COMPUTER VISION

The NMS threshold decides the counted number and therefore belongs in the operating documentation.

See also: IoU

Read it in the article

Norm

MATHEMATICS

Dividing out the norm is the difference between how similar and how similar and how long.

See also: Vector

Read it in the article

Open weights

FUNDAMENTALS

Open weights are not the same as open source: training data and training code are usually absent, and the licence may restrict use.

See also: Model

Read it in the article

Parameter

FUNDAMENTALS

Parameters are adjusted during training and fixed afterwards. Their count sets memory need directly: two bytes per parameter in 16-bit representation.

See also: Model, Quantisation

Read it in the article

Population stability index

OPERATIONS

Below 0.1 counts as unremarkable, above 0.2 as a clear shift that should trigger a check.

See also: Drift

Read it in the article

Prompt injection

FUNDAMENTALS

As soon as a system reads external content, that content can contain an instruction aimed at it. Read content is always data and never an instruction.

See also: Agent, Tool call

Read it in the article

Prompt library

FUNDAMENTALS

Without a library everyone reinvents their own prompts and nobody knows which version is in use where. It belongs under version control like code.

See also: Prompt, System prompt

Read it in the article

PUE

OPERATIONS

A PUE of 1.3 means a 30 percent surcharge for cooling and losses. It belongs in every consumption figure.

See also: Edge

Read it in the article

Representativeness

DATA

Deviations are normal; unnoticed ones are the problem. Art. 10 AI Act requires an assessment for high-risk systems.

See also: Sample, Bias

Read it in the article

Reranking

ARCHITECTURE

Once retrieval returns more than ten candidates, a cross-encoder over 50 lifts precision noticeably.

See also: RAG, Recall

Read it in the article

RLHF

FUNDAMENTALS

People choose between two answers, a reward model is trained on those preferences, and the policy is optimised against it. DPO achieves similar results without that detour.

See also: Alignment, DPO

Read it in the article

Scaling law

MATHEMATICS

It turned a research question into an investment decision: the effect of ten times the budget can be estimated without building.

See also: Training, Parameter

Read it in the article

System prompt

FUNDAMENTALS

The system prompt sets behaviour, tone and limits and is usually invisible to the user. It belongs under version control and documented.

See also: Prompt

Read it in the article

Text and data mining

LAW

Lawful access is a precondition. Outside research the exception applies only where no machine-readable reservation was made.

See also: Copyright

Read it in the article

Token price

OPERATIONS

German needs roughly 20 to 40 percent more tokens for the same content and therefore costs correspondingly more.

See also: Token, Quota

Read it in the article

Tokeniser

FUNDAMENTALS

A model is inseparable from its tokeniser. Two models with different vocabularies have no comparable perplexity figures.

See also: Token

Read it in the article

Tool call

FUNDAMENTALS

The model produces structured text read as a function call. A tool's description is itself a prompt and decides how often the right one is picked.

See also: Agent

Read it in the article

Training

FUNDAMENTALS

Training means making a prediction, measuring the error, shifting every parameter a tiny step in the direction that lowers it, and repeating that millions of times.

See also: Gradient, Loss function

Read it in the article

Zero-shot

FUNDAMENTALS

The counterpart is few-shot, where two to five examples sit in the prompt. Where a task fails zero-shot, one example is almost always cheaper than fine-tuning.

See also: Prompt, Fine-tuning

Read it in the article
AI OPERATING PLATFORM

You don't have to know all of this yourself.

This wiki explains what needs doing. The LumeSec platform does it: it brings together what your AI systems are doing, holds them inside the agreed limits, and records the evidence as it goes.

See the platform
  • CorrelateOne picture of what your systems are actually doing.
  • ContainThe agreed limits hold at runtime, not just on paper.
  • AttestEvidence accrues in operation, not the week before an audit.
Glossary