Active learning
DATAIt reaches the same quality with 30 to 60 percent of the annotations but makes the dataset unusable as a test set.
See also: Annotation, Test set
Read it in the articleLOOK UP
Short explanations of every term used in this wiki. Each entry links to the article that covers it in full.
It reaches the same quality with 30 to 60 percent of the annotations but makes the dataset unusable as a test set.
See also: Annotation, Test set
Read it in the articleAdam keeps two extra states per parameter. At seven billion parameters that is 56 gigabytes for the optimiser alone.
Read it in the articleAn agent is a loop: think, call a tool, read the result, repeat. Reliability falls exponentially with the number of steps.
See also: Tool call, Prompt injection
Read it in the articleWithout an inventory no duty assessment is possible, because it is unknown what it would apply to.
See also: Approval process, High risk
Read it in the articleThe duty falls on providers and deployers alike and is tied to the concrete use, not to generic training.
See also: Deployer, Human oversight
Read it in the articleAlignment teaches preferences, not knowledge and not truth. The guideline given to raters is the real specification of the behaviour.
Read it in the articleFrom a few seconds per case for classification to a quarter of an hour for segmentation. The factor between is about 200.
See also: Label, Cohen's kappa
Read it in the articlePseudonymisation is not anonymisation: pseudonymised data remains personal and stays subject to the GDPR.
See also: Pseudonymisation, Differential privacy
Read it in the articleWithout an approval process, shadow AI appears: tools in use that nobody knows about and nobody owns.
See also: Evidence, AI inventory
Read it in the articleAn umbrella term, not a single method. In current usage it almost always means systems that learn from data rather than following written rules.
Read it in the articleEvery token queries all keys and mixes the values accordingly. Cost grows quadratically with length.
See also: Transformer, KV cache
Read it in the articleUnder Art. 22 GDPR, rights then arise to human intervention, to express a view, and to contest.
See also: Human oversight, GDPR
Read it in the articleThe backward pass costs about twice the forward pass but needs a multiple of the memory for intermediates.
See also: Chain rule, Gradient
Read it in the articleA base model is not an assistant. Instruction fine-tuning and alignment are what turn it into a system that answers rather than continues.
See also: Fine-tuning, Alignment
Read it in the articleIgnoring the base rate is the most common mistake when judging AI results. For rare events it decides everything.
See also: Bayes' theorem, Precision
Read it in the articleContinuous batching admits a new request as soon as one finishes and often raises throughput five to twentyfold.
See also: Throughput, Latency
Read it in the articleHow often a test fires when someone is affected has little to do with how likely being affected is when the test fires.
See also: Base rate
Read it in the articleDifferences below three percentage points on a general benchmark are worthless for a decision.
See also: Contamination, Evaluation set
Read it in the articleDropping a protected attribute does not remove bias, only its measurability.
See also: Fairness, Representativeness
Read it in the articleThe remedies point in opposite directions. Get the diagnosis wrong and you make things worse.
See also: Variance, Overfitting
Read it in the articleThe vector derived from a face is itself biometric data, even after the image has been deleted.
See also: Embedding, High risk
Read it in the articleBM25 saturates term frequency: the twentieth occurrence adds little over the tenth. That matches reality.
See also: Hybrid search
Read it in the articleA calibration is valid for that rig only. Moving a camera invalidates it.
See also: Disparity
Read it in the articleThe hysteresis step joins broken lines while suppressing isolated noise points.
Read it in the articleThe product form of the chain rule immediately explains why gradients vanish or explode in deep networks.
See also: Gradient, Backpropagation
Read it in the articleChunking decides RAG quality more than the choice of model. Split at headings, not by character count.
Read it in the articleThat lets an image be compared with a sentence without any training for that class.
See also: Embedding, Vision transformer
Read it in the articleBelow 0.6 means the guideline is unclear. A model cannot reach high accuracy on such a task.
See also: Annotation, Label
Read it in the articleThe colour space is a decision, not a property of the image. For colour filters HSV is nearly always right.
See also: Pixel
Read it in the articleIt says how precise the measurement is, not how good the model is. An accuracy figure without an interval is a snapshot.
See also: Standard error
Read it in the articleA model that solves a known task and fails the same task with different numbers has not understood it but seen it.
Read it in the articleThe window covers the system instruction, history, attached documents and the answer so far. A larger window does not mean everything in it is used equally well.
Read it in the articleContours allow counting, measuring and shape checks with no trained model at all.
See also: Canny, Morphology
Read it in the articleDepending on the numbers in the matrix, the result traces edges, blurs, or highlights patterns.
See also: Convolutional network, Kernel
Read it in the articleA CNN brings two assumptions that hold for images: neighbourhood matters, and a pattern means the same everywhere.
See also: Convolution, Vision transformer
Read it in the articleIn the EU, training rests on the text and data mining exception; a purely machine-generated work generally enjoys no protection of its own.
See also: Text and data mining
Read it in the articleTwo texts on the same topic point the same way even if one is ten times longer.
See also: Dot product, Embedding
Read it in the articleCross-entropy includes the entropy of the true distribution as a constant floor; only the remainder depends on the model.
See also: Entropy, KL divergence, Loss function
Read it in the articleWith several cases per person you must split by group, or the same person appears in training and test.
Read it in the articleIt is a required document, not a nicety: the AI Act requires exactly this information for high-risk systems.
See also: Representativeness, High risk
Read it in the articleIt is required for large-scale processing of special categories, systematic monitoring, and profiling with significant effect.
See also: High risk, Fundamental rights impact assessment
Read it in the articleIn most projects, cleaning data delivers more than changing model, and it benefits every future model too.
See also: Duplicate, Cohen's kappa
Read it in the articleServer location alone is not enough: what also matters is whether a parent company in a third country can gain access.
See also: Processing on behalf, Third-country transfer
Read it in the articleDeployer duties are lighter than provider duties but they exist: human oversight, information, logging, suitable input data.
See also: Provider, Human oversight
Read it in the articleA derivative is purely local information. It says what a tiny step does, not where the path leads.
See also: Gradient, Chain rule
Read it in the articleA good descriptor stays the same under rotation and lighting change, enabling matching across images.
See also: Homography, RANSAC
Read it in the articleDice and IoU induce the same ordering and convert into each other, but their values are not comparable.
See also: Segmentation, IoU
Read it in the articleOnly with such a guarantee and a documented budget can anonymity be argued to a supervisory authority.
See also: Synthetic data, Anonymisation
Read it in the articleUnlike a transformer it does not work sequentially but refines the whole result at every step.
See also: Transformer
Read it in the articleDepth error grows with the square of distance: a system accurate to a centimetre at two metres is 25 centimetres out at ten.
See also: Calibration
Read it in the articleDistillation cuts cost and latency substantially and costs measurable quality. How much depends on the task and must be measured against your own evaluation set.
See also: Inference
Read it in the articleThe dot product is why similarity search and attention work at all.
See also: Cosine similarity, Vector
Read it in the articleBecause the optimal policy can be written in closed form through the reward model, that model drops out of the calculation. What remains is stable supervised training.
Read it in the articleThe most common failure is not a crash but silent degradation. Without measurement it surfaces weeks later.
See also: Population stability index, Model registry
Read it in the articleIt usually does not help in large-scale pre-training, because the data is seen only once. In fine-tuning it very much does.
See also: Regularisation
Read it in the articleDuplicates are the most common silent source of error: the same case can land in training and test.
Read it in the articleComputing on the device is a legal advantage before it is a technical one.
See also: Quantisation, Data residency
Read it in the articleAn embedding is not anonymised data: a sentence embedding partially reconstructs its source text.
See also: Vector, Cosine similarity, Vector database
Read it in the articleTranslation and speech recognition use it because the whole input is available before output begins. Language models mostly use the decoder alone, because they keep writing on.
See also: Transformer, Attention, Embedding
Read it in the articleA die has more entropy than a coin, because more can happen. In models, high output entropy indicates guessing.
See also: Cross-entropy, Perplexity
Read it in the articleThe AI Act regulates use rather than technology. The same model can be unproblematic or high risk depending on purpose.
See also: High risk, Provider, Deployer
Read it in the articleFifty of your own cases carry more weight than any external leaderboard. The set is versioned and frozen.
See also: Benchmark, Cross-validation
Read it in the articleEvidence does not arise retrospectively. What was not logged in operation cannot be evidenced later.
See also: Logging, Model registry
Read it in the articleF1 weights both equally. They almost never are: the cost per error type belongs in the metric, not in a footnote.
Read it in the articleSeveral common fairness criteria are provably not simultaneously satisfiable. The choice is normative and requires justification.
See also: Bias
Read it in the articleFine-tuning teaches form, not knowledge. For knowledge, retrieval is cheaper, more current and verifiable.
Read it in the articleMemory falls from quadratic to linear in sequence length. Without it, contexts above 8,000 tokens would be unaffordable.
Read it in the articlefloat16 overflows at 65,504, bfloat16 does not. Which is why bfloat16 won out for training.
See also: Quantisation
Read it in the articleIt sits alongside the data protection impact assessment rather than replacing it; the subject matter overlaps only partly.
See also: High risk, Data protection impact assessment
Read it in the articleThe AI Act does not replace the GDPR. Both apply side by side, with different triggers and different duties.
See also: Legal basis, Purpose limitation
Read it in the articleLearning means shifting every parameter a little against the gradient and repeating that millions of times.
See also: Derivative, Learning rate
Read it in the articleWhen buying, memory decides first, bandwidth second and compute only third.
See also: VRAM, Memory bandwidth
Read it in the articleGrounding is the most effective lever against invented statements. It consists of retrieval, a citation duty, and a check that the claims are actually supported.
See also: RAG, Hallucination
Read it in the articleWith 32 query heads and 8 key-value heads the KV cache is a quarter of the size, at barely measurable quality cost.
Read it in the articleA language model produces the most probable continuation. Where no information exists, a plausible one still emerges, and plausibility is not truth.
See also: Grounding, Temperature
Read it in the articleHigh risk follows from Annex I for safety components and Annex III for named fields such as employment and education.
See also: EU AI Act, Fundamental rights impact assessment
Read it in the articleThe histogram reveals exposure problems in a second and is the cheapest quality check there is.
See also: Pixel, Otsu threshold
Read it in the articleOn a million vectors HNSW finds around 96 percent of the true nearest neighbours in about one millisecond rather than 380.
See also: Vector database, Recall
Read it in the articleFour point pairs determine it uniquely. It is the tool that straightens a document photographed at an angle.
See also: Descriptor, RANSAC
Read it in the articleOversight is not a line in a policy. It needs authority, time, competence and a real way to intervene.
Read it in the articlePure vector search fails on case numbers and proper nouns. The combination regularly lifts recall@10 by 5 to 15 points.
Read it in the articleDuring inference nothing about the model changes. Per operation it is far cheaper than training, but over a lifetime it often costs more in total.
See also: Latency, Throughput
Read it in the articleIoU is stricter than it looks: two visibly overlapping rectangles can land well below the usual 0.5 threshold.
See also: mAP, Non-maximum suppression
Read it in the articleIn classical image processing a human picks the values; in a CNN they are learned.
See also: Convolution
Read it in the articleIn RLHF it measures how far a model has drifted from its starting point, and bounds exactly that.
See also: Cross-entropy, RLHF
Read it in the articleThe cache grows linearly with context length and concurrent requests and overtakes the model itself at long context.
See also: Attention, Context window
Read it in the articleLabel quality caps achievable model quality. Five percent wrong labels means at most 95 percent accuracy.
See also: Annotation, Cohen's kappa
Read it in the articleMore than two seconds to the first character reads as a hang, however fast the rest is.
See also: Throughput, Batching
Read it in the articleUnlike BatchNorm, LayerNorm is independent of batch size and therefore works at batch one.
See also: Residual connection
Read it in the articleLeakage produces excellent measurements and a worthless model. A result that is too good is nearly always a leak.
See also: Cross-validation, Contamination
Read it in the articleToo small means training crawls. Too large means the loss explodes within a few steps.
Read it in the articleWithout a legal basis, processing is unlawful however useful it is. Consent in employment is vulnerable.
See also: GDPR, Purpose limitation
Read it in the articleA complete log of all inputs is tempting and routinely a processing operation with its own legal basis.
See also: Evidence, Retention period
Read it in the articleTemperature divides the logits before they become probabilities, which is why it sharpens or flattens the distribution.
See also: Softmax, Temperature
Read it in the articleMemory need falls by a factor of five to ten, and the result is a file of a few tens of megabytes.
See also: Fine-tuning, Adam
Read it in the articleIt is where business priorities enter the model. Choose it unconsciously and you leave that to a library default.
See also: Cross-entropy, Gradient
Read it in the articlemAP at 0.5 to 0.95 is stricter than mAP at 0.5 and therefore the usual comparison figure.
Read it in the articleA 4096 by 4096 matrix has 16.8 million entries and takes about 33 megabytes in 16-bit representation.
Read it in the articleFor a single request bandwidth is the limit, not compute: the arithmetic units sit over 99 percent idle.
See also: VRAM, Roofline model
Read it in the articleGenerating a single token is memory-bound. Batching is how that changes.
See also: Memory bandwidth, Batching
Read it in the articleThat saves compute, not memory: every expert must be in memory. Which makes the design unsuitable for edge devices.
See also: Parameter, Quantisation
Read it in the articleA model is not a program in the usual sense. It consists of an architecture and the parameters stored in it, which came out of training.
Read it in the articleWithout a registry you can neither roll back nor evidence which version produced a given output.
Read it in the articleMomentum helps leave flat plateaus and saddle points, which are far more common in high dimensions than local minima.
Read it in the articleMorphology repairs what thresholding broke: closing holes, removing fringes, separating connections.
See also: Otsu threshold
Read it in the articleImages are translated into tokens and placed in the context like text. One image can cost more tokens than a page of text.
Read it in the articleThe NMS threshold decides the counted number and therefore belongs in the operating documentation.
See also: IoU
Read it in the articleDividing out the norm is the difference between how similar and how similar and how long.
See also: Vector
Read it in the articleErrors show up as double counting: a broken track becomes two objects.
See also: Optical flow, Non-maximum suppression
Read it in the articleAt 98 percent character accuracy, a 22-character IBAN is fully correct only 64 percent of the time.
See also: Rectification
Read it in the articleOpen weights are not the same as open source: training data and training code are usually absent, and the licence may restrict use.
See also: Model
Read it in the articleFrom a single point only motion across the edge can be determined. That is the aperture problem.
See also: Object tracking
Read it in the articleOtsu works well when the histogram has two clear humps, and otherwise picks an arbitrary point.
See also: Histogram, Morphology
Read it in the articleRecognisable when the training loss falls while the validation loss rises.
See also: Bias and variance, Regularisation
Read it in the articleParameters are adjusted during training and fixed afterwards. Their count sets memory need directly: two bytes per parameter in 16-bit representation.
See also: Model, Quantisation
Read it in the articlePerplexity 8 means as uncertain as guessing among eight equally good continuations. It is not comparable across tokenisers.
See also: Cross-entropy, Tokeniser
Read it in the articleA colour image is a three-dimensional array of height, width and channels. That is all it is.
See also: Colour space, Histogram
Read it in the articleBelow 0.1 counts as unremarkable, above 0.2 as a clear shift that should trigger a check.
See also: Drift
Read it in the articleFor rare events, precision decides usefulness rather than the detection rate.
Read it in the articlePrefill fully loads the card and sets time to first token. Decode afterwards barely loads it.
Read it in the articleEvery language model output is a probability distribution over the whole vocabulary.
See also: Base rate, Bayes' theorem
Read it in the articleIt requires a contract with the mandatory content of Art. 28(3), including sub-processors, deletion and audit rights.
See also: GDPR, Data residency
Read it in the articleIt goes beyond data protection: a processing agreement alone is not enough, the provider must be bound and informed.
See also: Processing on behalf, GDPR
Read it in the articleThis includes social scoring, emotion recognition in the workplace and in education, and biometric categorisation by sensitive attributes.
See also: EU AI Act, Biometric attribute
Read it in the articleA prompt is more than a question. It carries role, task, context, examples and constraints on form. Answer quality depends on it more than on the model.
See also: Context window, System prompt
Read it in the articleThe invariant part must sit at the very front or the cache does not engage. Savings in practice are 30 to 60 percent of input cost.
See also: KV cache, Context window
Read it in the articleAs soon as a system reads external content, that content can contain an instruction aimed at it. Read content is always data and never an instruction.
Read it in the articleWithout a library everyone reinvents their own prompts and nobody knows which version is in use where. It belongs under version control like code.
See also: Prompt, System prompt
Read it in the articleSubstantially modifying a third-party model or offering it under your own name can make you a provider.
Read it in the articlePseudonymising before the model call is the single most effective risk reduction and is usually possible without drawback.
See also: Anonymisation, Purpose limitation
Read it in the articleA PUE of 1.3 means a 30 percent surcharge for cooling and losses. It belongs in every consumption figure.
See also: Edge
Read it in the articleData collected to perform a contract is not thereby usable for model training. A change of purpose must be assessed and justified.
See also: Legal basis, Retention period
Read it in the articleint8 usually costs under one percent of quality, int4 one to three. On arithmetic tasks the loss is disproportionate.
See also: Floating point, VRAM
Read it in the articleWithout a quota the busiest user sets the response time for everyone. It belongs in the first version, not the second.
See also: Throughput, Latency
Read it in the articleThe practical gain is not better language but the ability to name a source and have it verified.
See also: Grounding, Vector database, Chunking
Read it in the articleWithout RANSAC a single bad match ruins the computed transform entirely.
See also: Homography, Descriptor
Read it in the articleIn vector search, recall measures how many of the true nearest neighbours an approximate search found.
Read it in the articleRectification typically lifts OCR accuracy on phone photos from below 60 to above 90 percent.
See also: OCR, Homography
Read it in the articleThe most effective regulariser is more and more varied data. Every algorithmic measure is a substitute for it.
See also: Overfitting, Dropout
Read it in the articleDeviations are normal; unnoticed ones are the problem. Art. 10 AI Act requires an assessment for high-risk systems.
Read it in the articleOnce retrieval returns more than ten candidates, a cross-encoder over 50 lifts precision noticeably.
Read it in the articleWithout residuals every layer has to regenerate its output. With them it only has to say what it wants to change.
See also: Transformer, Layer normalisation
Read it in the articleRetention duties and erasure duties routinely conflict. Both belong set separately per data category.
See also: Purpose limitation, Processing on behalf
Read it in the articlePeople choose between two answers, a reward model is trained on those preferences, and the policy is optimised against it. DPO achieves similar results without that detour.
Read it in the articleWith heavily imbalanced classes ROC AUC looks optimistic; the precision-recall curve is more informative.
Read it in the articleAn unrehearsed rollback takes about three times the estimate when it matters.
See also: Model registry
Read it in the articleIf arithmetic intensity is below the ratio of peak compute to bandwidth, more compute does not help.
See also: Memory bandwidth
Read it in the articleBecause RoPE encodes relative distances, context can be extended afterwards by stretching the frequencies.
See also: Transformer, Context window
Read it in the articleDeliberately collected hard cases improve training and make a test set unusable.
See also: Test set, Representativeness
Read it in the articleIt turned a research question into an investment decision: the effect of ten times the budget can be estimated without building.
Read it in the articleSegmentation costs a multiple of detection in annotation and is usually the bottleneck of a project.
See also: Dice coefficient, IoU
Read it in the articleShadow deployment is the only way to measure a new version on real traffic before it has any effect.
See also: Rollback, Model registry
Read it in the articleMagnitude and direction of the edge follow from the two directional derivatives, as for a vector.
See also: Canny, Convolution
Read it in the articleSoftmax is never computed without subtracting the maximum, or the exponent overflows and the result turns NaN.
See also: Temperature, Logit
Read it in the articleThe method is lossless: the output distribution stays exactly that of the large model.
See also: Batching, Distillation
Read it in the articleTo double measurement precision you need four times the data. That is the most commercially significant number in statistics.
See also: Variance, Confidence interval
Read it in the articleThey do not replace a test set. Evaluation is always against real data, and anonymity without a formal guarantee is unproven.
See also: Differential privacy, Data card
Read it in the articleThe system prompt sets behaviour, tone and limits and is usually invisible to the user. It belongs under version control and documented.
See also: Prompt
Read it in the articleAt zero the most probable token always wins. A high temperature produces not creativity but spread.
Read it in the articleThey engage only with matching formats and dimensions. A width of 4095 rather than 4096 costs noticeably.
See also: Graphics card, Floating point
Read it in the articleAs soon as the test set is used repeatedly to choose between variants it is a second validation set and its number is too optimistic.
See also: Cross-validation, Evaluation set
Read it in the articleLawful access is a precondition. Outside research the exception applies only where no machine-readable reservation was made.
See also: Copyright
Read it in the articleIt needs its own basis: an adequacy decision, standard contractual clauses with a transfer impact assessment, or a derogation.
See also: Data residency, Processing on behalf
Read it in the articleLatency and throughput trade off: batching raises throughput and slows the individual request.
Read it in the articleCost, context length and perplexity are all measured in tokens. German needs roughly 20 to 40 percent more tokens than English for the same content.
See also: Context window, Tokeniser
Read it in the articleGerman needs roughly 20 to 40 percent more tokens for the same content and therefore costs correspondingly more.
Read it in the articleA model is inseparable from its tokeniser. Two models with different vocabularies have no comparable perplexity figures.
See also: Token
Read it in the articleThe model produces structured text read as a function call. A tool's description is itself a prompt and decides how often the right one is picked.
See also: Agent
Read it in the articleTop-p beats top-k because the candidate set adapts to the situation rather than being fixed.
See also: Temperature
Read it in the articleTraining means making a prediction, measuring the error, shifting every parameter a tiny step in the direction that lowers it, and repeating that millions of times.
See also: Gradient, Loss function
Read it in the articleThe 2017 breakthrough is not intelligence but parallelisability: the whole sequence is processed at once.
See also: Attention, Residual connection
Read it in the articleUnder Art. 50 AI Act, providers must label machine-readably and deployers must disclose deepfakes.
See also: EU AI Act, Watermark
Read it in the articleWithout a spread figure, a comparison of two models cannot be told apart from noise.
See also: Standard error, Bias and variance
Read it in the articleWords, sentences and images are translated into vectors so that similarity becomes a question of distance.
See also: Dot product, Embedding
Read it in the articleBelow about 100,000 vectors an extension to your existing database suffices; a dedicated system pays only above that.
See also: Embedding, HNSW, Recall
Read it in the articleWithout a CNN's built-in assumptions, a ViT needs more data but scales better.
See also: Convolutional network, Transformer
Read it in the articleA seven-billion-parameter model needs 14 gigabytes in float16, 7 in int8 and 3.5 in int4, plus context memory.
See also: Quantisation, KV cache
Read it in the articleFor transformers, a warmup over 2 to 5 percent of the steps is effectively standard.
See also: Learning rate
Read it in the articleA watermark is a duty-of-care measure, not protection: it helps with cooperative providers and not against deliberate misuse.
See also: Transparency duty
Read it in the articleThe counterpart is few-shot, where two to five examples sit in the prompt. Where a task fails zero-shot, one example is almost always cheaper than fine-tuning.
See also: Prompt, Fine-tuning
Read it in the article