Troubleshooting
When results get worse: the order in which to check, and the five causes it almost always is.
The order of checks
- 01
Did the model change?
Hold the version and checksum against the log. Many providers swap the model behind the same label.
- 02
Did the prompt change?
System prompt, template, parameters. A change without a re-measurement is the second most common cause.
- 03
Did the inputs change?
New document type, different format, different source, different language.
- 04
Did the data situation change?
With retrieval: new versions, deleted documents, changed permissions.
- 05
Only then: is it the model itself?
That question comes last, not first.
The five most common causes
| Cause | Sign |
|---|---|
| Silent model swap | Everything was fine, then suddenly it was not |
| Changed prompt with no measurement | Degradation from one exact date on |
| Shifted input distribution | Only certain cases affected |
| Stale sources in retrieval | Answers are plausible and out of date |
| Review rate too low | The errors were always there and are only now noticed |
Proving a degradation
- Recompute the frozen evaluation set with the same scoring.
- Compare results per category, not as a single figure. One category can collapse while the mean holds.
- State confidence intervals. Two points at 200 cases is noise.
- Compare the input distribution with the one at the time of the last measurement.
Detecting a distribution shift
A simple and hard test: train a classifier that tells whether a case comes from the period of the last measurement or from current operation. If it reaches a ROC AUC clearly above 0.5, the distributions differ measurably, and every old quality figure no longer holds for today's operation. See MLOps.
When it is the retrieval
In systems with their own sources the cause is usually not the model but the search. Measure separately:
| Metric | What it shows |
|---|---|
| Recall@k | Whether the right passage was found at all |
| Faithfulness | Whether the answer is covered by the passages supplied |
| Refusal rate | Whether it guesses when coverage is missing |
A system with low recall and high faithfulness has a search problem, not a model problem. Any work on the prompt is wasted there. See RAG in depth.
What belongs in the incident report
- From when, the scope affected, how it was noticed.
- What changed, with the date and the source of the change.
- Which decisions rested on faulty outputs.
- Immediate measure, cause, lasting measure.
- What delayed detection, and what is being changed about that.