Prompt library
How scattered prompts become a maintained stock: structure, approval, measurement and upkeep.
The entry
| Field | Content |
|---|---|
| Identifier and version | procurement-quote-extract v3 |
| Purpose and boundary | What for, and expressly what not for |
| Owner | A person, not a mailbox |
| The prompt itself | With placeholders in capitals |
| Expected output | Schema or example |
| Tested against | Evaluation set with identifier and date |
| Result | Hit rate per field, with date |
| Approved for | Model and version, temperature, context length |
| Change history | Who, when, why |
Without the last four rows a template has not been demonstrably tested.
The sequence
- 01
Draft
From practice, by the person who does the task.
- 02
Measure
Against the evaluation set for that task. No number, no approval.
- 03
Approve
By the responsible person, with a date and a model reference.
- 04
Distribute
Through the tool, not by email. Otherwise versions drift apart.
Why a model change devalues the library
A new model holds to different formats, picks different default lengths, and deviates differently on edge cases. A template that reached 96 percent on the old model can sit at 88 on the new one without a word of the text having changed.
- Agree a notification duty for model changes with the provider.
- Before switching, compute every template against its evaluation set.
- Rework or block templates that drop, rather than letting them run on silently.
- Keep the old model version available until the switch has been checked.
What the library does beyond that
- Training. A tested template shows what a good prompt looks like better than an explanation does.
- Evidence. Which template in which version handled a task belongs in the log.
- Reuse. A pattern that works in procurement often works in accounting with different fields.
- Clearing out. Templates unused for six months get reviewed and usually removed.
Upkeep
| Trigger | What to do |
|---|---|
| Model change | Re-measure everything |
| New task | Check whether an existing template fits |
| Cluster of errors | Check the template before suspecting the model |
| Every six months | Review the stock, remove the unused |
| Legal change | Have templates with a legal bearing reviewed again |
The third row matters most in practice: on a degradation the template is the first suspect, not the model.
Related courses and sources
OWASP Top 10 for LLM applications
The list of weaknesses that actually occur in systems built on language models, from prompt injection to insecure tool integration.
For anyone building. The list replaces no review, but it is the best starting point for one.