Data protection basics
The six principles that apply to every AI deployment, and the questions they raise in practice.
The six principles and their questions
| Principle | The practical question |
|---|---|
| Lawfulness and transparency | On what basis, and do the people concerned know? |
| Purpose limitation | What was the data collected for, and is that still the purpose? |
| Data minimisation | Does the tool really need everything it gets? |
| Accuracy | What happens if the result is wrong? |
| Storage limitation | When is it deleted, and who ensures that? |
| Integrity and confidentiality | Who can access it, including at the provider? |
When personal data is involved at all
More often than expected. A personal reference arises not only from names but from any detail that makes a person identifiable: a customer number, an IP address, a device identifier, a combination of postcode, age and occupation. Business contacts are personal data too.
The assessment before deployment
- 01
Which data actually flows?
Not the planned data but what ends up in the input. A file note contains names even where nobody intended that.
- 02
Which purpose, and is it covered?
For a new purpose: a compatibility assessment under Art. 6(4) or a legal basis of its own.
- 03
How far can it be minimised?
Pseudonymisation before the model call is nearly always possible and the single most effective lever.
- 04
Who else processes?
Provider, sub-processors, group companies. All three, not just the first.
Data subject rights in the AI context
| Right | Particularity |
|---|---|
| Access | Covers inputs and outputs too, so far as stored |
| Rectification | Hard for model outputs; the source is rectified, not the model |
| Erasure | Reaches stores and indexes, including vector indexes |
| Objection | Always available where legitimate interest is the basis |
| Automated decision | Art. 22 with a right to human intervention |
The points that are new in the AI context
- Training data. An erasure request against a trained model is practically impossible to satisfy. Which is why personal data belongs removed before training, not after.
- Embeddings. They are derived from the source text and partly invertible. The vector index carries the same duties as the store.
- Logs. A complete log of all inputs is its own processing with its own basis and its own retention period.
- Outputs. A model statement about a person is personal data, even if invented. Accuracy under Art. 5(1)(d) applies to it too.
- Third-country element. Not only the server location but also the access possibility of a parent company in a third country.
The fourth point is becoming practically relevant: supervisory authorities have found that an invented statement about a person is inaccurate processing and that a right to rectification arises.
When an impact assessment is required
- Large-scale processing of special categories under Art. 9.
- Systematic large-scale monitoring of publicly accessible areas.
- Systematic and extensive evaluation of personal aspects with significant effect.
- Use of new technologies with high risk, which covers many AI applications.
- Always where a high-risk system under the AI Act processes personal data.
Related courses and sources
Austrian Data Protection Authority
The competent supervisory authority for Austria, with forms, decisions and guidance on reporting breaches.
For controllers in Austria: the competent authority for notifications and enquiries.
Guidelines of the European Data Protection Board
The GDPR as interpreted by the body of supervisory authorities. In a dispute about a legal basis, the most solid source after the text of the law itself.
For legal teams and data protection officers when an interpretation has to hold up.