A user ID that never changes keeps data personal under GDPR, even if hashed

Ulf Aslak Lai photo Ulf Aslak Lai Published 11 Oct 2026 AI drafted 8 min read
A user ID that never changes keeps data personal under GDPR, even if hashed

A common clause in SaaS data processing agreements lets the vendor keep "anonymised usage data" to improve its product, and AI clauses now say the same about training data. In many products that data is an event table where every row still carries the user's ID, hashed perhaps, but the same since signup. On 23 September 2026, Italy's data protection authority, the Garante, fined the health-data company IQVIA 7 million euros for a dataset built on the same idea: no names, only a random code per patient that stayed the same for years 1. The Garante held that a code which never changes ties one person's whole history together, so the data was personal data under the GDPR 1 4. Everything IQVIA had skipped because it believed the data was anonymous then counted as a separate infringement: no legal basis, no notice to the patients, no retention period, no completed risk assessment.

What did the Garante fine IQVIA for?

It fined IQVIA for running a database of health data as if the GDPR did not apply to it. An add-on to the practice software of about 800 family doctors extracted the records of about one million patients and sent them to IQVIA, which used them for studies, some commissioned by pharmaceutical companies 1 3. Each record held year of birth, sex, diagnoses, prescriptions, tests and location data, with the oldest going back to 2001 2. Because IQVIA treated all of this as anonymous, it never found a legal basis, told the patients, set a retention period or completed a risk assessment, and the Garante counted each of those, along with weak security, as a separate infringement 1.

IQVIA has to bring the database into line if it wants to keep it, and BleepingComputer reports that the company says it reserves the right to appeal.

Does your legal documentation match what you ship?

Lawcel watches your product changes and flags the moment your terms or privacy policy fall out of sync, then proposes the edits for your team to approve and publish.

Get started

Why did a random patient code not make the data anonymous?

The code itself revealed nothing, but it never changed, so it tied years of one person's medical events together into a single history. The add-on generated it with a random algorithm, using no personal characteristics, and it changed only when the patient switched doctor, the doctor replaced their computer, or the patient was deleted and entered again 1.

The EDPB, the board where the EU's data protection authorities agree common positions, explains why that matters in its July 2026 draft guidelines on anonymisation. A "record" is everything known about one person, even when it is spread over many rows in a table. The more a record contains, the more likely it is to be unique, and the easier it is to match against what someone already knows about that person. If the person can be picked out that way, the data is not anonymous 5.

The Garante applied that reasoning. Blurring the obvious fields, such as putting ages and locations into bands, does not help when the rest of each record stays detailed: a sequence of diagnoses, hospital stays and prescriptions can be unique on its own, and someone who knows part of it could identify the patient 1.

Two further points from the decision apply to almost any product.

  • Hashing. The Garante said that hashing the patient code would not remove the one-to-one link between a patient and their code 1. I find this the most useful sentence in the decision, because hashing the user ID is the first thing most engineering teams reach for. De Montjoye and colleagues' 2013 study of 1.5 million mobile phone users shows how little it takes once the history stays together: four points of location and time identified 95% of them uniquely.
  • Receiving coded data. IQVIA argued that the data was anonymous for it, because only the doctors' software could link a code back to a patient. It relied on a 2025 EU Court of Justice judgment (EDPS v SRB) that allows coded data to be non-personal for a recipient in some conditions. The Garante read those conditions as strict: the recipient must have no influence over how the data is coded, must not be able to re-identify anyone by reasonable means, and must not pass the data on. IQVIA failed on two counts: it had an active role in defining the anonymisation measures, and the long histories linked by one code made re-identification possible 1.

How did names end up in a database designed to have none?

Through free-text fields. Some doctors typed patients' names, tax codes, addresses and phone numbers into the notes fields of their practice software. The add-on extracted those fields along with everything else, so direct identifiers reached IQVIA's database and the copies it passed to a medical society it worked with. IQVIA notified a data breach covering 3,370 patients, 3,080 of them with health data attached 1.

IQVIA blamed the doctors. The Garante held that writing notes was legitimate patient care, and that it was IQVIA's job to stop its extractor collecting them, with checks before data left the practice and on the database afterwards 1.

This is the part of the case I expect to see repeated most often. Any integration that copies records out of someone else's software, and any analytics or AI pipeline fed from a product's own database, can carry a free-text field along with it. Notes, comments, descriptions and support messages are where people write names.

If your product calls a dataset anonymous, what should you check?

If you build or run a product that keeps user activity under an ID and calls some of that data anonymous, in your Privacy Policy, your DPA, an AI-training clause or a sales answer, that word is a claim a regulator can test against your database. If the test fails, the data was personal all along, and the legal basis, notice and retention period you skipped are each a separate infringement. Here is what I would check.

  • Find every use of the word. Search your legal documents and product copy for "anonymous" and "anonymised", and name the dataset each one refers to.
  • Look for an ID that never changes. If rows about the same user share an ID that lasts across sessions or months, and your company can link that ID to an account, the dataset is personal data for you. Hashing or randomising the ID does not change this.
  • Fix the wording. Replace "anonymised" with "pseudonymised" or a plain description of what you keep and why. EDPB guidance asks controllers not to call data anonymous while people are still identifiable 5. A clause that lets you reuse that data for analytics or model training then needs a legal basis and a line in your Privacy Policy like any other use of personal data.
  • Check data that arrives already coded. A partner's export or an enrichment feed with the names swapped for IDs is not anonymous for you just because you never saw the names. Under the Garante's reading it can only be anonymous if you had no say in how it was coded, cannot re-identify anyone with reasonable means, and do not pass it on.
  • Allowlist what your pipelines extract. Name the fields each integration, export or training pipeline copies, and exclude free text unless you need it. Then check the destination for names and email addresses, because a typed note gets past any schema.
  • If you still want to claim anonymity, test the dataset itself. IQVIA had commissioned two re-identification assessments and accepted their conclusion. The Garante faulted it for relying on those general assessments without checking the database it held 1.
  • Give the anonymisation step a legal basis. Anonymising is itself processing of personal data, so it needs one. The EDPB's draft guidelines presume it is the same basis as before only when the anonymisation is part of the same processing and serves the same purpose 5. Anonymising usage data so you can train a model on it is a new purpose, and needs its own basis or a check that the new purpose is compatible with the old one.

An anonymity claim also goes stale: a later release that adds a field to the event data or keeps a user ID for longer can make it false. That kind of change is what Lawcel, which reads a team's pull requests and flags the ones that affect its legal documents, is built to catch. My co-founder Kenneth's post on the EDPB's draft anonymisation guidelines covers why "anonymous" depends on who holds the data.

FAQ

Yes. Recital 26 of the GDPR says data that could be attributed to a person by using additional information counts as information on an identifiable person. Replacing names with codes lowers the risk and keeps the data inside the regulation.
Not on its own. A hashed ID that stays the same still ties every row about one person together. In the IQVIA decision the Garante said hashing the patient code would not remove that one-to-one link.
Yes. Anonymising is itself processing of personal data. The EDPB's draft anonymisation guidelines and the Garante's IQVIA decision both say so.
No. The Garante required groups of at least 10 only for one route in its IQVIA order, where the doctors anonymise the data themselves. The GDPR sets no fixed number.

References

  1. Garante per la protezione dei dati personali, Provvedimento n. 710 del 23 settembre 2026, IQVIA Solutions Italy S.r.l. (doc. web n. 10302112) - accessed 11 Oct 2026
  2. Garante per la protezione dei dati personali, Comunicato stampa del 2 ottobre 2026, Dati sanitari: 7 milioni di euro di sanzione (doc. web n. 10302141) - accessed 11 Oct 2026
  3. EDPB, Italian DPA fines IQVIA EUR 7 000 000 for unlawful processing of patients' health data, 9 October 2026 - accessed 11 Oct 2026
  4. Regulation (EU) 2016/679 (General Data Protection Regulation) - accessed 11 Oct 2026
  5. EDPB, Guidelines 02/2026 on Anonymisation, version 1.0, adopted 7 July 2026 (public consultation) - accessed 11 Oct 2026

About the author

Ulf Aslak Lai

Ulf Aslak Lai

Co-founder, Chief Technology Officer

Ulf is Co-founder and CTO at Lawcel. He leads engineering architecture for connectors, analysis pipelines, and the safeguards needed when automation touches regulated customer content.

  • Platform architecture
  • Data governance
  • ML/AI systems
GDPR

Your Privacy Policy says "anonymised". The EDPB just asked: for whom?

The EDPB's draft Guidelines 02/2026, out for consultation until 30 October 2026, treat anonymity as relative: the same data can be anonymous for one entity and personal data for another. So the question is not whether your data is anonymous, but for whom, and as of when. For a SaaS team that ships continuously, that turns the word anonymised in a Privacy Policy or DPA into a claim tied to a specific recipient and a date, which ordinary product change can quietly falsify.

GDPR

Does your new feature need a DPIA? Decide before you ship, and write the answer down.

A data protection impact assessment (DPIA) is mandatory under GDPR Article 35 where processing is likely to result in a high risk to people, and it belongs before that processing starts. Guidance from the EU's regulators gives nine criteria and says meeting two usually requires one. Your national regulator also publishes a mandatory list of its own. If a feature needs none, record why.

GDPR

How long can you keep user data? Write down the period, then make something enforce it.

Set a period for each category of data you hold, publish it, and build something that enforces it. GDPR Article 5(1)(e) lets you keep personal data no longer than is necessary for your purpose, and Article 13(2)(a) makes you publish either that period or the criteria you use to set it. Regulator guidance is explicit that "as long as necessary" does not satisfy it.

GDPR

A DSA illegal content report form should accept reports without a name or email

If people in the EU can upload, publish or share content through your product, the Digital Services Act requires a way for anyone to report it as illegal. Final EDPB guidelines of 17 September 2026 say the form should make name and email optional, ask for nothing beyond them, and keep the reporter's name from the uploader unless strictly necessary.