Healthcare · 21 Aug 2025
A patient is not a user id
Atalaya Digital
I have seen "anonymized" extracts that still joined back to a person in two hops. I have also seen pipelines that treated a withdrawn consent like a soft delete on a blog comment.
That is not a healthcare system. That is a product analytics stack wearing a white coat.
Three facts the model must hold
For each row that touches a person we store:
- Who — a durable internal key, not an email, not a national id in the mart.
- Purpose — why this processing exists. Care is not marketing. Quality audit is not research.
- Until when — a retention clock, not "we will clean this later."
dim_consent is not optional decoration. Downstream jobs join it. If consent is missing or expired, the row does not land in the board.
Deletes that actually delete
GDPR access and erasure are not a CSV you email to legal. They are jobs:
- find every copy of the person in raw, staging, and mart
- tombstone or drop according to the purpose
- prove it with an audit table your DPO can read
If the "lake" is a pile of untyped JSON, you cannot do this. That is why we are picky about what we ingest and how long raw lives.
Aerospace and automotive taught the same lesson
Telemetry on a test bench and a VIN on a dealer system are not "users" either. The cost of a bad row is not a slide in a deck. You model identity and retention with the same seriousness, even when the regulation is different.
The pattern is the same: know the grain, know the person-or-asset, know when the data must die.
What we refuse
We do not dump a production HIS into BigQuery "for later." We do not put a patient grain on a public dashboard because the chart looked good. We do not promise AI insights on data you do not have a legal basis to process.
If that sounds slow, it is slower than a lawsuit. It is faster than rebuilding the stack after the first audit.