building

De-identify first: how I handle health data around AI tools

I’m an AHPRA-registered pharmacist who builds AI systems, and this is the actual rule I run, not a compliance summary of one: no identifiable patient data goes near a general AI tool, ever. If it can’t be properly de-identified first, the workflow gets a manual step instead, every time. This is what that costs and why I hold it anyway.

What does “de-identified” actually mean here?

It means stripped clean, not lightly edited. No real names, no addresses, no dates of birth, no anything that lets a reader reverse-engineer who the case is about — not just the obvious fields, but the combination of details that could still point to one person if you know the context. A medication list with an unusual drug combination and a suburb mention can be identifying even with the name removed. Practice cases I build or run through AI tools are synthetic or stripped to that standard. “Anonymised” in the loose, everyday sense isn’t the bar. Properly de-identified is.

The Office of the Australian Information Commissioner, which administers the Privacy Act 1988 and the Australian Privacy Principles, draws the same line: information that has been appropriately de-identified is no longer personal information, which is exactly why the standard has to be strict rather than approximate. OAIC — De-identification and the Privacy Act

Why is this a hard line instead of a judgement call?

Because “just this once” is how the rule stops existing. I could point to plenty of individual cases where skipping the de-identification step would have been low-risk — a quick check, a one-off summary, nobody’s going to see it. But the rule doesn’t work if it’s case-by-case. The moment it becomes a judgement call, it becomes a judgement call I’ll eventually get wrong on a day I’m tired or rushed, and a patient’s data doesn’t get a second chance. So I don’t decide per-case. I decide once, and then I follow it every time, including the times it’s obviously overkill.

Claim: de-identification has to be a standing rule, not a per-instance decision. Evidence: the failure mode isn’t “the AI tool leaks data” — it’s a human deciding, under time pressure, that this particular case is fine to skip. A standing rule removes that decision point. Bottom line: the rule exists to protect me from myself on a bad day, not just to satisfy a policy.

What does this cost you in practice?

Time, mostly. De-identifying a medication list or a set of interview notes properly — checking for the indirect identifiers, not just the name field — takes longer than pasting the raw file in. Some workflows that could technically run end-to-end through an AI tool instead have a manual stripping step in the middle, done by me, before anything touches a model. That’s slower delivery, and it’s a real cost I pay on every relevant build, not a hypothetical one.

It also means some things I won’t do at all. I won’t run a real patient’s raw notes through a general AI tool to save five minutes, even when the output would probably be fine. “Probably fine” isn’t the bar for identifiable health data, and I’d rather be the person who’s slower and boring about this than the one who’s fast and wrong once.

Where does human sign-off sit in this?

Underneath de-identification, not instead of it. Even after data is properly stripped, a de-identified draft or summary that an AI tool produces still gets a human clinical check before it goes anywhere near a real decision. De-identification protects privacy; sign-off protects clinical accuracy. They’re two different gates, and both stay in place regardless of how good the tool is. When I do Home Medicines Review work, the report that goes to a GP carries my name and my AHPRA registration number — that accountability doesn’t move just because a model helped draft it.

Does this slow down everything you build?

No — most of what I build for clients has nothing to do with health data, and this rule doesn’t touch it. Scheduling, drafting, first-pass data entry, general business automation: none of that carries this constraint. It’s specifically identifiable health information that triggers the de-identify-first rule, because that’s the category where a privacy failure is also, often, a clinical trust failure. Outside that category, I move at normal speed.

I hold this line because I was a pharmacist with a professional obligation to protect patient information before I was someone building AI systems, and that ordering doesn’t change just because the tools got more capable. Evidence over hype applies to privacy too — the model being good at summarising isn’t a reason to hand it something it shouldn’t see.

If you’re weighing AI against real health data in your own business, the audit that maps where this rule would bite is on /work.

the build log

Get the build log.

One email when I ship something — a new AI system, a pharmacy workflow, a number from The 2040 Project.

No spam, no drip sequence, unsubscribe in one click.

The weekly dose: graded, sourced, 5 min.Subscribe