Article

Mopping with the tap running

Good AI in healthcare starts at the source. Because AI on messy data produces messy output. Guaranteed.

AI in healthcareData strategyFHIR & openEHRPrivacy by designEHDS
Mopping with the tap running: good AI in healthcare starts at the source

Challenge

  • A clinician sometimes retypes the same information three times over.
  • One client's data sits scattered across the EHR, Outlook, Karify, Minddistrict and a handful of loose files.
  • Nothing talks to anything else, so time that should have gone to the client goes to administration.

What goes wrong next

  • AI on messy data produces messy output. Guaranteed.
  • Without a single source of truth you keep correcting after the fact, forever.
  • Migrate everything first with no use case attached and support evaporates before anything is visible.

Result

€1.5-2morder of magnitude per year at around 250 clinicians, from automatically registering asynchronous care alone
3 phasesprove the value, expand, then scale and secure

Solution

  1. 01Pick one system that leads. In healthcare that is usually the EHR.
  2. 02Write everything created elsewhere back to that single source as fast as possible, in exactly the right format.
  3. 03Store raw and standardised data separately, in two layers.
  4. 04Start with one use case that carries hard value, then scale from there.

Everyone wants "something with AI". Almost nobody talks about data

Every week I speak with executives, IT people and clinicians in healthcare. Everyone wants "something with AI". Almost nobody talks about the one thing that determines whether that AI will ever work: data.

That is what I want to talk about today. In plain language, no jargon and no hype.

The real problem is not too little data. It is too much data, in too many places

A clinician sometimes retypes the same information three times over. Once in the EHR, once in an email, once in another system. One client's data sits scattered across the EHR, Outlook, Karify, Minddistrict and a handful of loose files.

Nothing talks to anything else. So valuable time, time that should have gone to the client, goes to administration and retyping instead.

Before you even think about AI, you have to solve this. Because AI on messy data produces messy output. Guaranteed.

The first law of data work: a single source of truth

Pick one system that leads. In healthcare that is usually the EHR. Everything created somewhere else, an email, a message, a form, gets written back to that one source as fast as possible, in exactly the right format.

If you do not, you are mopping with the tap running. You keep correcting after the fact, endlessly, because the source is never right.

It sounds simple. It is the foundation almost everything else rests on.

Two words you will hear more often: FHIR and openEHR

No panic, I will keep this short.

FHIR is the language healthcare systems use to talk to each other. It handles the exchange: system A asks for a medication overview, system B delivers it, quickly, in a standardised way and in real time.

openEHR is about how you store data so that it still carries meaning twenty years from now, whatever software you are using by then. It makes the record future-proof and decoupled from any single vendor.

It is not either/or. It is both: FHIR for the traffic, openEHR for durable storage.

Why does this matter to you as a healthcare professional? One sentence: it is how your data stays yours, rather than your software vendor's.

A secret from an entirely different world

Before I worked in healthcare I led data transformations in manufacturing, at Unilever, across 21 programmes and 190 countries. On data, that sector is years ahead of healthcare. And the most important lesson is surprisingly simple: work with two layers.

The two data layers

  • Layer 1, "As Is": all data from every system, stored exactly as it arrives. Raw and unfiltered.
  • Layer 2, "Core": that same data, cleaned up and converted to recognised standards such as FHIR and openEHR, ready for analysis and AI.

Raw data and clean data kept strictly separate. Nothing is lost, and you can always go back to the source. Healthcare can adopt this approach today.

And the AI? Do not start by boiling the ocean

This is where it usually goes wrong. An organisation decides to migrate and standardise all its data first, a project of a year or more. A great deal happens under the bonnet, but nobody sees a result. Support evaporates and the project dies quietly.

Do it the other way round. Start with one use case that carries hard value.

A concrete example we have already built: automatically registering asynchronous care, meaning email contact and digital interactions with clients. Less administration, better records and fewer missed billable activities.

The business case? For an organisation with around 250 clinicians it quickly adds up to a value in the order of 1.5 to 2 million euro per year. And it needs only a few tables out of the EHR. Small project, large result.

Once people see that value, trust appears. And from trust you scale, use case by use case. Not the other way round.

Privacy is not a brake. It is a design choice

"But what about GDPR?" A fair question. The answer is: build privacy in from the start.

Work with pseudonymisation and anonymisation, rule out the risk of re-identification as far as possible, and host AI models internally, inside your own environment, so sensitive data never leaves the organisation. Then you can reuse data responsibly, for instance for wider European research under the European Health Data Space (EHDS).

Good privacy and good AI are not opposites. They come out of the same considered architecture.

The plan, in three phases

  1. Phase 1, prove the value. Start with one concrete use case that pays for itself, such as automatically registering asynchronous care. Small, fast and visible. That is how you build trust and support.
  2. Phase 2, expand. Add use case after use case, partly reusable solutions and partly bespoke ones. The platform grows organically alongside the value it delivers.
  3. Phase 3, scale and secure. Standardise at scale (FHIR, openEHR), get privacy and governance right, and make the architecture ready for what is coming, such as secure data exchange under the European Health Data Space (EHDS).

The core of it, in four sentences

  1. AI in healthcare rarely fails because of the AI. It fails because the data is not in order.
  2. Pick a single source of truth and fix it at the source, or you are mopping with the tap running.
  3. Build on open standards, so your data stays yours.
  4. Start small, deliver visible value and scale from there.

In the end this is not a technology story. It is a care story: less time on administration, more time for the client.

Do you work in healthcare and recognise this? I am glad to think it through with you. Send me a message and I will happily share our approach.

The use case we start with

Automatically registering asynchronous care. Less administration, better records and fewer missed billable activities.

Watch the video · Automatic email registration

Frequently asked questions

The questions this article raises most often.

Why does AI in healthcare fail so often?

Rarely because of the AI itself. It fails because the data is not in order. One client's data sits scattered across the EHR, Outlook, Karify, Minddistrict and loose files, and those systems do not talk to each other. AI on messy data produces messy output.

What does "a single source of truth" mean in healthcare?

It means naming one system as the leading one, in healthcare usually the EHR. Everything created elsewhere, such as an email, a message or a form, is written back to that single source as fast as possible in exactly the right format. If you do not, you keep correcting after the fact, endlessly.

What is the difference between FHIR and openEHR?

FHIR is the language healthcare systems use to talk to each other and handles the exchange between systems, quickly and in real time. openEHR is about how you store data so it still carries meaning twenty years from now, whatever software you use by then. It is not either/or but both: FHIR for the traffic, openEHR for durable storage.

What are the two data layers "As Is" and "Core"?

Layer 1 "As Is" holds all data from every system, stored exactly as it arrives, raw and unfiltered. Layer 2 "Core" holds that same data, cleaned up and converted to recognised standards such as FHIR and openEHR, ready for analysis and AI. Raw and clean data stay strictly separate, so nothing is lost and you can always return to the source.

Where should a healthcare organisation start with AI?

Not by migrating all its data. Start with one use case that carries hard value. Automatically registering asynchronous care, such as email contact with clients, needs only a few tables out of the EHR and at around 250 clinicians quickly adds up to a value in the order of 1.5 to 2 million euro per year.

How does this sit with GDPR?

Privacy is not a brake but a design choice. Work with pseudonymisation and anonymisation, rule out the risk of re-identification as far as possible, and host AI models internally, inside your own environment, so sensitive data never leaves the organisation. That is what makes responsible reuse possible later, for instance for research under the European Health Data Space.

Next step

Do you work in healthcare and recognise this?

I am glad to think through where your data is leaking today and which use case pays for itself fastest. A first conversation is free of charge.