Writing

Building Healthcare Software for the Moments When the Workflow Cannot Fail

Healthcare software is not ordinary SaaS with clinical vocabulary. Reliability, authorization, auditability and failure behaviour carry different consequences.

Most software fails politely. A page does not load, someone refreshes, the day continues. I spent a long time building that kind of software and learned the habits that go with it: optimise for the common path, degrade quietly, let the retry sort it out.

Healthcare software is not that, and the difference is not vocabulary. You can take an ordinary SaaS product, rename the entities to patients and encounters, and still have built something that behaves exactly like a project tracker. The difference is what happens at the edges, because in a care setting the edges are the entire point.

The consequence of a dropped message changes everything upstream

In most products, a notification that arrives late is an annoyance. In a behavioural health unit, an alert is a request for a person to physically move toward a room. Late is not a degraded version of on time. It is a different outcome.

Once you accept that, a lot of ordinary architectural decisions stop being matters of taste.

Fire-and-forget delivery is no longer acceptable, because nobody can confirm that the intended responder ever saw the thing. Client-side state as the record is no longer acceptable, because the device that holds the truth is the one most likely to be locked in a pocket or out of signal. "We'll reconcile it later" is no longer acceptable, because the window in which reconciliation matters has already closed.

None of this is exotic engineering. It is mostly the discipline of refusing to let convenience decide, in a domain where the cost of the convenient choice is not paid by you.

One incident, several surfaces, one truth

The architectural decision I keep coming back to from SwiftCode is treating an incident as a single thing that several surfaces observe, rather than as a feature each client implements.

The platform spans an Apple Watch, an iPhone, an iPad and a web dashboard. The temptation, and the thing a normal product roadmap pushes you toward, is to build four applications that each know how to display an alert. Each one gets its own state handling, its own idea of what "acknowledged" means, its own reconnect behaviour. Each is reasonable alone. Together they are four slightly different opinions about the same emergency.

The alternative is less convenient and much calmer to reason about. One incident exists. It has an authoritative state held server-side. Each surface is a view onto it with a different interaction budget: the watch is glanceable and one-handed, the phone is the responder's working surface, the tablet is the station board, the web dashboard is for oversight after the fact. They differ in what they show and what they let you do. They do not differ in what is true.

The practical test is a question worth asking of any multi-client system: if two surfaces disagree, which one is wrong? If the answer requires thought, the truth is in the wrong place.

Authorization is workflow, not configuration

In a normal product, permissions are a settings page. In care software, who may do what is part of the clinical workflow itself, and it changes with shift, unit, role and the state of the incident.

This has a design consequence that took me a while to internalise: authorization checks that live only in application code eventually drift from the data. Different services, different code paths, different assumptions, and one of them is wrong in a way nobody notices until it matters. Pushing the rules down to where the data lives, rather than reimplementing them per service, is less elegant to write and considerably harder to get wrong.

It is also the only version that survives the thing that actually happens to software over time, which is that someone adds a new client.

Auditability is a feature of the architecture, not a log file

"Who acknowledged this, when, and what did they see at the time?" is a question care software has to answer on demand, sometimes long after the fact and sometimes to someone who is not an engineer.

You cannot answer it well from logs, because logs record what the system did rather than what was true. You answer it by making state transitions explicit and durable: the scan, the alert, the acknowledgement, the escalation, each an event that happened rather than a column that changed. The current state becomes something you can derive, and the path to it is the thing you kept.

This is the same append-only argument I would make for any system where the history matters, and it is not free. It costs more storage and more thought. In this domain it also means the compliance conversation is a query rather than a project.

Designing for the failure, not around it

The habit I would most want to carry into any future healthcare work is designing the failure path first.

Not "what if the network drops", answered with a spinner, but: the responder's phone lost signal mid-incident for ninety seconds, and when it returns, what does the device know, what does the server know, and how do they agree again? If the answer depends on the device having stayed awake, it is not an answer.

This is where realtime work stops being about transport and starts being about recovery, which is its own article. The short version is that an open connection is not synchronisation, and treating it as such is how systems quietly lose events they were trusted to deliver.

What software should not try to do

One more thing, because it is easy to get wrong in the other direction.

Good software in this domain reduces coordination friction. It gets the right alert to the right person faster, keeps the record straight, and removes the steps where a human has to remember something a machine could have held. What it should not do is present itself as clinical judgment. The system should make it easier for the right person to decide, and it should be unambiguous that a person is deciding.

The engineering ambition belongs in reliability, in authorization, in the audit trail, and in the behaviour of the system on its worst day. Those are the parts a care team will feel, usually without noticing, which is the point.