Writing

Realtime Is More Than Opening a WebSocket

An open connection is not a synchronised client. Reconnect, missed-event recovery, ordering and idempotency are where realtime systems actually get hard.

The first realtime feature anyone builds works immediately, and that is the problem. You open a WebSocket, push a message, the UI updates, and it feels finished. The demo is honest about what it shows and silent about everything it does not.

What it does not show is the ninety seconds the client spent in a tunnel. That is where realtime systems are actually built, and almost none of that work is about the transport.

An open socket is not a synchronised client

The most expensive assumption in this space is that connection state and data state are the same thing. They are not, and the gap between them is where events go missing.

A connection can be open while the client is behind, because it reconnected and nobody told it what it missed. It can be open while the client is ahead, holding an optimistic update the server rejected. It can be open and perfectly healthy while the client holds state from a session two hours ago that nothing has invalidated. Most alarmingly, a connection can appear open on the client while the server considers it long gone, and neither party finds out until something is sent.

So the question to design against is not "are we connected" but "is this client's view of the world correct, and how would it know if it were not".

The ninety seconds are the whole design

Pick a concrete failure and design the recovery before building the happy path. Mine is: a responder's phone loses signal for ninety seconds in the middle of an incident.

During that window the server does not stop. Alerts fire, someone else acknowledges something, an escalation moves on. When the phone comes back, it has stale state, no knowledge that it is stale, and a user who is about to act on what the screen says.

There are three common answers and only one of them holds up.

Push what you missed from memory. The server keeps a per-connection buffer of undelivered messages. This works until the server restarts, the client is away longer than the buffer, or the client reconnects to a different instance. It is the default because it is easy, and it fails exactly when the system is already under stress.

Refetch everything on reconnect. Simple and correct-ish, but it is a second code path that only runs in the bad case, which means it is the least tested code in the system. It also tends to grow: first you refetch the incident list, then you realise you also need the acknowledgements, then the staff roster.

Let the client say where it was. It reports the last event it durably processed and asks for everything after. Recovery and normal operation become the same mechanism. There is no second path to drift out of sync, which is the real win.

The third requires a durable, ordered trail server-side, which is why realtime work tends to drag you into event design whether you intended it or not.

Ordering is not free, and global ordering is usually not what you want

"Messages arrive in order" is true per connection and false in general, and systems break on that distinction. Two transports, a reconnect, or a fan-out through a broker and your guarantee is gone.

The useful move is to decide what ordering you actually need. Global ordering across an entire system is expensive and almost never the requirement. Ordering within one incident, or one conversation, or one device, is usually what matters and is far cheaper to provide. Scope the guarantee to the thing that needs it.

Then make the client resilient to the rest. If a client can handle events arriving slightly out of order within a scope, by sequence number or by ignoring anything older than what it has, a whole category of race condition stops being a race.

Two transports, one meaning

Working with MQTT and WebSockets together clarified something I would now treat as a rule: the transport is an implementation detail, and the moment it stops being one you have a problem.

Different clients want different transports for good reasons. Constrained devices and flaky mobile networks suit one, browsers suit another. That is fine, as long as every transport is carrying the same events with the same meaning and the same identity. The instant a message means something slightly different depending on how it arrived, you no longer have one system, and the bug will surface on whichever path you test least.

Concretely, that means the event shape and its identity are defined once, server-side, and the transports are routes rather than formats. A fallback path that is subtly less reliable than the primary one is worse than no fallback, because it will be trusted equally.

Idempotency, again

Everything above produces duplicates. Reconnect-and-replay sends things the client may already have. Retries re-send. Dual transport can deliver the same event twice by two routes.

So the client has to be built the way the server is: events carry identity, the client records what it has applied, and applying something twice is a no-op. This is not an optimisation to add later. A realtime client without idempotency is a system that is correct only while the network is.

What I would check in review

A short list that catches most of it:

None of those questions are about WebSockets. That is rather the point. The transport is the part that works on the first afternoon. Everything that makes a realtime system trustworthy is what you build around it for the moments the connection is not there.