Writing

Why I Prefer an Append-Only Event Trail for Critical Workflows

Critical systems have to answer how they got here, not just where they are. Postgres stays the source of truth and everything else reacts around it.

Most systems are built to answer one question: what is true right now? A row holds the current state, you update it, and the previous value is gone. For most software this is correct and anything else is overhead.

Then you work on something where the second question matters as much as the first. Not "what is the state of this alert" but "how did it get there, who moved it, what did each participant know at the time, and can you prove it". A current-state table cannot answer that, and bolting an audit log onto the side produces two records that disagree the first time someone writes to one and not the other.

The alternative I keep returning to is making the transitions themselves the durable thing, and deriving current state from them.

The boring database is the source of truth

The instinct when you say "event-driven" is to reach for a broker and let it carry the system. I would argue almost the opposite: put the events in Postgres, in the same transaction as the state change, and treat every realtime transport as a delivery mechanism rather than the record.

The reason is ordering and durability. When a state change and the event describing it commit together, there is no window where one exists without the other. No "we updated the row but the publish failed". No "the broker has an event for a transition the database never made". Those windows are where the genuinely awful bugs live, because they produce a system whose history is a plausible-looking lie.

A broker is excellent at fanning a message out to interested parties quickly. It is not a good custodian of what happened. Keeping those two jobs separate is most of the design.

The pattern in practice: the state mutation, the durable event, and the intent to publish all commit in one transaction. A separate process reads committed intents and does the network work. If that process dies mid-flight, it restarts and picks up where it left off, because the thing it is working from is in the database rather than in memory.

Idempotency is the price of admission

Once delivery is a separate concern from truth, you will deliver some things twice. Not occasionally, as a bug, but routinely, as a consequence of retrying anything.

This means every consumer has to be able to see the same event twice and behave as though it saw it once. In practice that means events carry an identity, consumers record what they have processed, and processing checks before acting.

It sounds like bookkeeping, and it is, but it buys something large: retries become safe. Once retries are safe, you can be aggressive about them, and once you are aggressive about retries, a large class of transient failure stops being an incident and becomes a log line.

The failure mode to watch for is the consumer that is idempotent in the happy path and not in the one that matters. Sending a notification twice is mildly embarrassing. Escalating an alert twice, or double-charging, is not. Those are the consumers worth being pedantic about.

Replay is the feature you do not know you need

The quiet benefit of a durable event trail is that a client which has been away can be brought back to correctness without a special code path.

A device drops off the network for ninety seconds. It reconnects. With current-state-only thinking, you now need a reconciliation routine: fetch the current state of everything it might care about, diff it against what it has, hope you got the set right. That routine is bespoke, it is hard to test, and it is usually wrong at the edges.

With an append-only trail, the device says which event it last saw and asks for everything after it. The recovery path and the normal path are the same path. There is no second implementation to keep in sync, which means there is no second implementation to be subtly wrong.

That property is worth more than the audit trail, and it is the one people tend not to anticipate when they decide this is over-engineering.

Explicit transitions beat inferred ones

A related discipline, and the one I would push hardest on in review: write the transition, not the conclusion.

It is tempting to store a status column and let the application work out how it got there. Status went from pending to resolved, so presumably someone resolved it. Except sometimes the reaper resolved it. Sometimes an administrator overrode it. Sometimes it was an automated escalation that found nobody and gave up. Those are different events with different meanings, and a status column flattens them into the same two characters.

When the transitions are explicit, questions you did not plan for become queries. How often does an alert get acknowledged by someone other than the intended responder? How long does the second escalation take compared to the first? You did not design for those questions, but you kept enough information to answer them.

What it costs

I would rather be honest about the trade, because there is one.

You write more code per feature. A change that would be one UPDATE becomes a transition, an event, and a consumer. You store substantially more, and you will eventually need a story for how long you keep it. Queries over derived state need care, and the naive version gets slow sooner than you expect. Debugging requires people to think in terms of a sequence rather than a row, which is a real onboarding cost for anyone who has not worked this way.

Against that: you can answer what happened, you can recover a client without a bespoke routine, and you can retry almost anything without fear.

That trade is clearly wrong for a CRUD admin panel. It is clearly right when a lost transition means a person did not get told something they needed to know. Most systems are somewhere in between, and the useful question is not "should we be event-driven" but "which of our workflows actually need their history, and can we be disciplined about only paying for those".

The architecture is not the goal. Being able to explain, a month later and under scrutiny, exactly how the system reached the state it is in, is the goal.