A tiny gap with a large blast radius
Imagine an order service that writes a new order to Postgres, then publishes an event so another service can reserve stock. Both calls work during testing. In production, the process can stop after the database commit and before the publish. Now the order exists, but the downstream service has no reason to act.
Reversing the calls changes the failure rather than removing it. Publish first, fail the database write, and the downstream service may act on an order that never committed. A retry loop is useful, but it needs a durable record of what to retry. A process that has disappeared cannot remember its unfinished work.
The boxes in this hypothetical order flow are ordinary. That makes it a good architecture exercise: the difficult part is the promise made between them. What, exactly, survives when the process does not?
Text version
The dual-write failure
- Commit order: The database now contains order 42.
- Process stops: No durable publication intent exists.
- No event: Stock reservation never starts.
Changing the order of the calls creates the opposite inconsistency; it does not make the pair atomic.
Store the intent next to the business change
A transactional outbox changes the first step. Write the order and an outbox row in the same database transaction. Commit both or neither. PostgreSQL transactions provide that all-or-nothing boundary within the database; they do not automatically extend it to an unrelated message broker.
The outbox row is the durable intention to publish an event. A relay can read committed rows and deliver them later. It may be a polling worker, or it may use change data capture. The transaction is the important part of the pattern, not the choice of relay.
AWS’s outbox guidance describes this separation and warns that duplicate delivery still needs handling. An outbox fixes the lost intent between the business write and the event. It does not promise that every downstream effect will happen exactly once.
Text version
Database boundary
- Begin transaction: Validate the business operation.
- Order + outbox: Write both records in one transaction.
- Commit: Both become durable, or neither does.
Delivery boundary · after commit
- Relay → broker: Read committed outbox changes; retry delivery.
- Consumer transaction: Deduplicate event ID + apply database effect together.
- Acknowledge: Only after the consumer commit succeeds.
Duplicates are part of the contract
Give each event a stable ID. Do not generate a new ID whenever a publish is retried. At a database-backed consumer, one approach is to insert that ID into a table with a unique constraint and apply the business update in the same transaction. If the ID already exists, skip the repeated effect. A separate ‘seen it’ check followed by a later write leaves another race.
Debezium’s outbox event router exposes an event ID for duplicate detection and an aggregate ID as the message key. That key matters for partitioning and ordering. It is not a global ordering guarantee across all orders, all partitions, or all consumer activity.
There is a boundary worth stating out loud: recording an event ID in your database does not make sending an email or charging a card part of that transaction. For an external effect, use the destination’s idempotency mechanism where available, or design another durable handoff with reconciliation. Otherwise you have simply moved the original two-write problem.
What I would put on the first dashboard
Throughput is useful, but I would start with the age of the oldest undelivered event. Ten thousand events per second can coexist with one order that has been stuck since breakfast. Also watch the retry count, the size of any quarantine queue, and the time from source commit to the consumer’s visible result. These are suggested operating signals, not universal service-level targets.
Write down the replay procedure before it is needed. Which event version will be replayed? Does the consumer still understand it? How long are deduplication records retained? Replaying events older than that retention window can repeat effects you thought were protected. A retention policy is part of the delivery contract.
For a small workload, I would be comfortable starting with a polling relay if someone owns its locking, retries and cleanup. CDC becomes more attractive when existing infrastructure and operational experience make it cheaper to run. ‘Streaming’ on the architecture slide does not settle that trade-off.
The design is ready for a serious review when the team can explain what happens after each possible crash. Not just the happy path from left to right, but the half-finished work somebody has to find tomorrow.

04 / Reader discussion
Continue the conversation.
Add a thoughtful question or perspective. Every comment is reviewed before it appears publicly.
Loading discussion…