Realtime application · 2025

Relaydesk

Live
  • Realtime
  • Backend
  • Full-stack
RoleLead engineer
Duration5 months
Team3 engineers, 1 designer
StackTypeScript, Node, WebSocket, Redis, Postgres, React

01 — The problemA shared support inbox where four agents can work the same queue without stepping on each other, colliding on replies, or refreshing to find out what changed.

The team had outgrown a shared email account. Two agents would answer the same ticket, a third would reopen something already resolved, and nobody could tell whether a conversation was being handled right now or had been sitting untouched since Tuesday.

02 — ConstraintsWhat I had to work inside

  • Presence had to be accurate within about a second, or agents stop trusting it and go back to asking in chat.
  • Reconnects are constant — laptops sleep, wifi drops, tabs get suspended. Recovery had to be invisible.
  • A dropped event must never silently lose a message. Slow is recoverable; wrong is not.
  • Existing Postgres data model could not be thrown away.

03 — ApproachPick the dumbest transport that works

We started with server-sent events because the traffic is overwhelmingly server-to-client. That held until typing indicators and presence heartbeats made the client chatty, at which point maintaining a separate POST path for every upstream signal was more code than just opening a socket. We moved to plain WebSocket — no framework — with a hand-rolled heartbeat.

04 — ApproachGive every event a sequence number

Every mutation writes to Postgres and appends to a per-conversation Redis stream with a monotonic id. Clients track the last id they processed. On reconnect they send it back and get the gap replayed rather than refetching the world — so a thirty-second tunnel outage costs a few kilobytes instead of a full resync.

client resume
socket.addEventListener('open', () => {
  // resume from the last event we durably applied, not from "now"
  socket.send(JSON.stringify({ type: 'resume', since: store.lastEventId }));
});

socket.addEventListener('message', (e) => {
  const msg = JSON.parse(e.data);

  // a gap means we missed something: fall back to a full refetch
  if (store.lastEventId && msg.prevId !== store.lastEventId) {
    return resyncConversation(msg.conversationId);
  }

  apply(msg);
  store.lastEventId = msg.id;
});

05 — ApproachPresence is a cache, not a database

Presence lives entirely in Redis with a short TTL, refreshed by heartbeat. Nothing about it is durable and nothing needs to be — if Redis restarts, every agent reappears within one heartbeat interval. Trying to persist presence to Postgres was the first thing we built and the first thing we deleted.

  • Key per agent, 30s TTL, refreshed every 10s. Expiry is the disconnect signal.
  • Draft indicators are throttled to one event per 800ms per agent — a keystroke-rate firehose helps nobody.
  • Locks are advisory. A colliding reply warns rather than blocks; agents route around it socially.

06 — DecisionsWhat I chose, and what it cost

WebSocket over SSE

ChoseRaw WebSocket with a custom heartbeat
OverServer-sent events plus POST for upstream
BecauseOnce typing and presence made the client talk back constantly, two transports were more moving parts than one.
What it costWe gave up automatic HTTP reconnection and had to write backoff, resume and gap detection ourselves — roughly two weeks.

Redis Streams over a message broker

ChoseRedis Streams
OverKafka or NATS
BecauseRedis was already deployed for sessions, and streams gave us ordered replay with consumer groups at zero new operational cost.
What it costRetention is memory-bound. We cap at 10k events per conversation and fall back to a Postgres replay past that.

Advisory locks, not hard locks

ChoseWarn on collision
OverBlock the second agent from replying
BecauseHard locks strand conversations whenever a browser dies mid-draft, and the lock-release edge cases outnumbered the collisions.
What it costDouble replies are still possible. They happen a few times a month and are self-correcting.

07 — OutcomesWhere it landed

~180msp95 event deliveryWrite commit to remote client paint.
99.95%Socket uptimeRolling 90 days, excluding client network.
4→11Agents supportedSame architecture, no rewrite.
−70%Duplicate repliesCompared to the shared-mailbox baseline.

08 — RetrospectiveWhat I would tell myself at the start

  • Build reconnection before you build features. Every realtime bug we shipped was really a resume bug.
  • Sequence numbers on events are the cheapest debugging tool in a distributed system. Add them on day one.
  • Ephemeral state wants an ephemeral store. Persisting presence bought us nothing and cost us a migration.
  • Social protocols beat technical locks when the users all sit in the same team chat.