← All writing

Who watches the watchman? Building alerts that can't fail silently

A Database monitor went down. No alert was sent. No incident was opened. The only trace was a console.error in a terminal nobody reads.

The monitor was working perfectly the whole time. It checked, it failed, it recorded the failure. What broke was everything *after* that — and nothing was watching that part.

This is the failure mode nobody designs for, because it doesn't look like a failure. Your dashboard is green. Your monitor list says "last checked 30 seconds ago." Every graph is flat and calm. And your API has been down for forty minutes.

Monitoring tools are judged on the wrong axis. Everyone compares check intervals and integration counts. But a monitoring tool has exactly one job, and it happens once every few months at 3am: get the message to a human. Right up until that moment, all monitoring tools look identical.

So the useful question isn't "what can it check?" It's: can you prove the alert will arrive?

Here are the five ways it silently doesn't — and how to test for each one.


1. The alert is sent once, into the void

The most common design: a check fails, and the code calls the webhook. If that call throws, it's caught, logged, and forgotten. If the process restarts mid-send, the alert never existed.

This is fine 99.9% of the time. The problem is that outages *correlate* with the conditions that break outbound requests — network trouble, a struggling host, a provider having a bad day. The moment you most need the alert is the moment it's most likely to fail.

The fix is an outbox. Write the alert to durable storage *before* attempting to send it. A separate worker drains the outbox and retries with backoff. A crash, a restart, a dead webhook, a DNS hiccup — the alert survives all of it, because it exists independently of the attempt to send.

Uplance writes every alert to the outbox first and retries on a ladder: 15 s → 1 m → 5 m → 15 m → 30 m, roughly fifty minutes of trying. If your Discord is having an outage at the same moment your API is, the alert still lands when Discord comes back.

How to test it: point an alert channel at a URL that returns 500, trigger a down event, then fix the URL. Did the alert eventually arrive? If not, your alerts are fire-and-forget.


2. The channel broke weeks ago

You set up a Discord webhook in March. In June, someone deleted that channel, or rotated the webhook, or the integration got removed during a server cleanup.

Your monitoring tool has been sending alerts into a 404 ever since. Nothing told you, because nothing was checking — and no outage happened in between to reveal it.

This one is genuinely dangerous, because it's invisible for an arbitrarily long time and it fails *exactly* when you're relying on it.

The fix is a drill. Once a week, send a real test event through every configured channel and report which ones failed. Not a config validation — an actual delivery, end to end, the same code path a real alert takes.

Uplance runs this weekly and shows you the result: which channels delivered, which didn't. The webhook you rotated gets found on a Tuesday afternoon, not during an incident.

How to test it: delete one of your alert channels at its destination — remove the Discord webhook, revoke the Slack app — and leave it. How long until your monitoring tells you? If the answer is "the next outage," that's a problem.


3. The alert pipeline is degraded and the UI looks fine

A queue is backed up. The mail server is refusing connections. The outbox has two hundred undelivered messages.

Most tools express this as a metric, or a log line, or nothing. You'd have to go looking, and you have no reason to go looking, because the dashboard is green.

The fix is refusing to look healthy when you aren't. If deliveries are failing, the interface that people actually look at should say so, in words, at the top.

Uplance shows "Alerts are at risk" in the dashboard when the pipeline is in trouble. It's a sentence on a screen, not a number on a graph you'd need to already suspect something to check.

How to test it: break every configured channel at once. Does the product's main screen change? Or does it keep showing you a calm, green, entirely fictional picture?


4. The checker itself stopped

This is the deepest version of the problem: the monitoring process died.

No checks are running. No checks are failing. No alerts are firing — correctly, from the system's point of view, because nothing has been observed to fail. The status page shows the last known state, which was green, and will keep showing it forever.

Stale green is worse than red. Red tells you something. Stale green actively lies.

There are two halves to fixing it:

Expose liveness to the outside. Uplance's /api/health returns 503 when the checker has stopped. Point a free external heartbeat service at it — one of the many that exist precisely for this — and now something outside your infrastructure can page you about your monitoring. Monitoring that can only report on itself is a closed loop with an obvious hole in it.

Say so publicly. When the checker has stopped, an Uplance status page stops claiming everything is fine and says "Monitoring has stopped." Your customers would rather read that than a green tick that means nothing. So would you.

How to test it: kill the monitoring process. Within a few minutes, does anything anywhere tell you? Does the public page still claim you're up?


5. The alert loses a race with the bookkeeping

Subtle, and it bites well-built systems specifically.

A check fails. The code opens an incident, captures diagnostic evidence, updates the status page, computes the impact — and then sends the alert. It's clean, ordered, sensible code.

Then one of those earlier steps throws. A traceroute hangs. A database write deadlocks. The exception propagates, and the alert — the one thing that actually mattered — never gets sent. The richer your incident tooling, the more places there are for this to happen.

The fix is ordering by importance, not by narrative. Send the alert first and unconditionally. Everything else is enrichment and happens afterwards, where its failure is survivable.

Uplance sends the down alert before opening the incident and before capturing DNS, TLS and traceroute evidence. If the evidence capture fails, you get a slightly thinner incident. You still get woken up.

How to test it: make incident creation throw. Does the alert still go out?


The general principle

Every one of these is the same bug wearing a different hat: the system reports on everything except itself.

It's an easy mistake, because self-monitoring feels redundant — surely the monitoring is the thing that works? But monitoring is software, and software fails, and software that fails silently while displaying green is the worst kind.

So the questions worth asking about any monitoring tool, including this one:

  1. If the send fails, does the alert survive? *(outbox + retries)*
  2. If a channel breaks quietly, when do I find out? *(drills)*
  3. If the pipeline degrades, does the screen I look at change? *(honest UI)*
  4. If the checker dies, can anything outside it tell me? *(503 health + a public admission)*
  5. If the bookkeeping fails, does the alert still go? *(ordering)*

Run those five tests against whatever you use today. Most tools fail at least three of them, and they fail invisibly, which is why you don't already know.


Uplance is built around those five answers. It also does the ordinary things well — fifteen check types including databases, cron heartbeats, browser flows and game servers; status pages on your own domain with 90 days of history; incident response inside your own Discord bot; monthly SLA PDFs.

But the ordinary things are why people try it. The five answers above are why they stay.

[Free for 10 monitors, no card →](https://uplance.cloud)

Uplance is built this way.

An outbox before every send, a retry ladder up to about fifty minutes, a weekly test event through every channel, and a health endpoint that returns 503 when the checker itself stops.

10 monitors, no credit card