Alerts: finding the setting between silence and spam
Two settings look similar and do completely different jobs. How often you check decides how quickly a problem could be noticed: that is the interval question. What turns a check into a notification decides whether you will still be reading those notifications in three months.
This page is about the second one, and it is the one that gets abandoned.
The four settings, and the value that works
| Setting | Value that works | What it prevents |
|---|---|---|
| Confirmation before alerting | 2 or 3 consecutive failures | Network hiccups, most of the noise |
| Second observation point | Required before alerting | Path outages mistaken for site outages |
| Recovery notice | Always on | Manual checking "just in case" |
| Reminders during an incident | One message, not twelve | The habit of swiping without reading |
Those four rows settle very nearly the whole problem, and none of them costs more money. They are checkboxes, not a feature. Which is worth stating plainly, because the usual reaction to too much noise is to look for a better tool, when the tool you already have almost certainly offers all four and has them switched off.
The failure mode is not too few alerts
Nobody stops monitoring because they missed something. They stop because their phone buzzed eleven times over a weekend for problems that fixed themselves, and by the third weekend they swipe the notification away without reading it.
That is when the real outage arrives, and it looks exactly like the ten before it.
The mechanism deserves naming, because it is silent: monitoring never dies at once, it degrades into background noise. On the day it stops being useful nothing changes on screen, and nobody decides it.
The four settings, in detail
Confirm before alerting. Require two or three consecutive failures rather than one. A single failed request is often a passing network problem; the cost of waiting is a minute or two of delay, and the benefit is that most false alarms never reach you. It is the setting with the largest effect, and the one people forget to look for.
Check from somewhere else before believing it. If a second location can still reach the site, the problem may be the path rather than the site. Not every tool offers this, and where it exists it removes a whole category of noise: the regional network incidents that have nothing to do with you.
Send the recovery notice. The alert that says it is back is what closes the loop. Without it, every notification leaves a question open, and you go and look manually, which is the habit you were trying to replace. The recovery message should carry the total duration, which is the one piece of information you will need afterwards, not least for the monthly report.
One notification per incident, not per check. If a site is down for six hours, that is one problem. Twelve reminders do not make it more fixed. A single reminder after an hour of continuous downtime is defensible; one every five minutes guarantees the channel gets muted.
What deserves an alert, and what does not
The question comes before the settings, and it is decided site by site.
Deserves an immediate alert: the site stopped answering, it returns a server error, the certificate expires in under seven days, the domain name no longer resolves.
Deserves a daily digest: a slowdown without an outage, a secondary page in error, a certificate at thirty days, a broken link found by a crawl.
Deserves nothing at all: a single failed check followed by immediate recovery, an expected redirect, planned maintenance you triggered yourself.
The practical test, applied to each type of event: if this happens while I am at lunch, do I get up? If the answer is no, it is not an alert, it is a line in a digest. That test also settles arguments quickly, because it moves the question away from how serious an event sounds and towards what anyone would actually do about it.
The channels, and what each is good for
Email suits anything that can wait an hour. It is searchable, archivable, and it wakes nobody. It is the default channel, and for most sites it is the only one needed.
Team chat suits the case where several people need to see the same thing. Mind the trap: a shared channel where everyone sees the alert is a channel where nobody owns it. If you use one, name the responsible person in the message itself.
SMS and push notifications exist to interrupt, and that is their only legitimate use. Keep them for the handful of sites where somebody would genuinely get up, and never put a digest in them. The fastest way to ruin a phone channel is to send one thing through it that could have waited until morning.
Who receives what
Alerts should go to someone who can act, and the list should be short. A notification sent to five people is a notification nobody owns.
For clients, the rule is stricter: send only what they can do something about. A client woken at 3 a.m. who calls you at 9 has been given anxiety, not information, and the monthly report is the right place for everything else.
As soon as the estate runs past a handful of sites, this question stops being individual and becomes a matter of groups: you configure by criticality, not site by site, otherwise the configuration itself becomes impossible to maintain.
About silencing at night
The honest question is not "do I want to be woken", it is "would I actually get up and fix it?".
If the answer is yes, alert at night. If it is no, the alert is noise dressed as diligence: it costs you sleep and changes nothing, and it trains you to ignore the channel.
Different sites deserve different answers here, which is fine: a client shop on a Saturday and an internal tool are not the same thing. The outage will happen either way, and the first fifteen minutes go the same way at three in the morning as at nine, except that at nine you are capable of working through them.
The setting worth revisiting
Alert rules decay. A site gets slower, a threshold that was reasonable starts firing weekly, and nobody adjusts it: they just learn to ignore that one.
Once a quarter, look at what fired and ask a single question about each: did I do anything about this? Everything that fired and changed nothing is either a threshold to fix or an alert to switch off.
Two numbers tell you where you stand. The share of alerts followed by an action should be over half; below a quarter, your monitoring is dying. And the delay between the alert and the first action, which measures your organisation rather than your tool.
The certificate alert deserves a setting of its own, incidentally. Thirty days, seven days, one day: three thresholds, not a daily reminder that ends up ignored. And the thirty-day threshold is useless if you watch the job rather than the date, which is the most common way to see nothing coming.
One last thing is worth checking at the same time, and almost nobody does: that the alerts still go out. A changed address, an expired sending key, an overzealous spam filter, and the channel goes quiet with nothing to report it. It is the same silent failure as the one your host will not warn you about, except that it affects the tool supposed to warn you about everything else.
Frequently asked questions
Should I be alerted on the first failed check?
Usually not. A single failure is often a network hiccup that resolves itself in forty seconds. Two or three consecutive failures from the same place is a real problem, and costs you only a couple of minutes of delay.
Do I want a notification when the site comes back?
Yes, and it is the one people forget to enable. Without it you never know whether an outage lasted two minutes or all night, and you check manually every time.
Should alerts be silenced at night?
Only if you genuinely will not act. Silencing an alert you would have fixed is a loss; sending one you cannot act on until morning is noise. Be honest about which you are.
Never lose a backlink again
Add your sites and links, and let Expansel watch them for you.
Start free