
I mean this sincerely when I say, you have got to get your alerting under control. Excessive alerting is a sign either of your service being out of control, or your alerts being misconfigured. Alert fatigue reduces reliability and destroys morale of on-call engineers. Here is why.
Why over alerting is an enemy of SRE
It reduces reliability
- Alerts that you have seen before, possibly multiple times per week, lack importance. If you get twice daily alerts about a replication lag, it’s tempting to simply acknowledge it and expect it to self resolve. Especially if the alert comes in out of hours.
- If it’s the case that certain types of events always self heal, they are rarely a good candidate for alerting. These should be downgraded to email notifications or silently logged in order that history can be searched and trended. They should not page on-call engineers.
- If an on-call engineer receives more than two alerts per on-call shift, they have a limited amount of time to investigate the root cause and put in proactive fixes. Engineers with a busy pager are always firefighting, and end up reworking the same incidents time after time.
Over alerting destroys morale
- Out of hours alerts interrupt sleep patterns and this makes engineers unwell.
- Over alerting causes the engineer to be in a state of high alert, especially if they are responsible for a mission critical service.
- Regular out of hours alerting can make life difficult for the engineer’s family. Especially their partner who gets woken up as well due to the blaring page notification.
- All of these lead to a fear of the next out of hours rotation.
Why does alert creep happen
In my experience, I put this down to two factors. Good intentions and incident post mortems. Starting with good intentions, I’ve done it myself. I’ve thought, you know what will help, a really fine grained alert about X condition. There’s a serious flaw with this. You are motivated for the information at the point in time you create the alert, but when it starts firing multiple times per day or week, you know it is safe to ignore and it stops receiving the action it needs.
The second and most prolific factor in my opinion is incident post mortems. In SRE, post mortems should not be about blame. That being said, when you have to present at an incident review meeting, often to people who don’t work closely with you, it’s natural to want to be seen as competent and not the cause of maxing out the error budget. You also really don’t want repeat incidents so the tendency is to identify reporting gaps. This insures you for the future and makes you look competent, the trouble is that the burden on the on-call engineer grows. More benign alerts, especially overnight, are not the friend of on-call engineers.
Simple steps to fix alert fatigue
Trim and downgrade
Do your alerts line up with your Service Level Objectives (SLO)? There is no need to alert on issues that don’t interrupt your service or uptime. They should be focussed on issues that are eating into your error budget, and thresholds need regular evaluation.
It’s important to trim alerts regularly. During a recent “fixit day” (where meetings are banned and engineers should deal with nagging problems), I relegated 4 alerts from pages to Slack alerts. There is no expectation to check Slack when on call, meaning only the important issues come through the paging service. Don’t let alerts die in Slack or email. Even these need regular evaluation and lifecycle management.
As well as this, some alerts are just no longer useful and should be cut altogether. Don’t be hasty about this but evaluate, can anything actually be done to resolve the issue or is this just information? If it’s information, it’s better for everyone to cut it.
Fair rotation and fair notice
This advice is for engineering managers. Most rotations are fair. If you have x people, they are on every x weeks. Possibly you run a secondary or a working hours schedule as well. Whatever you do, rotations for employees should be equally spaced. The nuance is when it comes to longer holidays such as Christmas and New Year. Split the usual weekly cover into smaller chunks to ensure someone isn’t on high alert the whole time.
When changing schedules, give people fair notice, especially if your jurisdiction doesn’t legally demand it. It’s the right way to treat people and they will reward you by being more engaged and happy at work.
Pay people - Some companies don’t pay for on-call time. I disagree with this but if the salary and other benefits mean the job is still worthwhile, that’s a decision for individual engineers to make. However, a small amount of compensation goes a long way to making the rotation feel less burdensome. You may even get volunteers when someone is sick and you need emergency cover.