John McCormack Engineering Manager

Engineering management, SRE and database administration

  • Posts
    • Hire me
      • Take a look at my Sessionize speaker’s profile
      • Let me solve your SQL Server problems – No longer available
    • Guides
    • cost-optimization
    • SQL Server
    • Training
    • Personal
    • Azure
    • T-SQL
    • AWS RDS
    • AWS SQL Server
  • Resume
  • Free Training
    • SQL Server on Amazon RDS (Free Course)
    • Free practice questions to help you pass DP-900
  • Cost Optimization
    • Azure IaaS SQL Backups – Stop burning money
    • Your Azure SQL Database and Managed Instance is too big
    • Turn the cloud off at bedtime to save 70%
    • Your Azure SQL Virtual Machine might be too big
    • Save money with Azure SQL DB serverless
    • Save up to 73% with reserved instances
    • Delete unused instances to save money in Azure
  • Personal
    • About

Alert fatigue

6th September 2026 By John McCormack Leave a Comment

Red button with the word HELP in bold capital letters
Image by OpenClipart-Vectors from Pixabay

I mean this sincerely when I say, you have got to get your alerting under control. Excessive alerting is a sign either of your service being out of control, or your alerts being misconfigured. Alert fatigue reduces reliability and destroys morale of on-call engineers. Here is why.

Why over alerting is an enemy of SRE

It reduces reliability

  • Alerts that you have seen before, possibly multiple times per week, lack importance. If you get twice daily alerts about a replication lag, it’s tempting to simply acknowledge it and expect it to self resolve. Especially if the alert comes in out of hours.
    • If it’s the case that certain types of events always self heal, they are rarely a good candidate for alerting. These should be downgraded to email notifications or silently logged in order that history can be searched and trended. They should not page on-call engineers.
  • If an on-call engineer receives more than two alerts per on-call shift, they have a limited amount of time to investigate the root cause and put in proactive fixes. Engineers with a busy pager are always firefighting, and end up reworking the same incidents time after time.

Over alerting destroys morale

  • Out of hours alerts interrupt sleep patterns and this makes engineers unwell.
  • Over alerting causes the engineer to be in a state of high alert, especially if they are responsible for a mission critical service.
  • Regular out of hours alerting can make life difficult for the engineer’s family. Especially their partner who gets woken up as well due to the blaring page notification.
  • All of these lead to a fear of the next out of hours rotation.

Why does alert creep happen

In my experience, I put this down to two factors. Good intentions and incident post mortems. Starting with good intentions, I’ve done it myself. I’ve thought, you know what will help, a really fine grained alert about X condition. There’s a serious flaw with this. You are motivated for the information at the point in time you create the alert, but when it starts firing multiple times per day or week, you know it is safe to ignore and it stops receiving the action it needs.

The second and most prolific factor in my opinion is incident post mortems. In SRE, post mortems should not be about blame. That being said, when you have to present at an incident review meeting, often to people who don’t work closely with you, it’s natural to want to be seen as competent and not the cause of maxing out the error budget. You also really don’t want repeat incidents so the tendency is to identify reporting gaps. This insures you for the future and makes you look competent, the trouble is that the burden on the on-call engineer grows. More benign alerts, especially overnight, are not the friend of on-call engineers.

Simple steps to fix alert fatigue

Trim and downgrade

Do your alerts line up with your Service Level Objectives (SLO)? There is no need to alert on issues that don’t interrupt your service or uptime. They should be focussed on issues that are eating into your error budget, and thresholds need regular evaluation.

It’s important to trim alerts regularly. During a recent “fixit day” (where meetings are banned and engineers should deal with nagging problems), I relegated 4 alerts from pages to Slack alerts. There is no expectation to check Slack when on call, meaning only the important issues come through the paging service. Don’t let alerts die in Slack or email. Even these need regular evaluation and lifecycle management.

As well as this, some alerts are just no longer useful and should be cut altogether. Don’t be hasty about this but evaluate, can anything actually be done to resolve the issue or is this just information? If it’s information, it’s better for everyone to cut it.

Fair rotation and fair notice

This advice is for engineering managers. Most rotations are fair. If you have x people, they are on every x weeks. Possibly you run a secondary or a working hours schedule as well. Whatever you do, rotations for employees should be equally spaced. The nuance is when it comes to longer holidays such as Christmas and New Year. Split the usual weekly cover into smaller chunks to ensure someone isn’t on high alert the whole time.

When changing schedules, give people fair notice, especially if your jurisdiction doesn’t legally demand it. It’s the right way to treat people and they will reward you by being more engaged and happy at work.

Pay people - Some companies don’t pay for on-call time. I disagree with this but if the salary and other benefits mean the job is still worthwhile, that’s a decision for individual engineers to make. However, a small amount of compensation goes a long way to making the rotation feel less burdensome. You may even get volunteers when someone is sick and you need emergency cover.

Filed Under: front-page, SRE Tagged With: alerting, burnout, firefighting, on-call, sre

Why SRE teams must eliminate toil

24th August 2026 By John McCormack Leave a Comment

A red-handled axe resting on a wooden chopping block
Image by markusspiske from Pixabay

Introduction

Near the start of my tech career, I remember handling a ticket and thinking, why am I doing this again? It was the sort of request that came in from developers to DBAs, multiple times per week and easily took 30 to 60 minutes each time. That started a process where I suggested and implemented a self service system for devs, that allowed them to bypass DBAs for specific repeatable tasks where sometimes all that changed were some parameters. It eliminated dev waiting time on DBA availability, prevented us from doing monotonous repeatable tasks, and made the service consistent each time. I didn’t know it in 2014 but I was working on toil elimination.

Definition of toil

Toil as an SRE concept was introduced to the world by Google. It was defined as “the repetitive, predictable, constant stream of tasks related to maintaining a service”.

Toil adds no enduring value. I remember thinking, once I’ve done this a few times already, I’m not learning anything each time I do it. The agent jobs, metadata tables and SSIS packages had already been developed, I was simply pushing buttons, running scripts and waiting for one part to finish before going on to the next bit.

I’ll go on to talk about the impacts of toil on employee morale. This was the thing that struck me most. I was late starting in tech. I wanted to play catch up to colleagues of a similar age, and this toil was preventing me from learning the skills to work on interesting projects. So I made the automation of these tasks my project.

Why toil elimination matters

Reliability

At the heart of SRE is reliability. Reliable services endeavouring for high nines levels of uptime need to be self diagnosing and self healing as much as possible. Whilst you should have a playbook for when things go wrong, and standard operating procedures for manual processes, if you have to touch it to fix it then it’s likely toil. Self healing processes also reduce mean time to recovery resulting in higher availability, more satisfied customers and happier engineers (who weren’t called into a 3am Sev 1 incident).

Employee morale

Some people might get some comfort out of working through established SOPs, but everyone has a tipping point. Add to that, no one likes being paged out of hours, or even during office hours so it really is to the benefit of everyone that toil is eliminated wherever possible.

Furthermore, the work gets boring and most people generally like to learn new things and deliver value at work. Some don’t but they don’t tend to be high performing engineers, and possibly aren’t the best fit for a modern SRE organisation.

Scalability

Google talks about sublinear scaling. In essence, if you double your output without doubling support requirements, you have achieved this.

A traditional operations role where all tasks are done manually such as patching has to be automated in order to scale effectively. The first step in this is writing scripts to do the work. The logical evolution is for the scripts to belong to a software solution that handles scheduling, makes the changes and ensures smooth failovers.

What isn’t toil

Not all manual tasks are toil. If there is a derived benefit either to a person’s skill set or the service they are supporting, manual work can and does add value. It also stops skill atrophy which is a real effect of over-automation.

Toil isn’t just about saving FTE either. Of course it’s an added benefit to organisations that as they scale, their costs don’t need to increase linearly however I hate to see SRE as a means to cutting jobs. It’s just a better way of doing things and it leads to opportunities to increase personal skill level. That makes the engineers more employable and that’s never a bad thing.

How does toil elimination fit into SRE

I mentioned earlier that reliability is at the heart of SRE. The same is true of engineering. SREs should primarily be able to write code. SREs are like software engineers but with competing priorities to their colleagues that ship products and features. By writing effective solutions in code, they are quite literally engineering reliability into the service.

Toil is often a result of technical debt. A consequence of saying I know how to fix this so I’ll write an SOP, or I’ll put a debugging script on confluence. That isn’t SRE, it’s operations and it doesn’t scale.

How does an organisation drive this

There needs to be buy in from management. I was lucky early in my career that my manager encouraged my self service project and the benefits were clear. However not everyone will have that support. If you don’t have it, it’s an opportunity to learn how to influence and here’s how I would do it.

Measuring toil and your elimination efforts

Measurement is essential, with or without buy in but where you don’t have it, it adds validity to your efforts.

First of all, work out how much time is spent on specific tasks. Maybe start with one task as a proof of concept. Look for tasks that are needed to keep the lights on (ktlo) but just don’t contribute to personal or business growth. Even if your team doesn’t track time spent on tickets, search historic tickets and put an estimate on it. Remember to include time for context switching.

Secondly, measure the impact of your engineered solution once it is operating. How many hours of toil did you eliminate? How did it affect service uptime? Possibly harder to quantify but worth thinking about is, how did it improve team morale?

When you have the numbers, the benefit cannot be denied. Good luck with your toil elimination efforts.

Filed Under: front-page, SRE Tagged With: automation, sre, toil

John McCormack · Copyright © 2026