Most on-call rotations start with good intentions and end with someone updating their LinkedIn. The schedule ships as a shared Google Sheet, the team agrees to weekly handoffs, and for the first month things feel manageable. Then the 3 AM pages start compounding. Alert noise grows. The engineer on rotation spends their weekdays recovering from their nights, and the backlog doesn't care.
On-call engineer burnout is not a character flaw. It's a structural problem, and structural problems respond to structural fixes.
What makes on-call rotations unsustainable
The most common failure isn't the volume of incidents. It's the volume of noise that doesn't lead to incidents. A , the on-call engineer becomes a human priority sorter, doing work that a well-configured system should handle automatically.
Recovery time doesn't exist in the schedule. The engineer finishes a rough on-call week and is expected to hit sprint velocity on Monday morning. There is no decompression built into the rotation, no acknowledgment that interrupted sleep for seven nights has a cognitive cost that doesn't reset with a weekend.
Reducing alert noise before it reaches a person
The fastest way to improve on-call quality is to reduce the number of alerts that wake someone up. Not by suppressing them, but by routing them correctly.
Start with an audit. Pull every alert that fired in the past 30 days and categorize each one: did it require human action within an hour? If the answer is no, it should not page. It might still warrant a Slack message or a ticket for morning review, but it should not vibrate someone's phone at 3 AM.
Severity-based routing is the mechanism that makes this practical. A warning-level check failure goes to a dashboard. A critical failure goes to the on-call channel. A P0 pages immediately. Most guide.
Follow-the-sun where possible. If your team spans time zones, route pages to whoever is awake. Overnight pages are the single largest contributor to burnout, and even partial coverage (handling pages until midnight local time instead of through the full night) makes a meaningful difference. An engineer in Berlin shouldn't get paged at 4 AM when a colleague in San Francisco is eating lunch.
Define an escalation path. The on-call engineer should never feel alone. If they can't resolve an issue within a defined window, escalation to a secondary responder should be automatic. Knowing that backup exists, and that escalation is expected rather than a sign of failure, changes how people experience on-call entirely.
Compensation and recognition
On-call is real work. It costs real things: sleep, personal time, cognitive capacity the next day. If your organization doesn't compensate for it, the people who can leave will leave, and the rotation pool shrinks, which accelerates burnout for everyone remaining.
Compensation takes different shapes depending on the company. Some teams pay a flat weekly stipend for carrying the pager. Others add per-incident bonuses. A few give compensatory time off after high-severity weeks. The specific model matters less than the principle: on-call should never be invisible, unpaid labor.
Beyond money, recognition matters. Teams that review on-call load in retrospectives, that track tying alerts to SLO burn rates rather than raw threshold violations. A single failed health check is noise. A pattern of failures that puts the error budget at risk is signal. This distinction, applied consistently, can cut actionable pages by half or more.
Automated can help reduce the noise before it reaches your on-call channel. But the structural fixes come first. No tool fixes a four-person rotation with no recovery time.
Originally published on DevHelm.
SOCIAL SHARE CARD GENERATOR