My health check recorded for 3 months and never told anyone
31 failures piled up with 0 alerts. Alerting on every failure meant 58 emails, 2 in a row meant 4. I counted from the records first, then added alerts on state change.
#Monitoring #Health check #Alerts #Node.js #Incident response
the last post I built a status page out of my health check records. If you look at the bars, there are a few yellow cells over the 90 days. But at the moments those failures happened, not a single signal reached me. The health check had been running since June, but it only recorded the results and never told anyone. In this post I'll go through how I added down and recovery alerts to it, and how I first used the accumulated records to work out what the alerting rule should be. It's not that there was no alerting The monitoring server did have an alert feature. If you put a rule in the alert_rules table, checkAlerts() runs every 60 seconds and sends email. But it only looked at three things. if (rule.metric === 'cpu') currentValue = cpu.currentLoad; else if (rule.metric === 'memory') currentValue = memPercent; else if (rule.metric === 'disk') currentValue = maxDiskUse; else continue; On the health check side, runHealthChecks() put the results into two tables and stopped there. When a service dies the light on the dashboard does turn red, but to see that I'd have to be sitting with the dashboard open. alert_history had just three rows in it from before September when I opened it. 2026-06-06 19:27 [Cpu usage] cpu 10% (current: 14.6%) 2026-06-06 19:46 [cpu usage ] cpu 1% (current: 3.7%) 2026-08-18 14:39 [cpu] cpu 40% (current: 62.4%) The first two rows are rules I added on the day I built the alert feature, to see whether the email arrived. In 3 months, not one alert had gone out for a service outage. What happens if I alert on every failure? The simplest approach is to send an email every time a check fails. Before building it, I checked whether that would be okay by running it against the accumulated records. I cut the records into runs of consecutive failures per endpoint, and this is what came out. Total checks about 270,000 Failures 31 Failure runs 29 1 failure, recovered right away 27 2 failures in a row 2 (8/14 Couples app API, 8/18 Blog) Out of 29 runs, 27 failed once and came back on the very next check. Most of them are timeout of 10000ms exceeded . Sending a down email for every failure and a recovery email for every recovery comes to 58 emails, and for 54 of them everything would already be fine by the time I saw the email and went to look. After getting alerts like that a few times you stop reading them, and the email for a real outage gets buried along with them. A consecutive failure threshold So I added fail_threshold . The default is 2 consecutive failures. Run against the same records, that gives 2 downs, or 4 emails including the recovery emails. The cost is that detection gets slower. With a 5 minute interval, it can take up to 10 minutes to reach the second failure. On a personal server, I figured getting numb to false alarm emails is worse than finding out 10 minutes late. The threshold can be changed per endpoint, so for an urgent service I can lower it to 1. Failures that fall short of the threshold still need to be visible on screen though, so I split the status in the health check table into three. Healthy, Unstable 1/2 (yellow), and Down for N min (red). This is what the health check table on the admin screen looks like now. All ten are healthy so there's no yellow or red to see, but in the alert column on the right I can turn email on and off per endpoint, and each one has its own threshold. So do I send one on every check while it's down? If I sent an email on every check once the threshold was passed, an outage lasting 1 hour would bring twelve emails at 5 minute intervals. So each endpoint holds an up or down state, and I send only at the moment the state changes . It's the way services like UptimeRobot do it. async function evaluateAlert(ep, healthy, detail) { const threshold = ep.fail_threshold ?? 2; const prevState = ep.alert_state || 'up'; const failures = healthy ? 0 : (ep.consecutive_failures ?? 0) + 1; let downSince = ep.down_since ? new Date(ep.down_since) : null; let state = prevState; if (!healthy) { if (!downSince) downSince = new Date(); if (prevState === 'up' failures = threshold) state = 'down'; } else { state = 'up'; } await pool.execute( 'UPDATE health_endpoints SET consecutive_failures=?, alert_state=?, down_since=? WHERE id=?', [failures, state, healthy ? null : downSince, ep.id] ); if (state === prevState) return; // if nothing changed, do nothing if (!ep.alerts_enabled) return; // update the state but skip the mail if (state === 'down') await sendDownMail(to, ep, detail); else await sendRecoveryMail(to, ep, Date.now() - downSince.getTime()); } The key is the single line if (state === prevState) return . One email when the down is confirmed, one email when it recovers, and nothing in between. alerts_enabled can be off and the state still keeps updating, which I also made a point of. If an endpoint goes down with alerts off and I then turn them on, the state is already down , so a down email doesn't suddenly pop out the moment I turn them on. When does downtime start counting? When I went to put the downtime in the recovery email, I had to decide one thing. If down_since is recorded at the point the down is confirmed, then with a threshold of 2 the 5 minutes from the first failure to the confirmation go missing. The outage already started at the first failure, but the email would show it 5 minutes short. So down_since is recorded at the first failure and only cleared when the endpoint goes back to healthy. That's the if (!downSince) downSince = new Date() part in the code above. If it recovers right after a single failure, it's quietly erased with no email. Adding columns to an existing table health_endpoints needed seven columns. alerts_enabled , fail_threshold , consecutive_failures , alert_state , down_since , notify_email , last_error . In this project, initDb() creates the tables with CREATE TABLE IF NOT EXISTS when the server starts. But that does nothing if the table already exists, so the columns don't get added. Bringing in a migration tool felt like overkill, so I wrote one helper that ignores only the duplicate column error. async function ensureColumn(conn, table, column, definition) { try { await conn.query(`ALTER TABLE ${table} ADD COLUMN ${column} ${definition}`); console.log(`[DB] ${table}.${column} column added`); } catch (err) { if (err.code !== 'ER_DUP_FIELDNAME') throw err; } } await ensureColumn(conn, 'health_endpoints', 'fail_threshold', 'INT NOT NULL DEFAULT 2'); await ensureColumn(conn, 'health_endpoints', 'alert_state', "ENUM('up','down') NOT NULL DEFAULT 'up'"); It's only half a solution since it can't rename a column or change its type, but when all I'm doing is appending like now, this is enough. The recipient is picked in this order: the endpoint's own notify_email , if there isn't one the alert email in the global settings, and if that's missing too the first of the emails allowed to log in. Even with nothing configured at all, it still reaches me. Testing You can't trust an alert until you've actually received one. I registered a test endpoint pointing at a port with nothing running on it, with a threshold of 1. [Health] [Test] alert check down - request failed: connect ECONNREFUSED 127.0.0.1:59999 (1 consecutive failure) The down email arrived, and when I changed the URL to a working address, the recovery email arrived 1 minute later. [Health] [Test] alert check recovered - downtime about 1 min After confirming that, I deleted the test endpoint. And while doing this work I also found that two services, Ieum and Dapjeongneo, had been missing from the health check targets altogether. Only after registering them did the 10 PM2 apps, 10 Apache vhosts and 10 health checks match up 1:1. The first real case On the night of September 13, four days after adding the alerts, a failure was recorded on the travel API. 23:07:04 Travel API HTTP 503 (expected 200) 23:10:04 Travel API Healthy It failed once and came back 3 minutes later. That falls short of the threshold of 2, so no health check email went out. Just as designed. But at the same time, a different alert went out. 23:08:07 travel-server crash loop - 3 restarts in the last 10 min (94 total) 23:12:08 travel-server stabilized - restarts stopped (lasted 4 min) 23:39:07 travel-server crash loop - 3 restarts in the last 10 min (99 total) 23:43:07 travel-server stabilized - restarts stopped (lasted 4 min) This was sent by the PM2 crash loop monitor I built on the same day. travel-server was dying three times every 10 minutes, and the health check happened to land on just one of the moments it was dead and saw the 503. During the second crash loop at 23:39, the health check didn't fail even once. PM2 brought the process back right away, so it was fine by the time the check arrived. I wrote up this blind spot separately in the PM2 crash loop post . With the two side by side, the roles split. The health check looks at "is it responding right now", and the PM2 monitor looks at "is it quietly dying over and over". That day the PM2 monitor caught what the health check let slip by. On the other hand, if the problem is DNS or Apache, the process is fine, so only the health check catches it. What I'm left with When you build monitoring, it's easy to stop at piling up graphs and records. I left it that way for 3 months myself. But records only mean something if someone looks at them, and there's nobody to look at them in the middle of the night. That said, if you alert on every failure, you end up not looking at those emails either. Before adding alerts, I'd recommend first counting from the accumulated records "how many emails would have come under this rule" . In my case it was the difference between 58 and 4.