I deleted the guest login and built a status page
I kept adding guest permission branches to show people my dashboard, then saw the only thing worth showing was "is it alive" and built a separate public status page.
#Monitoring #Health check #React #Express #Operations
My one server has ten sites running on it. The portfolio, the blog, the couples app, the travel planner, the coin tracker, a game and so on. They all run under PM2, and Apache splits them out to ports by subdomain. To manage all this I built a monitoring server in June. I put in CPU, memory and disk graphs, a process list, a log viewer for /var/log , a web terminal, and HTTP health checks as well. The problem was that all of it sat behind a login. In this post I'll go through the things I decided while removing the guest login and replacing it with a public status page. The guest login kept getting messier To let people look without logging in, I first made a guest login. You press one button and come in as read-only. But read-only didn't mean I could show everything. The log viewer shows server paths, the visit stats show real URLs, and the terminal goes without saying. So every screen got a branch that hides things if you're a guest. Server requireNonGuest middleware masking branches in sysinfo, analytics, terminal, services Frontend isGuest, GuestBlock, GuestBlur For example, the top pages in the visit stats were blocked in two layers: blurred on the frontend, and also emptied out as topPages: [] on the backend before being sent down. Block only one side and it's all visible the moment someone opens the response in dev tools. With things working this way, every time I added a feature I first had to work out "is it okay for a guest to see this". Looking back, the only thing worth showing a guest came down to one thing, whether the services are alive . So instead of making the permission branches more elaborate, I decided to build a completely separate screen that can be viewed without logging in. A status page, in the style of status.claude.com . The raw material had already piled up There was nothing new to collect. The health check scheduler had been putting results into the health_history table since June 6, and over three months more than 250,000 rows had piled up. When I counted on September 15, as I was revising this post, it was 271,360 rows. The scheduler wakes up every 60 seconds and picks only the endpoints that are due for a check. The interval differs per endpoint: 3 minutes for the couples app and the travel API, and 5 minutes for the rest. SELECT * FROM health_endpoints WHERE last_checked IS NULL OR last_checked = DATE_SUB(NOW(), INTERVAL interval_minutes MINUTE) For uptime, all I have to do is count this per service and per date. SELECT endpoint_id, DATE_FORMAT(checked_at, '%Y-%m-%d') AS day, COUNT(*) AS total, SUM(is_healthy = 1) AS ok FROM health_history WHERE checked_at = DATE_SUB(CURDATE(), INTERVAL 89 DAY) GROUP BY endpoint_id, day What color should a day be? To draw a 90 day bar, each day has to be summarized as one cell. At first I set green at 99.9% or higher, but when I did the math, that didn't work. At a 5 minute interval there are 288 checks a day. With that, even a single failure gives 99.65% . On a 99.9% cutoff, one timeout turns the whole day yellow. And most of the real failure records are a single 10 second timeout that came right back on the next check, so the bar would end up covered in yellow dots. function dayStatus(uptime, total) { if (!total) return 'nodata'; if (uptime = 99.5) return 'operational'; if (uptime = 95) return 'degraded'; if (uptime 0) return 'partial'; return 'down'; } So I lowered it to 99.5%. For a service on a 5 minute interval, one slip in a day is green, and two or more is yellow. In fact, on August 14 the blog failed 2 out of 285 checks, so it's marked yellow at 99.30%. Honestly, I don't know if this cutoff is exactly right. The couples app is on a 3 minute interval, so it gets 480 checks a day, and the same two failures come out at 99.58%, which is green. I'm cutting services with different check intervals at the same percentage, so it's not textbook-clean. For now there are only ten services, so I'm leaving it. Grouping failure records into incidents With only the bars, you start wondering "it's yellow, so what happened that day". So I made the incidents for that day show up when you hover over a cell (or tap it on mobile). Listing the failure rows as they are would show the same incident as several lines. So I grouped them so that failures with the same reason that continue within 20 minutes count as one . The check interval is 5 minutes, so up to three or four consecutive failures are treated as the same incident. const INCIDENT_GAP_SEC = 20 * 60; for (const f of failures) { const reason = failureReason(f.status_code, f.error_message); const last = incidents[incidents.length - 1]; if (last last.reason === reason f.ts - last.endTs = INCIDENT_GAP_SEC) { last.end = f.time; last.endTs = f.ts; last.count += 1; continue; } incidents.push({ start: f.time, end: f.time, endTs: f.ts, count: 1, reason }); } There were only 31 failures in 90 days, so I didn't split this into a separate API and put it in the /api/status response along with everything else. There's no need to send a request every time the mouse goes over a cell. Is it okay to put raw error text on a public page? I stopped here for a moment. The errors stored in health_history look like this. getaddrinfo EAI_AGAIN blog.jaeyonging.com timeout of 10000ms exceeded connect ECONNREFUSED 127.0.0.1:59999 A health check target could be an internal address, and the raw error text includes the host or port as it is. That's useful to me when I'm logged in, but it's not information to put on a public page. So the public response only gets a categorized reason. function failureReason(statusCode, errorMessage) { const err = (errorMessage || '').toLowerCase(); if (statusCode == null) { if (err.includes('timeout') || err.includes('etimedout')) return 'Response timed out'; if (err.includes('eai_again') || err.includes('enotfound')) return 'DNS lookup failed'; if (err.includes('econnrefused')) return 'Connection refused'; if (err.includes('econnreset') || err.includes('socket hang up')) return 'Connection dropped'; if (err.includes('cert') || err.includes('ssl')) return 'Certificate error'; return 'No response from server'; } if (statusCode = 500) return `Server error (${statusCode})`; if (statusCode = 400) return `Request error (${statusCode})`; return `Unexpected response (${statusCode})`; } The raw text is only sent down when you're logged in. I build two versions from the same data, and give out the one with the raw text if the token is valid, and the cleaned up one if not. router.get('/', async (req, res) = { if (!cache.full || Date.now() - cache.at = CACHE_TTL) { const full = await buildStatus(); cache = { at: Date.now(), full, publicData: sanitize(full) }; } res.json(isAuthenticated(req) ? cache.full : cache.publicData); }); Under each service name I show the domain too. With only "Travel API" written there, even I got confused about which project it was. This also gets filtered to null if it's localhost or a private IP range, so internal addresses don't go out. An API with no login gets a cache first With a public endpoint you don't know who will call it or how often. Running a GROUP BY over more than 250,000 rows of history every time is a burden, so I added a 30 second in-memory cache. The frontend polls every 30 seconds too, so it lines up with the refresh cycle. I measured it just now, and the first request with an empty cache took 0.75 seconds, and after that it was 0.035 seconds. $ curl -so /dev/null -w '%{http_code} %{time_total}s %{size_download}B\n' https://monitor.jaeyonging.com/api/status 200 0.745856s 87067B 200 0.037002s 87067B 200 0.034586s 87067B There was no need for something like Redis. It's one server and one process, so a single module variable is enough. The screen At the very top is the overall status banner, and below it the 90 day bar, uptime and current response time for each service. I cut down the number of days shown depending on the width: 30 days under 480px, 45 days under 768px, and 60 days under 1100px. If you fit 90 cells into a phone's width, one cell is a little over 2px and you can't press it with a finger. Touch devices have no hover, so tapping a cell opens the tooltip and tapping again closes it. And then I deleted the guest login Once the status page existed, I deleted all the guest-related code. POST /api/auth/guest , requireNonGuest , the four masking branches on the server side, and on the frontend isGuest , GuestBlock and GuestBlur too. I moved the admin screen under /dashboard , and / became the status page. With that gone, I don't have to think about permissions when adding a feature. I'm the only one who logs in, and everyone else only sees the status page. What I'm left with Compared with hiding bits of one screen to fit several permission levels, building a separate screen that holds only what should be shown was far simpler. Guest mode started from "show them the dashboard", but what was actually needed wasn't the dashboard. And while building this page, I laid out the accumulated records by date and looked at them for the first time. There were 3 months of records, yet nothing was set up to tell me when a service actually died. I wrote that story up in the next post .