Analyzing one link took 17 seconds, so I dropped the request-response approach
One link analysis API took 17.5 seconds, so I turned the request into a job. Progress stages, joining the same link, a rate limit trap, and parallel work to cut the wait.
#Backend #Node #Express #API #Performance #Troubleshooting
In the last post I wrote about pulling travel spots out of YouTube videos by going as far as reading the captions. The feature came out the way I wanted, but once it was hooked up a different problem showed up. It takes too long . Thinking about it, that was only natural. It scrapes the metadata from YouTube, downloads captions with yt-dlp, calls the AI, matches the names that come out against public data, and runs a Kakao Local search for the ones it couldn't find. That's four external services. I had set the timeout for caption fetching alone to 45 seconds, which means that in the worst case a single request stays alive that long. So I'll go through, in order, how I split that request up and how I cut down the waiting time itself. The limits of waiting for everything at once The first API I built was ordinary. It takes a link by POST and returns the result when everything is done. router.post('/analyze', asyncHandler(async (req, res) = { const url = String(req.body?.url ?? '').trim(); // ... fetch -> extract -> verify -> geocode ... res.json({ content, places }); // takes ten-odd seconds to get here })); When I actually measured it, one YouTube video took 17.5 seconds. Taking that long is a problem in itself, but the bigger problem is that nothing else can be done during that time . The user has to wait looking at a spinning icon on a white screen, and if they move to another screen the request is cut off. There's an Apache reverse proxy in front, so I have to worry about the proxy timeout too, and on mobile the screen turning off for a moment is a risk of its own. Above all, from the user's side there's no way to know how far along it is . A loading icon spins the same whether it takes 3 seconds or 30. I turned the request into a job So I went with the usual approach. The start request responds right away with just a job id, and progress is asked about separately using that id. const JOB_TTL_MS = 10 * 60 * 1000; const jobs = new Map(); // jobId -> { stage, result, error, at } const runningByContent = new Map(); // contentId -> jobId (the same link joins one job) // POST /api/sns/analyze/start -> { result } (already analyzed) or { jobId, stage } // GET /api/sns/analyze/jobs/:id -> { stage } or { result } The important part here is the second Map. If several people paste the same video link at the same time, there's no reason to run the same analysis several times. So if a job is already running for that content id, it joins that job instead of starting a new one . That saves quite a lot of YouTube and AI calls. Job results stay in the in-memory Map for only 10 minutes. Once the analysis finishes it's saved to the DB anyway, and if the same link comes in again it's pulled straight from the DB. There was no reason to hold on to it for long. I matched the progress stages to the real progress While splitting it into a job I also added a progress stage display. Not the fake progress bar you often see, but one that tells you when it has actually entered that stage . /** * The actual analysis: fetch -> extract -> save. Reports the progress stage through onStage * fetch(collect content, captions) -> extract(find places) -> verify(check public data) * -> geocode(fill in coordinates) -> done */ async function runAnalysis(url, platform, id, onStage = () = {}) { onStage('fetch'); const meta = platform === 'youtube' ? await fetchYoutubeMeta(url) : await fetchWebMeta(url); const { placeIds, signals } = await extractPlaces(meta, { onStage }); ... onStage('done'); } On screen it looks like this. The four stages light up one by one as the real work progresses. The button at the bottom matters more than it looks. If you press "I'll keep analyzing while you look around", the analysis keeps running on the server even after you move to another screen, and a notification tells you when it's done. With the request-response approach this wouldn't have been possible in the first place. I also wrote "Long videos can take over a minute" up front. If I can't hide that it really takes a long time, I figured it's better to say so in advance . The rate limit almost killed the polling After switching to asking for status, I got caught once by something silly. I had put a request limit in place so that one user couldn't use up the external API quota, and it ended up catching the status checks too. I had limited analysis to 20 requests per minute. But checking status every 2 seconds is 30 requests per minute. The polling I built was hitting the limit I set . So I put the limit only on the start request and took the status check out of it. const limiter = (max, windowMs, message) = rateLimit({ windowMs, max, standardHeaders: true, legacyHeaders: false, message: { error: message }, }); app.use('/api/', limiter(600, 60 * 1000, 'Too many requests. Please try again in a moment.')); app.use('/api/sns/analyze', limiter(20, 60 * 1000, 'Too many analysis requests. Please try again in 1 minute.')); app.use('/api/itinerary/generate', limiter(15, 60 * 1000, 'Too many itinerary generation requests.')); app.use('/api/auth/', limiter(30, 60 * 1000, 'Too many login attempts.')); There was one more trap that came along with this. This server sits behind an Apache reverse proxy, so left alone, every request looks like it comes from the single proxy IP . If you set a limit in that state, all visitors are treated as one person and everyone gets blocked after only a few people use it. app.set('trust proxy', 1); // rate limit by the real client IP behind Apache One line fixes it, but if you don't know about it, the question isn't "why is it only me" but "why is everyone blocked together", so finding the cause gets pretty confusing. Before adding a rate limit, it's better to check once what shows up in req.ip. I cut the waiting time itself too Splitting it into a job fixed the experience of waiting, not the speed. 17 seconds is still 17 seconds. So I worked on the inside as well. When I took the analysis apart, there were things that had no reason to run in sequence. The AI call takes a few seconds and the server just sits idle during that time, and picking candidates by rules has nothing to do with the AI result. So I changed it to run the two at the same time and merge them afterward . // The AI call (a few seconds) and the rule-based candidate calculation are independent, so they run at the same time const aiPromise = aiEnabled() ? aiCandidates(meta) : Promise.resolve({ cands: [], method: 'rules' }); const basePromise = (async () = { const r = await matchCandidates(baseCands, hints); return { matched: r.matched, extra: await geocodeUnmatched(r.unmatched, hints, 'rules') }; })(); const [ai, base] = await Promise.all([aiPromise, basePromise]); The Kakao Local search was the same. I was searching one candidate at a time in order, but the candidates have nothing to do with each other, so they can be sent together. // Kakao searches per candidate are independent, so handle 5 at a time (16 sequential calls about 4s -> around 1s) const results = []; for (let i = 0; i targets.length; i += 5) { results.push(...(await Promise.all(targets.slice(i, i + 5).map(lookup)))); } The reason I send them 5 at a time instead of all at once is the other API's side of things. Throw dozens at it in one go and you're likely to get rejected. Timing each stage, this is what it looks like now. ===== AI on : 4,647ms stages: extract+0ms -> verify+3,594ms -> geocode+3,640ms ===== rules only : 1,522ms stages: extract+1ms -> verify+1,449ms With AI the whole thing is about 4.6 seconds, and 3.6 seconds of that is the part that matches candidates against public data and fills in coordinates. The 17.5 seconds I mentioned earlier is this plus caption fetching. In the end most of the remaining time is spent waiting on external services, so to cut it further I'd have to work on that side. As a bonus, I cached the first screen too Separately from the analysis, I also looked at the app's slow first screen. There's an API that sends down 3,000 places plus region and festival info in one go on first load, and this data only changes once a day, at midnight, through cron . There was no reason to build it fresh from the DB every time. const TTL_MS = Number(process.env.BOOTSTRAP_CACHE_MINUTES || 10) * 60 * 1000; let cache = { body: '', etag: '', at: 0 }; router.get('/', asyncHandler(async (req, res) = { const fresh = Date.now() - cache.at TTL_MS ? cache : await buildBootstrap(); res.setHeader('ETag', fresh.etag); res.setHeader('Cache-Control', 'public, max-age=0, must-revalidate'); ... })); The server keeps the finished JSON string in memory for 10 minutes and gives the browser an ETag. If the content hasn't changed it ends with a 304 and the body isn't sent at all. For a list-type API whose data doesn't change often, even this much makes a noticeable difference. What came up after switching to polling Once I changed to a structure that asks for status, the number of requests jumped. A single analysis comes with dozens of status checks. And the next day's commits include this. 08-27 fix(auth): login expiry handling and a safeguard for festival loading If the token expires while polling is running, 401s keep piling up, but on screen it just looked like "Analyzing" had frozen. That's when I learned that the worst thing is a failure that doesn't look like a failure . The polling code is now organized like this. It asks for status every 0.9 seconds, but once 4 minutes have passed since the start it stops with "The analysis is taking too long", and if the server returns 404 it treats the job as expired and tells the user to try again. Any other temporary network error is just ignored, and it asks again on the next poll. const POLL_MS = 900 const MAX_WAIT_MS = 4 * 60 * 1000 // after this long, give up and tell the user if (Date.now() - job.startedAt MAX_WAIT_MS) { setError(job.url, { status: 408, message: 'The analysis is taking too long. Please try again in a moment.' }) continue } try { const s = await api.getSnsAnalysisJob(job.jobId) if (s.result) complete(job.url, s.result) else if (s.error) setError(job.url, s.error) else if (s.stage !== job.stage) setStage(job.url, s.stage) } catch (e) { if (toApiError(e).status === 404) setError(job.url, { status: 404, message: 'The analysis job has expired. Please try again.' }) // other temporary errors are retried on the next poll } The 4 minute figure is a margin that still leaves room after adding the AI call and the searches to the 45 second caption timeout, not a value that came from measurement. I also thought about counting how many times in a row it failed, but if a brief drop on the subway counts as a failure, the screen might give up first on an analysis that's running fine, so I left it at a single time limit. Whether this is the right criterion, I'll only know after using it more.