I ended up reading the captions to pull travel spots out of YouTube videos
To pull travel spots from a YouTube link, I read auto captions with yt-dlp when the description was empty. Why merging AI and rules made results worse, and how I split their roles.
#Backend #Node #AI #OpenAI #YouTube #Troubleshooting
These days I'm building a travel planner for Gangwon Province. The feature fits in one line. You paste a YouTube or blog link, and it finds the places that appear in it and builds an itinerary for you. It sounds simple, but when I actually built it, the very first step, working out the place names from a link, took the longest. I want to write down two things I ran into along the way. One is how I handled videos with an empty description, and the other is how adding AI actually made the results worse. All this screen does is take one link. The problems were behind it. Too many videos had an empty description At first I only read the title and the description. Travel YouTubers often lay out chapters in the description like "0:00 Opening / 1:20 Baekchon Makguksu", so scraping just that worked pretty well. Vlogs were a different story, though. It's common for the description to have nothing but a greeting and a sponsorship notice, without a single line of place names. The shop sign is clearly on screen and the name is said out loud, but nothing is left as text. That left only one source. The names people say out loud do stay in the captions, so it came down to reading the captions . The YouTube API can't fetch auto-generated captions I assumed the official YouTube API would obviously have a way to get captions. It does, but only the owner of the video can download the caption text itself. For someone else's video you can see the list of captions and it won't give you the contents. So in the end I put the yt-dlp executable on the server and ran it directly. By the textbook it's not a pretty choice. A Node process running an external binary means one more dependency, and I also have to decide separately what happens when that binary is missing or broken. Still, there was no other way, so instead I made it a condition that the server keeps running fine even when the captions can't be fetched . The code looks like this. const YTDLP = process.env.YTDLP_BIN || path.join(os.homedir(), '.local/bin/yt-dlp'); const TRANSCRIPT_MAX = 12000; async function fetchYoutubeTranscript(videoId) { if (!fs.existsSync(YTDLP)) return ''; // if it's missing, go on without captions const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'yt-cap-')); const out = path.join(dir, 'cap'); try { await new Promise((resolve) = { execFile(YTDLP, [ '--skip-download', '--write-auto-sub', '--write-sub', '--sub-lang', 'ko,ko-KR', '--sub-format', 'json3', '--no-warnings', '--no-playlist', '-o', out, `https://www.youtube.com/watch?v=${videoId}`, ], { timeout: 45000, maxBuffer: 4 * 1024 * 1024, env: { ...process.env, PATH: `${process.env.PATH || ''}:/usr/bin:/usr/local/bin:/bin` }, }, (err, _stdout, stderr) = { // even if some languages fail, use the files that did download if (err) console.warn('[sns] some captions failed:', String(stderr || err.message).trim().split('\n').pop()); resolve(); }); }); ... } finally { fs.rmSync(dir, { recursive: true, force: true }); } } There were three things I paid attention to here. First, a process started by PM2 doesn't inherit the PATH I use in my terminal. So I filled in the PATH myself and passed it along . I lost a good while by not doing this, because it worked fine when I ran it by hand in the terminal and failed only on the server, which made the cause surprisingly tricky to find. Next is the temp folder. yt-dlp drops the captions as files, so I create a new folder with mkdtemp every time and delete the whole thing in finally. Otherwise junk piles up in /tmp. The last one matters most: even when there's an error, it resolves instead of rejecting . Why I did that comes right next. I shouldn't have treated failure as failure In production, things like this really do show up in the log. [sns] some captions failed: ERROR: Unable to download video subtitles for 'ko': HTTP Error 429: Too Many Requests [sns] some captions failed: ERROR: [youtube] zzzzzzzzzzz: Video unavailable At first I handled this as an error and aborted the whole analysis. As a result, a video whose captions couldn't be fetched produced no result at all, even though its title and description were perfectly fine. When I thought about it, captions are a nice-to-have ingredient, not a required one . So I changed caption fetching to always resolve, success or failure, and to return an empty string when it fails. Whether a 429 comes back, the video has been taken down, or yt-dlp itself is missing, the analysis carries on with just the title and description. As a bonus, caption fetching starts at the same time as metadata fetching. There's no reason for them to wait on each other anyway. const transcriptPromise = fetchYoutubeTranscript(videoId); // start alongside the metadata // ... fetch metadata ... return { transcript: await transcriptPromise, ... }; The captions I actually got were pretty messy I thought adding captions would solve everything, but the quality of auto-generated captions was worse than I expected. When I actually measured one video, the captions came to 1,067 characters, and the beginning looked like this. เฮ [music] [music] [music] [music] [music] [music] เฮ [music] [music] [music] [music] [music] [music] [music] [music] because it's that kind of pochagi it can and also at the party level they can't respond could be or it could be the candidate's difference, you know. All sorts of problems It's a vlog with long stretches of background music, so half of it is filled with [music], and in between, scraps of some unrelated language and even completely different content get mixed in. It's text produced by speech recognition, so that can't be helped, and if I trusted it as is and pulled place names from it, the wrong places would pop out. So I settled on using captions only to widen the pool of candidates, without trusting the names that come out of them as they are . That leads into the next part. Adding AI actually made the results worse I was using two methods together to find place names. One is rules. They scrape names by pattern from things like chapters, hashtags and subheadings, and scan the text for names that exist in the Korea Tourism Organization's public data catalog. The other is AI. I hand over the whole text and get back a list of places as JSON. At first I just merged the two. I figured the more candidates, the better. But when I looked at the results after merging, odd things kept getting mixed in. The cause was simple. The AI reads all of the body text and subheadings anyway. If the rules scan the body on top of that, it amounts to scraping the same text twice , and since the rules don't understand context, they put forward even names that were only mentioned in passing. And because the final result keeps only the top 20 places, that noise pushed out places that were actually visited. So I split the roles. When AI is on, the rules contribute only reliable sources . const strongOnly = (c) = c.card || c.chapter || c.title; // map card, chapter, title const baseCands = rules.filter((c) = { if (useAi) return strongOnly(c); // AI reads the body and subheadings, so rules keep only the sure ones return !longText || strongOnly(c) || c.heading || c.source === 'hashtag' || c.count = 2; }); // The catalog scan raises its bar in AI mode too // A name that only appears in the body is a candidate only at 3 or more mentions (2 or more for long names) if (useAi !inHead (count 3 !(e.key.length = 6 count = 2))) continue; And if the AI call fails, I bring the subheading and body rule candidates back in to fill the gap. It's a safety net I kept because the result shouldn't come back completely empty just because the AI died. So how much did it change? Words alone didn't seem convincing, so I took one video and ran it three different ways myself. It's a single vlog with well-organized chapters. ===== Rules only : 1,522ms, 17 places Yukgu Bangatgan, Sokcho Corn Salt Bread, Kokkiri Bunsik, Jorongbak, Jogae Jupging, Kare no Kare, Sampo Beach, Mandong Jegwa, Gangneung Gil Gamja, Jeong Coffee, Chilsadang, Saebarami Oneun Geuneul, Gangneung Dakgangjeong, Sageunjin Beach, Gangneung Gimbap, Sokcho IPARK Suite Hotel and Resort, Jumunjin Beach ===== AI on : 4,647ms, 20 places (the 17 above) + Kisa, Noren, Rito Rules alone already give 17 places. That's because the video has well-organized chapters. Turning AI on spends 3 more seconds and finds 3 more places. The added ones, Kisa, Noren and Rito, are all cafes, and their names weren't in the chapters, only inside sentences in the description. Next is the noise problem I mentioned earlier. I ran the code from before the role split and the code from after it on the same video and compared. ===== Before the fix (all rules + AI merged) : 20 places ... Sokcho Beach ... ===== After the fix (rules add only reliable sources when AI is on) : 20 places ... Jumunjin Beach ... Only in before: Sokcho Beach Only in after: Jumunjin Beach The count is the same 20 places, but one entry changed. Before the fix, Sokcho Beach , mentioned once in passing in the body, took up a slot, and Jumunjin Beach , which the video actually went to, had been pushed out of the top 20. This is where I confirmed that my idea that blindly adding candidates would make things better was wrong. Things like this piled up one by one too. It's a list that stops common nouns from being picked up as places, and every one of them is a word that actually got caught. const GENERIC_CORES = new Set([ 'famous restaurant', 'good eats', 'restaurant', 'cafe', 'seaside', 'beach', 'market', 'hotel', 'pension', 'resort', 'observatory', 'museum', 'park', 'harbor', 'station', 'terminal', 'rest area', 'parking lot', 'tourist spot', 'travel destination', 'attraction', 'lodging', 'guesthouse', 'travel', 'sea trip', 'domestic travel', 'full roundup', 'ocean view', 'ocean view hotel', 'course', 'tour', 'travel course', 'vlog', 'spring summer fall winter', 'four seasons', 'one day', 'two days', 'weekend', ].map(normalize)); As for why things like "spring summer fall winter" or "weekend" are in there, the video title contained that word and there happened to be a shop with the same name in the public data. False positives still remain Even after all this, it isn't perfect. If you look at an actual result screen, things like this are mixed in. The third item is classified as a tourist spot. It's a drain and toilet repair company. Names I can't find in the public data get looked up once more through Kakao Local search, and in that process a business called Sokcho Goseong Yangyang Drain Toilet Leak Detection Plumbing Master got pulled in as a tourist spot. It has three region names in its name, which seems to be why the search matched it. I decided that blocking all of this in code is practically impossible. Add one rule to block something and a perfectly fine shop gets caught by it somewhere else. So I changed direction, and instead of filtering, I chose to show how trustworthy each result is right on the screen . An exact match in the public data is "confirmed", something similar is "likely", and anything that only came from Kakao search is "needs checking". On the screen above, only 2 of the 10 places are confirmed, and the other 8 carry a "Couldn't find this in the public data" label. I judged this to be better than hiding it. Automatically extracted results can't be 100% right anyway, and marking them honestly and letting the user choose is less risky than pretending to be right and sending someone to the wrong place. This is how the places pulled from a link end up in the itinerary. There are eight commits from that day The commit log for the day I added this feature shows eight commits in a single day. 08-25 Analyze web page links that aren't YouTube too 08-25 Improve place extraction accuracy for blog and video links 08-25 Unify AI place extraction on OpenAI 08-25 Fix AI analysis showing up twice in the analysis source label 08-25 Fix places being dropped from list-style blog posts 08-25 Run the AI call and the search at the s