After moving to React, every KakaoTalk link preview looked the same
My React SPA showed the same KakaoTalk preview for every post. How Express inserts per-post meta tags, JSON-LD and a noscript body without SSR, and how canonical burned me later.
#SEO #Frontend #SPA #OG #React
I shared one of my blog posts on KakaoTalk. A preview card did show up, but the title, the thumbnail and the description were the same as for every other post. Sending a link to a different post made no difference. Whatever post I sent, it showed up with the preview for the blog's main page. Search was similar. Typing a post title into Google rarely found it, and when it did, the description was off. This never happened when PHP was rendering the HTML on the server. The problem appeared after I moved to React. This post is a record of how I solved that without an SSR framework like Next.js, by having the server insert only the per-post meta tags. I'm also writing down how that same structure burned me badly later on. Why did everything look the same? My Express app ends with this one route. // app.js // Handle React SPA routing app.get('*', (req, res) = { res.sendFile(path.join(__dirname, 'dist', 'index.html')); }); Whether it's /blog/56, /blog/60 or the main page, every path gets the same single index.html. That's what an SPA is. The actual post content gets drawn on screen afterwards, once the JavaScript runs. The problem is that the KakaoTalk bot and search crawlers don't wait for JavaScript. They look only at the HTML the server sends first and leave. And that HTML doesn't contain a single character that differs from post to post. The title and the og tags are all the defaults baked into index.html, so from a crawler's point of view every post on my blog is the same page. It doesn't matter that JS changes the title later. The KakaoTalk bot is already gone by then. So insert it on the server instead of in JS There were two options. Switch to an SSR framework like Next.js, or have the server put the per-post meta into that first HTML the crawler sees before sending it. The first meant rewriting the whole frontend for the sake of one blog's SEO, which was too much. So I went with the second. When a request comes in for /blog/:bid, instead of just sending index.html, the server reads that post's title, description and thumbnail from the DB, inserts the meta tags, and then sends it. It isn't running React on the server, it's editing an HTML string, so not one line of frontend code gets touched. Order mattered most The first thing I learned doing this was the route registration order. The app.get('*') from above grabs everything, so the handler for posts has to go before it. Put it after and it never gets called. // app.js const { blogMetaHandler, sitemapHandler } = require('./util/seo'); // sitemap goes before the static dist/sitemap.xml app.get('/sitemap.xml', sitemapHandler); app.use(express.static(path.join(__dirname, 'dist'))); // Intercept only post detail pages and inject per-post meta app.get('/blog/:bid', blogMetaHandler); // Everything else is the SPA app.get('*', (req, res) = { /* index.html */ }); The point is that /blog/:bid sits between static and '*'. Put it before static and it becomes a problem when there are static files under /blog/, and put it after '*' and it never gets called at all. If this order is off, the handler never runs no matter how well the code is written. Swapping the meta tags What the handler does is simple. It reads the post from the DB, strips the default meta out of index.html, and swaps in the per-post tags. // util/seo.js // Strip all the existing title/og/twitter tags html = html .replace(/ title [\s\S]*? \/title /i, '') .replace(/ meta\s+property="og:[^"]*"[^ ]* /gi, '') .replace(/ meta\s+name="twitter:[^"]*"[^ ]* /gi, ''); // Insert the per-post tags right before /head const tags = ` title ${esc(meta.title)} /title meta property="og:title" content="${esc(meta.title)}" / meta property="og:description" content="${esc(meta.description)}" / meta property="og:image" content="${esc(meta.image)}" / meta name="twitter:card" content="summary_large_image" / meta name="twitter:image" content="${esc(meta.image)}" / `; html = html.replace(' /head ', tags + ' /head '); The thumbnail in the KakaoTalk preview is this og:image. Setting twitter:card to summary_large_image makes it show up as a large image card. If I only added tags without removing the existing ones, there would be two of each and no telling which one wins, so I strip first and then insert the new ones. And every value that comes from user input is wrapped in esc(). A quote or an angle bracket in a post title would break the whole tag, so this is a security issue and a stability issue at the same time. Fetching this post with curl right now gives this. Each post goes out with its own og:title, og:description, og:image and og:url. What matters is that it's visible with curl, without a browser. That's exactly the state a crawler sees. The thumbnail is the post's first image, or the logo if there isn't one For the thumbnail in og:image, I use the first image attached to the post. If there isn't one, it falls back to the blog logo. const fileRows = await q( `SELECT f.path FROM BoardFile bf JOIN Files f ON f.id = bf.file_id WHERE bf.board_id = ? ORDER BY bf.seq ASC LIMIT 1`, [bid] ); const image = fileRows.length ? `${BASE}/${fileRows[0].path}` : `${BASE}/logo.png`; The image address here has to be an absolute URL, no exceptions. The KakaoTalk bot fetches this URL separately from outside my site, so a relative path like /files/... means no thumbnail. That's why I always put the domain (BASE) in front. I left that one thing out once and got a card where the title was fine but the thumbnail was blank. Going one step further for Google, JSON-LD For KakaoTalk the og tags are enough, but if Google is given structured data (JSON-LD), it shows richer search results with things like the author and the publish date. I embedded the post information as JSON with the BlogPosting type. I got burned once here. This JSON goes inside a script tag, and if the post body contains the string for a closing script tag as is, the browser thinks the script ends there and closes the tag. The rest of the body spills out of the script and the page breaks. So I replaced every opening angle bracket in the JSON string with a Unicode escape. // util/seo.js function buildJsonLd(meta) { const data = { '@context': 'https://schema.org', '@type': 'BlogPosting', headline: (meta.h1 || meta.title).slice(0, 110), description: meta.description, ...(meta.bodyText ? { articleBody: meta.bodyText } : {}), image: meta.image, datePublished: meta.published, author: { '@type': 'Person', name: 'Jaeyong Choi', url: meta.base }, mainEntityOfPage: { '@type': 'WebPage', '@id': meta.url }, }; // Escape so a sequence like /script can't break the script return JSON.stringify(data).replace(/ /g, '\\u003c'); } Inside JSON, \u003c is read as the same character as a plain angle bracket, so the data doesn't change, and the HTML parser never sees a closing tag. Without this, everything looks fine most of the time, and then one post that happens to have that string in its body breaks the entire page. It's exactly the kind of trap I hate most. Invisible most of the time, and it only goes off on specific data. Crawlers couldn't read the body at all The meta was solved, but the body was the next problem. Because it's an SPA, the root div in the initial HTML is completely empty. There are no body keywords for a crawler to read. So I put the actual post text in a noscript and sent it along. const body = ` noscript article h1 ${esc(meta.h1)} /h1 p ${esc(meta.description)} /p div ${esc(meta.bodyText)} /div /article /noscript `; Pulling the body straight out of the HTML drags the p and strong tags along, so I stripped all the tags with a regex and put in plain text only. Too long and the page gets bloated, so I cut it at 12,000 characters. In a browser it's not visible because it's noscript, and only crawlers that don't run JavaScript see it. The sitemap is built from the DB too sitemap.xml tells crawlers "here's the list of my posts", and if it's a static file I have to edit it by hand every time I publish. So when a request comes in, the server just reads the post list from the DB and builds the XML on the spot. A newly published post goes into the sitemap automatically. That's also why the sitemap handler sits before static in the route order above. If an old sitemap.xml is left in dist, static serves that one first. One step further, the crypto site draws its thumbnail in real time On the blog, what differs from post to post is the "content", so using an attached image as the thumbnail was enough. The crypto site was a different story. What changes there every time is the price. The Bitcoin price and the kimchi premium, the price gap on Korean exchanges, change every second, and there's no way to show that with one static image. So on the crypto side the thumbnail is drawn on the fly every time. /og-image takes the current price, builds a single card as SVG, converts it to PNG and responds with that. app.get('/og-image', async (req, res) = { if (!btcPrice) btcPrice = await fetchLiveBtc(); // fetch the price const svg = buildOgSvg({ btcDisplay, kimpDisplay, kimpColor }); // draw it as an SVG card const png = await sharp(Buffer.from(svg)).png().toBuffer(); // convert to PNG res.setHeader('Content-Type', 'image/png'); res.send(png); }); Drawing it as SVG means the text position, font size and color can all be set freely in code. I paint the kimchi premium green when it's positive and red when it's negative, so the thumbnail alone shows what the mood is right now. And the share HTML (/share) is a page meant only for crawlers. Its og:image points to the dynamic image above, and when a real user comes in, JS sends them on to the main site. But drawing a new image for every share is expensive. Social media bots sometimes scrape the same link several times. So I added a cache that reuses the generated PNG for 60 seconds when the parameters are the same. if (ogCache.key === cacheKey ogCache.png (now - ogCache.at) OG_CACHE_TTL) { return res.send(ogCache.png); // within 60 seconds, don't redraw } In the end the blog and the crypto site solved the same problem in opposite ways. On the blog the content is the variable, so I swapped the meta, and on the crypto site the data is the variable, so I drew the image. The approach depended on which part of what's being shown changes. This code burned me badly later The structure I built here had a trap in it. I had the server insert the meta tags, but for canonical I was still using the value hardcoded into the built index.html. link rel="canonical" href="https://blog.mydomain/" While it was only the blog, there was no problem. But later the same server started serving the root domain too, and crawlers coming in through the root began seeing this tag. The site itself was effectively declaring "the canonical version of this page is at a different address". The sitemap submitted root addresses while canonical pointed to the subdomain, so to Google the root became a page with nothing to evaluate. I only found the cause after AdSense warned me about "screens without publisher content". If the server was going to render the meta tags, it should have rendered all of them. The halfway state, where the server renders a few and the rest use build-time values, was the most dangerous one. I've since fixed it so that canonical is built from the host being accessed, and the og:url in the curl output above showing the domain being accessed comes from that same fix. If an SPA needs KakaoTalk previews or search visibility and changing frameworks feels like too much, this approach is worth trying once. I'd just suggest deciding at the start that "the server is responsible for every tag inside head". Do only half of it and, like me, it blows up later somewhere that's hard to find.