Engineering
Programmatic SEO in Next.js: JSON-LD and Sitemaps
Abhishek Bahukhandi

Programmatic SEO in Next.js is mostly an exercise in not repeating yourself. One module holds every article on this blog, and three separate build-time consumers read it: the route generator, the sitemap, and the image script. Nobody hand-writes a meta tag, a JSON-LD block, or a sitemap entry.
This post is a walk through that pipeline as it actually exists in this repo — what the data module feeds, how the structured-data graph is wired together by @id, and the two places where the convenient option turned out to be the wrong one.
What programmatic SEO in Next.js actually means here
The phrase gets used for spam farms, so it is worth being precise. We mean one narrow thing: pages and their metadata are derived from a data source instead of authored one file at a time. The content is still written by a human. What is generated is the plumbing around it — routes, canonicals, Open Graph tags, the structured-data graph, the sitemap row, and the social card.
The payoff is that publishing an article is a single append to one array. The cost is the thing that makes programmatic SEO risky at any scale: every mistake in the plumbing is replicated across every page at once, which is why the verification step at the end of this post exists.
One data module, three build-time consumers
The data module is a plain JavaScript file exporting an array of post objects, plus a handful of derived helpers. It is 410 KB on disk and it never reaches the browser — a point worth dwelling on, because it is the thing that makes the approach viable at all.
generateStaticParams turns rows into routes
The dynamic segment maps every post to a slug, and the build emits one prerendered HTML file per article:
export function generateStaticParams() {
return getAllBlogPosts().map((post) => ({ slug: post.slug }));
}
In the build output that route shows up as SSG rather than server-rendered, which is exactly what you want for content that changes once a day. We went through the server/client boundary reasoning for this route in detail in where to draw the server vs client component line; the short version is that the article body is rendered to HTML on the server and handed to the client component as a string.
The data module stays on the server
The numbers from our own build make the case better than the theory does. The /Blogs/[slug] route reports a 2.23 kB route size against 107 kB first-load JavaScript, while the module it reads is 410 KB. None of those 410 KB are in the bundle, because nothing in the client graph imports the module — the server reads it, renders one article, and passes finished HTML across the boundary.
Get this wrong and the failure is quiet. Import the data module from anything marked with the client directive and every reader downloads all 25 articles to read one. The same discipline applies to heavier dependencies; we wrote about keeping a code editor's bundle under control in building the in-browser editor with CodeMirror 6.
Building a cross-linked JSON-LD @graph
Structured data is where most programmatic SEO setups get sloppy, usually by emitting four unrelated script tags that each redescribe the same author and the same organisation.
Why a graph rather than separate blocks
A @graph is a flat array of nodes where each node gets an @id, and relationships are expressed as references to those ids. The article does not inline its author; it says the author is the node at that id. The practical effect is that one person is described once, and twenty-five posts point at that single entity rather than at twenty-five lookalike copies.
The @id convention that does the linking
Ids only work if they are stable and predictable. Ours are derived from the site URL and the post URL, never hand-typed:
Node ids we emit per post
{url}#article— the BlogPosting itself.{url}— the WebPage. The article points at it withmainEntityOfPage.{url}#primaryimage— an ImageObject, referenced by both the page and the article.{url}#breadcrumb— the BreadcrumbList for Home, Blog and this post.{url}#faq— the FAQPage built from the post's question list.{site}#author-{name}— a Person node, shared across every post that person wrote.
Nodes that live in the root layout
Two nodes are not per-post at all. {site}#organization and {site}#website are emitted once from the root layout, on every page of the site. The per-post graph then refers to them: the article's publisher is a reference to the organization id, and the author's worksFor is the same reference.
That split is deliberate, and it is visible in the generated HTML. A rendered post contains two JSON-LD blocks — a 770-byte block from the layout holding Organization and WebSite, and a 6,242-byte block from the page holding the other six types. Eight @type values in total, about 8% of the page's HTML.
Where the script tag belongs
The Metadata API has no field for arbitrary script tags, so structured data cannot ride along with the title and canonical. The Next.js JSON-LD guide puts it plainly: the current recommendation is to render structured data as a script tag in layout.js or page.js. Both of ours are server components, so the markup is in the initial HTML and needs no JavaScript to be read.
Escaping the payload
The same guide flags something easy to miss. JSON.stringify does not escape the < character, so any string that reaches the payload is, in principle, able to close the script tag early. The recommended fix is a single replace, swapping that character for its unicode escape — which is valid JSON and parses back to the same character, so consumers see identical data:
<script
type="application/ld+json"
dangerouslySetInnerHTML={{
__html: JSON.stringify(data).replace(/</g, "\u003c"),
}}
/>
Our schema fields are short strings — titles, descriptions, FAQ answers, alt text — and none of them currently contain markup. That is precisely why the escape belongs in the shared component rather than in a reviewer's head: the field that eventually carries an angle bracket will be added by someone who has never read this file.
The one thing to take away
Derive every SEO artifact from the same data, and give the structured-data nodes stable ids so they can reference each other instead of repeating themselves. One append to one array should produce the route, the metadata, the graph, the sitemap row and the social card — and a verification script should assert all of it before anything ships.
Generating sitemap.xml from data instead of maintaining it
This blog used to carry a hand-maintained public/sitemap.xml. It was wrong within a week, which is the normal fate of such files.
The shape Next.js expects
The sitemap.xml file convention wants a default-exported function returning an array of objects with url, lastModified, changeFrequency and priority. Ours concatenates a short list of static routes with one entry per post:
...posts.map((post) => ({
url: SITE_URL + "/Blogs/" + post.slug,
lastModified: new Date(post.updatedAt || post.publishedAt),
changeFrequency: "monthly",
priority: 0.85,
}))
Our current build emits 33 URLs into a 6.0 KB document from a 174 B route. The /Blogs index gets a small refinement: rather than today's date, its lastModified is the newest post's date, so the index claims to have changed exactly when it did.
A static file in public/ shadows the route
The migration has one trap worth stating loudly. A file at public/sitemap.xml and a sitemap.js route both answer /sitemap.xml, and the static file wins. Delete the old file, and leave a comment in the route saying why it must stay deleted — otherwise someone restores it from an old branch and the generated sitemap silently stops being served.
The lastmod churn trade-off
Our static routes use new Date() for lastModified, which means every build stamps them with the build time whether or not those pages changed. It is honest about the deploy and dishonest about the content. Post entries do not have this problem because they read the post's own date. If we were starting over, the static routes would carry hard-coded dates bumped by hand when the page actually changes.
When one sitemap stops being enough
At 33 URLs this is a non-issue, but the ceiling is worth knowing before you hit it. Google's limit is 50,000 URLs per sitemap file, and Next.js provides generateSitemaps for splitting the output into numbered files that you slice from your data by id. A daily-publishing blog will not get there this decade; a product catalogue might on day one.
Per-post OG images at build time with sharp
Every post needs a 1200x630 social card. Two options exist and the obvious one is not the one we took.
Why not ImageResponse
Next.js ships ImageResponse from next/og, which renders JSX to PNG and is genuinely the fastest way to start. The constraint is in its API reference: it converts HTML and CSS through Satori and Resvg, and only flexbox plus a subset of CSS properties are supported, with advanced layouts such as display: grid explicitly not working.
For our cards that subset was the wrong shape. The artwork is a generated geometric field — concentric rings, diagonal rules, a dot matrix — which is naturally SVG, not flexbox. So the generator writes SVG directly and composites the brand mark with sharp.
Deterministic artwork from the slug
The script takes three arguments and nothing else:
node scripts/generate-og-image.mjs \
--slug="programmatic-seo-nextjs-app-router" \
--title="Programmatic SEO in Next.js: JSON-LD and Sitemaps" \
--category="engineering"
Two choices make it boring in the good way. The pattern variant is picked by hashing the slug, so a given post's artwork never changes between runs. And there is no network call, no API key and no image model in the path — the publishing routine cannot fail to produce artwork, which is the entire reason it was built this way rather than calling an image service.
Two files per post, one titled and one not
One command writes both cards. og.png carries the headline and is what a social preview unfurls. thumb.png is deliberately title-free, because the listing grid prints the headline next to the image and a title baked into the artwork would render twice.
Three practical notes on committing generated images:
- SVG has no text wrapping, so line breaks are computed from an estimated glyph width before the SVG is built. Bold sans-serif averages around 0.52em per character.
- The brand mark is composited as a raster layer on a white chip rather than drawn into the dark card, because the mark is designed for light backgrounds.
- They add up. 50 PNGs across 25 posts is about 7.8 MB in the repo, averaging 162 KB per card. Fine now; worth revisiting at a few hundred posts.
The verification that runs before anything ships
Generated SEO plumbing fails silently. A broken @id reference does not throw, a missing FAQ block just quietly drops a schema type, and a second h1 in a body is invisible until an audit flags it months later. So the publish step asserts the output rather than trusting it:
const h = fs.readFileSync(".next/server/app/Blogs/" + slug + ".html", "utf8");
const blocks = [...h.matchAll(/<script type="application\/ld\+json">([\s\S]*?)<\/script>/g)];
const types = blocks.flatMap((x) => {
const d = JSON.parse(x[1]);
return (d["@graph"] || [d]).map((n) => n["@type"]);
});
if (types.length < 6) throw new Error("fewer than 6 schema types");
if ((h.match(/<h1/g) || []).length !== 1) throw new Error("expected exactly one h1");
It reads the prerendered HTML, parses every JSON-LD block, counts the types and the heading levels, and throws rather than publishing. Three things make it worth the twenty lines:
- It parses the real build output, not the source. Anything that breaks during rendering is caught.
JSON.parseon each block is itself the test. Malformed JSON-LD fails loudly instead of being ignored by crawlers.- Counting
h1tags catches the most common body-content mistake, since the page renders the post title as the only top-level heading and a stray one in the body is stripped at render time.
If you are building something similar, the lesson that transfers is not the specific schema types. It is that programmatic SEO needs an assertion attached to every generated artifact, because the whole point of generating twenty-five pages from one array is that nobody is reading those twenty-five pages.
The same pipeline publishes every article here, including this one. If you want to see what the platform it serves actually does, Taqari runs free AI mock interviews — and the engineering posts on this blog are mostly notes from building it.
Frequently asked questions
What is programmatic SEO in Next.js?
+
It means deriving pages and their metadata from a data source rather than hand-authoring each one. In the App Router that usually combines generateStaticParams to turn rows into routes, generateMetadata for per-page titles and canonicals, and a sitemap.js that reads the same data.
Should JSON-LD go in generateMetadata or in the page?
+
In the page. The Metadata API has no field for arbitrary script tags, so structured data needs its own element. The Next.js JSON-LD guide recommends rendering it as a script tag in layout.js or page.js, which also puts it in the initial HTML where crawlers see it without running JavaScript.
Why use a JSON-LD @graph instead of separate script tags?
+
A @graph lets nodes reference each other by @id instead of repeating themselves. The article node points at one author node and one image node rather than inlining both, so the author is described once and every post refers to that same entity.
Does a sitemap.xml file in public/ conflict with sitemap.js?
+
Yes. Both resolve to the same /sitemap.xml path and the static file in public/ wins, so a stale hand-maintained file silently shadows the generated route. If you migrate to sitemap.js, delete the old file and keep it deleted.
How many URLs can one Next.js sitemap hold?
+
Google's limit is 50,000 URLs per sitemap file. Below that a single sitemap.js is fine. Above it, Next.js provides generateSitemaps to split the output into numbered sitemap files that you slice from your data by id.
Should I generate OG images with ImageResponse or sharp?
+
ImageResponse is the quickest path and renders JSX, but it runs on Satori and supports only flexbox plus a subset of CSS, so grid layouts will not work. A sharp script trades convenience for full SVG control and files you can commit and inspect.