All articles
SEO & GrowthAugust 24, 2026 6 min read

Schema Markup at Scale: Generating JSON-LD From Your Content Model Without Breaking Rich Results

Hand-writing JSON-LD is fine for ten pages. At ten thousand, schema becomes a compile step. Here's how we generate, validate, and monitor structured data without triggering manual actions.

Schema markup is one of those things that looks trivial in a tutorial and turns into a swamp the moment you have 40,000 URLs generated from a database. You start with a clean Product snippet, then marketing wants FAQs, then legal wants review disclaimers, then someone ships a template change that silently drops priceValidUntil and two weeks later Search Console lights up with "Invalid object" errors on 12,000 pages.

This is the article I wish we had when we started treating structured data as a build artifact instead of a copy-paste job.

Why hand-authored schema breaks at scale

The common failure mode is that JSON-LD lives in the template as a string. A designer edits the template, forgets to update the schema block, and now your Product says the item costs $49 while the visible page says $59. Google's Merchant guidelines are not gentle about that mismatch — expect a structured data manual action, not a warning.

The other failure mode is field drift. Your CMS adds a new attribute called warranty_months. Half your templates map it into Product.warranty.WarrantyPromise.durationOfWarranty, the other half ignore it. Six months later, nobody can answer "which pages have valid warranty schema?" without running a crawl.

Both of these disappear when schema is generated from the same content model that renders the page.

Treat schema as a derived artifact

The mental shift: your database rows are the source of truth, your page HTML is one projection of those rows, and your JSON-LD is another projection. They should share a validator, not a copy-paste history.

In practice this means three things:

  1. A typed content model (Zod, Pydantic, Prisma, whatever your stack uses) defines the shape of a page's data.
  2. A pure function maps that model to a JSON-LD object. No string templating.
  3. That function is tested and validated in CI, before deploy.

Here's the shape we use on most Next.js projects:

import { z } from "zod";

const ProductPageModel = z.object({
  sku: z.string(),
  name: z.string().min(3),
  description: z.string().min(50),
  priceCents: z.number().int().positive(),
  currency: z.enum(["USD", "EUR", "GBP"]),
  availability: z.enum(["InStock", "OutOfStock", "PreOrder"]),
  priceValidUntil: z.string().date(),
  brand: z.string(),
  images: z.array(z.string().url()).min(1),
  aggregateRating: z
    .object({ value: z.number().min(1).max(5), count: z.number().int() })
    .optional(),
});

type ProductPage = z.infer<typeof ProductPageModel>;

export function toProductJsonLd(p: ProductPage, url: string) {
  return {
    "@context": "https://schema.org",
    "@type": "Product",
    name: p.name,
    description: p.description,
    sku: p.sku,
    brand: { "@type": "Brand", name: p.brand },
    image: p.images,
    offers: {
      "@type": "Offer",
      url,
      price: (p.priceCents / 100).toFixed(2),
      priceCurrency: p.currency,
      priceValidUntil: p.priceValidUntil,
      availability: `https://schema.org/${p.availability}`,
    },
    ...(p.aggregateRating && {
      aggregateRating: {
        "@type": "AggregateRating",
        ratingValue: p.aggregateRating.value,
        reviewCount: p.aggregateRating.count,
      },
    }),
  };
}

The key detail: aggregateRating is only emitted when present. Google will flag AggregateRating with a rating count of zero as invalid, and empty strings are worse than omission. Optionality has to be a first-class concept in your generator.

The one-way rule

Data flows from model → HTML and model → JSON-LD. Never model → HTML → JSON-LD by scraping your own DOM. We've seen teams do it and it always ends the same way: a whitespace change in the template breaks the schema extractor and nobody notices for a month.

Validate at build time, not in Search Console

Search Console tells you about broken schema three to fourteen days after you ship. That's not a feedback loop; that's an incident.

We run two layers of validation in CI:

Layer 1: Schema.org shape validation. The typed model already handles this. If priceCents is missing, the build fails before a page renders.

Layer 2: Google Rich Results eligibility. This is subtler. A JSON-LD document can be valid Schema.org and still be ineligible for rich results because Google requires specific fields. Product needs offers OR review OR aggregateRating. Recipe needs image with specific aspect ratios. These rules change.

We use the Schema Markup Validator for structural checks and Google's Rich Results Test API for eligibility. In CI, we sample: run 20 randomly selected URLs from each template against the API on every PR that touches schema code, plus a full sweep nightly.

# Nightly sweep, parallelised
cat urls.txt | xargs -P 8 -I {} \
  node scripts/rich-results-check.js {} \
  >> reports/$(date +%F).jsonl

The output goes into BigQuery so we can trend eligibility over time per template. When a template's eligibility rate drops below 98%, that's a Slack alert, not a quarterly review item.

Handling the FAQ and HowTo landmines

In August 2023 Google gutted FAQPage and HowTo rich results. Many teams had built entire content strategies on those snippets and woke up to flat CTR. The lesson wasn't "don't use schema" — it was "don't build growth strategy on a single snippet type you don't control."

Our current rule for programmatic pages: emit schema for entities Google is unlikely to deprecate — Product, Article, BreadcrumbList, Organization, LocalBusiness, Event. These map to durable commercial intent. Emit FAQPage only when the page genuinely is a FAQ, not as a CTR hack.

This is also an AdSense-adjacent concern. Pages stuffed with fake FAQ blocks to game rich results tend to score poorly on brand safety reviews. If you're monetising with display ads, keep your schema honest.

Monitoring: the three signals that actually matter

GSC's Enhancements report is useful but slow. We supplement it with three internal signals:

1. Coverage rate per template

For each page template, what percentage of live URLs emit valid schema for the expected type? This should be 99%+. A drop means either data quality regressed in the CMS (missing brand, empty images) or the generator has a bug.

2. Rich result impressions in GSC

Query searchAppearance in the GSC API and split by page group. If Product snippets on /shop/* drop 30% week over week without a corresponding traffic drop, Google likely tightened eligibility for that subset. Investigate before the traffic follows.

3. Manual actions and structured data warnings

Poll the GSC API daily. Route both manualActions and Enhancement issues into your incident channel. A Structured Data: Product warning affecting 4,000 URLs is a P1, not a ticket someone picks up next sprint.

The migration path for existing sites

If you're staring at a codebase with schema strings sprinkled across 30 templates, don't rewrite everything in one PR. We do this in stages:

  1. Audit. Crawl the site, extract every JSON-LD block, classify by @type and template. You'll usually find that 80% of pages use 3–4 schema types.
  2. Model the top type first. Usually Product or Article. Build the generator, run it in shadow mode for a week — emit both the old and new schema, log diffs.
  3. Cut over one template. Ship. Watch GSC for 14 days. No regression? Move to the next template.
  4. Delete the string templates. Only after every template is migrated. Leaving both around invites the drift you were trying to eliminate.

Shadow mode is the underrated part. It catches cases where the old hand-written schema was subtly wrong in ways that were somehow working — and you don't want to "fix" those on a Friday afternoon.

Where AdSense and structured data intersect

One thing we've learned running content sites with both organic and display revenue: brand-safe schema matters. If your Article schema declares an author that doesn't exist on the page, or an datePublished that's five years off from the visible date, you're sending mixed signals to both Google Search and Google's ad quality systems. Ads served on pages with inconsistent metadata tend to get lower-value fills over time. Correlation isn't causation, but the correlation is consistent enough that we treat it as a rule.

Keep author entities real, dates accurate, and publisher information consistent across your Organization block and your ads.txt. Boring, but it compounds.

Where we'd start

If you have a programmatic site with more than a few hundred pages and hand-authored schema, do this in the next two weeks:

  • Pick your highest-traffic template. Extract its current JSON-LD and diff it against your database. Count the mismatches.
  • Build a typed generator for that one template. Ship it in shadow mode.
  • Wire up the Rich Results Test API in CI, even if it only samples 10 URLs per PR.
  • Add a daily job that pulls GSC Enhancement issues into Slack.

That's a week of engineering work and it will save you the next manual action email. If you want help auditing an existing programmatic setup, our SEO engineering services team does exactly this kind of work — but you can get most of the value from the checklist above without hiring anyone.

#Programmatic SEO#Structured Data#JSON-LD#Technical SEO#Rich Results

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project