All articles
SEO & GrowthSeptember 30, 2026 6 min read

Internal Linking at Scale: The Graph Model That Beats Related-Posts Widgets

Related-posts widgets are how programmatic sites bleed authority. Here's how we model internal links as a graph, score edges, and ship a link plan that actually moves rankings.

Internal Linking at Scale: The Graph Model That Beats Related-Posts Widgets

Most programmatic sites we audit have thousands of pages and a related-posts widget doing 90% of the internal linking work. That widget is almost always picking neighbours by shared tag or vector similarity, which is fine for engagement and terrible for ranking distribution. If you treat internal linking as a graph problem instead of a template problem, you can push authority into the pages that actually convert.

This is how we model it, score it, and ship it on sites with 10k to 500k URLs.

Why related-posts widgets underperform

The default pattern on most CMS and headless stacks looks like this: on each page, query the last N posts sharing a taxonomy or the top N vector-nearest neighbours, render them at the bottom. It feels smart. It is not.

Three failure modes we see repeatedly:

  • Reciprocal loops. Page A links to B, B links back to A, and neither passes authority anywhere new. On a big programmatic tree this creates dense clusters of mutually-linking siblings and starves the pages one hop away.
  • Orphan drift. New pages don't get incoming links until they've earned tags or embeddings similar to existing content. Cold-start pages sit at depth 5+ from the homepage for weeks.
  • No intent weighting. A comparison page ("X vs Y") and a definition page ("what is X") get treated as equal candidates, even though one converts 20x better and deserves more inbound juice.

Related-posts widgets optimise for "is this relevant to the reader on this page." Internal linking should optimise for "where should authority flow across the entire site." Different question, different answer.

Model the site as a directed weighted graph

Start with the obvious representation. Every URL is a node. Every link between two URLs is a directed edge. Attach weights to both.

Node attributes we actually use

At minimum, store per node:

  • url
  • entity_type (e.g. product, category, comparison, guide, location)
  • depth_from_home (BFS from /)
  • impressions_28d, clicks_28d, avg_position from GSC
  • conversion_rate or a proxy (email signups, add-to-cart, whatever)
  • indexable (boolean — noindex, canonical target, robots status)
  • last_modified

We pull this into a Postgres table nightly. Nothing exotic.

Edge attributes that matter

For each link, store:

  • source_url, target_url
  • anchor_text
  • position (nav, body, footer, sidebar, related-widget)
  • is_followed
  • rendered (did it exist in pre-JS HTML, or only after hydration?)

The position field is the one most teams skip. A body link inside prose is worth substantially more than a link in a footer megamenu — Google has been open about position-based weighting for years. If you can't tell them apart, you can't reason about your own graph.

Score the edges you want to exist

Once you have the current graph, the interesting work is deciding which edges should exist that don't. This is a scoring problem, not a similarity problem.

For a candidate edge from source s to target t, we compute a rough score:

def edge_score(s, t, sim):
    if not t.indexable or s.url == t.url:
        return 0
    # topical fit — embedding cosine or shared entities
    topical = sim(s, t)  # 0..1
    # target needs help: high impressions, mediocre position
    opportunity = 0
    if 6 <= t.avg_position <= 20 and t.impressions_28d > 100:
        opportunity = min(t.impressions_28d / 1000, 1.0)
    # commercial value of the target
    value = t.conversion_rate_percentile  # 0..1
    # don't over-link deep pages from deep pages
    depth_penalty = 1 / (1 + max(s.depth - 3, 0))
    # avoid reciprocal
    reciprocal_penalty = 0.3 if edge_exists(t, s) else 1.0
    return topical * (0.4 + 0.3*opportunity + 0.3*value) \
           * depth_penalty * reciprocal_penalty

The exact weights are less important than the shape. You are rewarding topical fit, ranking opportunity, and commercial value; you are penalising depth and reciprocity. Tune on your own data.

Run this for every (source, target) pair where topical similarity clears a threshold. On a 50k-page site that's a few million candidate edges, which fits comfortably in memory or a single warehouse query.

Cap the fan-out per page

Give every source a link budget — say 8 to 15 contextual internal links beyond nav and footer. Take the top-scoring targets per source under that cap. This prevents the classic mistake where high-authority hub pages link to 200 leaves and dilute themselves.

Ship the link plan without a manual editor

A scored edge list is not internal linking. You still have to render the links on the page. Two patterns work depending on your stack:

Pattern A: contextual injection

For prose-heavy pages, run a small service at build time that takes the page's markdown, finds anchor-text candidates matching your planned targets, and rewrites the first sensible mention as a link. Keep a strict allowlist of anchors per target so you don't turn every mention of "pricing" into a link.

// pseudo
for (const edge of plannedEdges[pageId]) {
  const anchor = pickAnchor(edge.target, edge.allowed_anchors);
  markdown = injectFirstMatch(markdown, anchor, edge.target.url);
}

This works beautifully for guides, glossary entries, and comparison pages. It falls apart on thin templated pages where there's no prose to inject into.

Pattern B: structured link modules

For templated pages (product listings, location pages, spec sheets), replace the generic related-posts widget with 2 – 3 purpose-built modules driven by the edge list:

  • "Compare with" — targets of type comparison
  • "Related in {category}" — targets one taxonomy level up or sideways
  • "Popular in {region}" — targets sharing a location entity

Each module reads from the same scored edge table, filtered by relation type. The template stays static; the links change as scores change.

Measure what actually moved

Internal linking changes are the easiest SEO experiment to fake yourself out on, because seasonality and unrelated updates swamp the signal. A few things we insist on:

  1. Cohort the target pages. Split changed pages into deciles by how much incoming internal PageRank they gained. Compare clicks and position changes per decile against unchanged pages over 28 and 56 days.
  2. Watch depth, not just links. Track median crawl depth from / for indexable pages. If your link plan is working, this number goes down.
  3. Log the diff. Every deploy, store which edges were added or removed. When something moves in GSC six weeks later, you want to be able to answer "what did we change on this page?" in one query.

If you're already piping GSC into a warehouse (worth doing — we wrote about the GSC-to-warehouse growth loop separately), joining it to your edge diff is trivial.

Common traps

A few things we've walked into so you don't have to:

  • Recomputing the whole graph on every build. On large sites this is 20+ minute builds. Recompute the scored plan weekly, cache it, and let builds just read from it.
  • Ignoring rendered vs raw HTML. If your links only appear after client-side hydration, treat them as worth roughly nothing for planning purposes. Render them server-side or don't count them.
  • Anchor text monoculture. If every link to /pricing says "pricing", you're leaving topical context on the table. Rotate through 3 – 5 approved anchors per target.
  • Linking to noindex or canonicalised URLs. Sounds obvious. We find it on 80% of audits. Your edge scorer should hard-filter these.

Where we'd start

If you have a programmatic site and a related-posts widget doing the heavy lifting, do this in order:

  1. Crawl your own site and dump the current edge list to a table. Just seeing it is usually a small horror show.
  2. Join it with GSC data and mark every target that sits at positions 6 – 20 with real impressions. Those are your priority targets.
  3. Build the scoring function above, generate a plan, and ship it to one page type first — usually comparison or category pages, because they benefit fastest.
  4. Kill or heavily constrain the related-posts widget on those pages.
  5. Wait six weeks before judging results, and diff your edges every deploy so you can attribute movement.

Internal linking is one of the few SEO levers you fully control, don't need Google's cooperation for, and can ship on your own release cadence. Treat it like the graph problem it is and it will out-perform any widget you can bolt on.

#Programmatic SEO#Internal Linking#Site Architecture#Technical SEO

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project