Internal Linking at Scale: The Graph Model That Beats Related-Posts Widgets
Related-posts widgets are how programmatic sites bleed authority. Here's how we model internal links as a graph, score edges, and ship a link plan that actually moves rankings.

Most programmatic sites we audit have thousands of pages and a related-posts widget doing 90% of the internal linking work. That widget is almost always picking neighbours by shared tag or vector similarity, which is fine for engagement and terrible for ranking distribution. If you treat internal linking as a graph problem instead of a template problem, you can push authority into the pages that actually convert.
This is how we model it, score it, and ship it on sites with 10k to 500k URLs.
Why related-posts widgets underperform
The default pattern on most CMS and headless stacks looks like this: on each page, query the last N posts sharing a taxonomy or the top N vector-nearest neighbours, render them at the bottom. It feels smart. It is not.
Three failure modes we see repeatedly:
- Reciprocal loops. Page A links to B, B links back to A, and neither passes authority anywhere new. On a big programmatic tree this creates dense clusters of mutually-linking siblings and starves the pages one hop away.
- Orphan drift. New pages don't get incoming links until they've earned tags or embeddings similar to existing content. Cold-start pages sit at depth 5+ from the homepage for weeks.
- No intent weighting. A comparison page ("X vs Y") and a definition page ("what is X") get treated as equal candidates, even though one converts 20x better and deserves more inbound juice.
Related-posts widgets optimise for "is this relevant to the reader on this page." Internal linking should optimise for "where should authority flow across the entire site." Different question, different answer.
Model the site as a directed weighted graph
Start with the obvious representation. Every URL is a node. Every link between two URLs is a directed edge. Attach weights to both.
Node attributes we actually use
At minimum, store per node:
urlentity_type(e.g.product,category,comparison,guide,location)depth_from_home(BFS from/)impressions_28d,clicks_28d,avg_positionfrom GSCconversion_rateor a proxy (email signups, add-to-cart, whatever)indexable(boolean — noindex, canonical target, robots status)last_modified
We pull this into a Postgres table nightly. Nothing exotic.
Edge attributes that matter
For each link, store:
source_url,target_urlanchor_textposition(nav, body, footer, sidebar, related-widget)is_followedrendered(did it exist in pre-JS HTML, or only after hydration?)
The position field is the one most teams skip. A body link inside prose is worth substantially more than a link in a footer megamenu — Google has been open about position-based weighting for years. If you can't tell them apart, you can't reason about your own graph.
Score the edges you want to exist
Once you have the current graph, the interesting work is deciding which edges should exist that don't. This is a scoring problem, not a similarity problem.
For a candidate edge from source s to target t, we compute a rough score:
def edge_score(s, t, sim):
if not t.indexable or s.url == t.url:
return 0
# topical fit — embedding cosine or shared entities
topical = sim(s, t) # 0..1
# target needs help: high impressions, mediocre position
opportunity = 0
if 6 <= t.avg_position <= 20 and t.impressions_28d > 100:
opportunity = min(t.impressions_28d / 1000, 1.0)
# commercial value of the target
value = t.conversion_rate_percentile # 0..1
# don't over-link deep pages from deep pages
depth_penalty = 1 / (1 + max(s.depth - 3, 0))
# avoid reciprocal
reciprocal_penalty = 0.3 if edge_exists(t, s) else 1.0
return topical * (0.4 + 0.3*opportunity + 0.3*value) \
* depth_penalty * reciprocal_penalty
The exact weights are less important than the shape. You are rewarding topical fit, ranking opportunity, and commercial value; you are penalising depth and reciprocity. Tune on your own data.
Run this for every (source, target) pair where topical similarity clears a threshold. On a 50k-page site that's a few million candidate edges, which fits comfortably in memory or a single warehouse query.
Cap the fan-out per page
Give every source a link budget — say 8 to 15 contextual internal links beyond nav and footer. Take the top-scoring targets per source under that cap. This prevents the classic mistake where high-authority hub pages link to 200 leaves and dilute themselves.
Ship the link plan without a manual editor
A scored edge list is not internal linking. You still have to render the links on the page. Two patterns work depending on your stack:
Pattern A: contextual injection
For prose-heavy pages, run a small service at build time that takes the page's markdown, finds anchor-text candidates matching your planned targets, and rewrites the first sensible mention as a link. Keep a strict allowlist of anchors per target so you don't turn every mention of "pricing" into a link.
// pseudo
for (const edge of plannedEdges[pageId]) {
const anchor = pickAnchor(edge.target, edge.allowed_anchors);
markdown = injectFirstMatch(markdown, anchor, edge.target.url);
}
This works beautifully for guides, glossary entries, and comparison pages. It falls apart on thin templated pages where there's no prose to inject into.
Pattern B: structured link modules
For templated pages (product listings, location pages, spec sheets), replace the generic related-posts widget with 2 – 3 purpose-built modules driven by the edge list:
- "Compare with" — targets of type
comparison - "Related in {category}" — targets one taxonomy level up or sideways
- "Popular in {region}" — targets sharing a location entity
Each module reads from the same scored edge table, filtered by relation type. The template stays static; the links change as scores change.
Measure what actually moved
Internal linking changes are the easiest SEO experiment to fake yourself out on, because seasonality and unrelated updates swamp the signal. A few things we insist on:
- Cohort the target pages. Split changed pages into deciles by how much incoming internal PageRank they gained. Compare clicks and position changes per decile against unchanged pages over 28 and 56 days.
- Watch depth, not just links. Track median crawl depth from
/for indexable pages. If your link plan is working, this number goes down. - Log the diff. Every deploy, store which edges were added or removed. When something moves in GSC six weeks later, you want to be able to answer "what did we change on this page?" in one query.
If you're already piping GSC into a warehouse (worth doing — we wrote about the GSC-to-warehouse growth loop separately), joining it to your edge diff is trivial.
Common traps
A few things we've walked into so you don't have to:
- Recomputing the whole graph on every build. On large sites this is 20+ minute builds. Recompute the scored plan weekly, cache it, and let builds just read from it.
- Ignoring rendered vs raw HTML. If your links only appear after client-side hydration, treat them as worth roughly nothing for planning purposes. Render them server-side or don't count them.
- Anchor text monoculture. If every link to
/pricingsays "pricing", you're leaving topical context on the table. Rotate through 3 – 5 approved anchors per target. - Linking to noindex or canonicalised URLs. Sounds obvious. We find it on 80% of audits. Your edge scorer should hard-filter these.
Where we'd start
If you have a programmatic site and a related-posts widget doing the heavy lifting, do this in order:
- Crawl your own site and dump the current edge list to a table. Just seeing it is usually a small horror show.
- Join it with GSC data and mark every target that sits at positions 6 – 20 with real impressions. Those are your priority targets.
- Build the scoring function above, generate a plan, and ship it to one page type first — usually comparison or category pages, because they benefit fastest.
- Kill or heavily constrain the related-posts widget on those pages.
- Wait six weeks before judging results, and diff your edges every deploy so you can attribute movement.
Internal linking is one of the few SEO levers you fully control, don't need Google's cooperation for, and can ship on your own release cadence. Treat it like the graph problem it is and it will out-perform any widget you can bolt on.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

Schema Markup for Programmatic Pages: A Validation Pipeline That Catches Drift Before Google Does
Structured data on programmatic pages breaks silently. Here's the validation pipeline we run in CI to catch schema drift before Search Console flags 40,000 URLs at once.
Canonical Tags on Programmatic Pages: The Duplicate Content Traps We Keep Finding
Canonical tags look trivial until you're running 200k programmatic pages and Google decides half of them are duplicates. Here's what actually breaks, how to diagnose it, and the rules we now enforce at template time.
GSC API to Warehouse: Building a Query-Level Growth Loop That Beats the UI
The Google Search Console UI hides your best growth signals behind 1,000-row caps and aggregated data. Here's how we pipe GSC into a warehouse and turn it into a weekly query-level growth loop.
