All articles
SEO & GrowthAugust 26, 2026 6 min read

Internal Linking at Scale: Building a Link Graph That Actually Distributes PageRank

Most programmatic sites treat internal links as a footer widget. Here's how to model your site as a directed graph, compute link equity flow, and inject links that move rankings — not just fill space.

Internal linking is the last thing most programmatic SEO teams take seriously, and the first thing that starts leaking equity once you cross a few thousand URLs. "Related posts" widgets and breadcrumb trails aren't a strategy — they're wallpaper. If you're running a real content engine, your internal link structure needs to be modelled, computed, and deployed like any other data product.

This is how we think about it when we're brought in to fix a stalled programmatic site.

The mental model: your site is a directed graph

Every page is a node. Every <a href> on that page is a directed edge. Google's crawler walks this graph, and a simplified version of PageRank still describes how authority flows through it. You don't need to reimplement Google — you need a good-enough model that tells you which pages are starving and which are hoarding.

The common failure modes we see:

  • Orphan clusters: whole content silos with zero inbound internal links because nobody wired them into navigation.
  • Equity sinks: pages with 400 inbound links and one outbound link to the homepage, wasting the flow.
  • Reciprocal loops: category A links to category B links to category A, and neither reaches the money pages.
  • Homepage hoarding: every template links to / in the logo, so 60% of computed authority pools at a page that doesn't need it.

You can't see any of this without building the graph.

Build the graph before you optimise it

Step one is a crawl that outputs edges, not just URLs. We usually run this weekly as part of the SEO data warehouse. A minimal Python approach with httpx and selectolax:

import httpx
from selectolax.parser import HTMLParser
from urllib.parse import urljoin, urlparse

def extract_edges(url: str, html: str, domain: str):
    tree = HTMLParser(html)
    edges = []
    for a in tree.css('a[href]'):
        href = a.attributes.get('href', '')
        target = urljoin(url, href)
        if urlparse(target).netloc != domain:
            continue
        # strip fragments and tracking
        target = target.split('#')[0].split('?')[0]
        edges.append({
            'source': url,
            'target': target,
            'anchor': (a.text() or '').strip()[:200],
            'rel': a.attributes.get('rel', ''),
            'in_nav': _is_in_nav(a),
        })
    return edges

Push the edges into a table — Postgres works fine up to a few million rows, BigQuery beyond that. You now have a queryable link graph.

Compute equity flow, not just link counts

Internal link count is a lazy metric. A link from a page with 20 outbound links is worth more than a link from a page with 200. That's the whole point of PageRank.

You don't need to run the full damping-factor iteration in production — networkx will do it on a few hundred thousand nodes in under a minute:

import networkx as nx
import pandas as pd

edges = pd.read_sql('SELECT source, target FROM internal_edges', conn)
G = nx.DiGraph()
G.add_edges_from(edges.itertuples(index=False, name=None))

pr = nx.pagerank(G, alpha=0.85, max_iter=100)

scores = pd.DataFrame(pr.items(), columns=['url', 'internal_pr'])
scores.to_sql('page_equity', conn, if_exists='replace')

Join internal_pr against your GSC impressions and clicks data. Two patterns jump out immediately:

  1. Pages with high internal PR but low impressions — you're overweighting pages that don't rank for anything valuable.
  2. Pages with high impressions but low internal PR — pages Google likes that your own site treats as second-class. These are the easiest wins on the site.

The mismatch report

We generate a weekly mismatch report ranked by opportunity. The formula is rough but effective:

opportunity = impressions_last_28d * (avg_ctr_position_1_to_3 - current_ctr)
              * (1 / (internal_pr_percentile + 0.1))

Sort descending. The top 200 rows are your internal linking backlog for the quarter.

Injecting links: templated vs. contextual

Programmatic sites usually get internal linking wrong in the same way: they bolt on a "Related" module that pulls from the same category. That's not linking, that's clustering. Google can tell the difference between a link that sits in a sidebar module across 40,000 pages and a link that appears inside the body copy with a relevant anchor.

We split internal link injection into three tiers:

  • Structural links: nav, breadcrumbs, footer. Set once, rarely changed. Optimise for crawl paths, not equity.
  • Templated modules: "Related X" blocks. Useful for topical clustering but heavily discounted by Google. Keep them, don't rely on them.
  • Contextual body links: entity-aware links inserted into paragraph text. This is where the equity actually moves.

The contextual layer is where engineering pays off. You need a phrase-to-URL index — essentially a dictionary of anchor candidates mapped to the canonical page that should own that entity. When a page is rendered, you scan the body for the highest-value unlinked phrases and inject links up to a per-page cap.

A sane phrase-to-URL index

# Simplified: real version handles stemming, casing, and priority
phrase_index = {
    'react server components': '/blog/react-server-components-guide',
    'postgres full-text search': '/blog/postgres-fts-vs-elasticsearch',
    'headless commerce': '/services/ecommerce',
}

def inject_links(html: str, current_url: str, max_links: int = 4):
    injected = 0
    for phrase, target in sorted(phrase_index.items(), key=lambda x: -len(x[0])):
        if target == current_url or injected >= max_links:
            continue
        # only replace first occurrence, only if not already inside an <a>
        html, replaced = _safe_replace(html, phrase, target)
        if replaced:
            injected += 1
    return html

Rules we've learned the hard way:

  • One link per target per page. Google ignores the second and third occurrence anyway, and repeated anchors look manipulative.
  • Longest-phrase-first matching. Otherwise "React" eats "React Server Components".
  • Never inject inside headings, code blocks, or existing anchors. This bug ships to production more often than you'd think.
  • Cap total contextual links per page. In our experience, three to six body links beats twenty.

Anchor diversity without keyword stuffing

One trap of automated injection: every internal link to /services/mobile-app-development uses the anchor "mobile app development". Google's had spam signals for exact-match anchor patterns since roughly forever, and internal links aren't immune.

Store each target with a rotating list of acceptable anchors:

{
  "url": "/services/mobile-app-development",
  "anchors": [
    "mobile app development",
    "native iOS and Android builds",
    "our mobile engineering team",
    "cross-platform app work"
  ]
}

Rotate deterministically based on a hash of the source URL so the same page always renders the same anchor — stable for diffs and QA — but the overall distribution across the site looks natural.

Measuring whether it worked

Don't ship internal linking changes without a measurement plan. The signal is slow and noisy. What we track:

  • Crawl frequency on the boosted pages from server logs. Should rise within 2 – 4 weeks.
  • Impressions in GSC for the target URLs, segmented against a control cohort of similar pages you didn't touch.
  • Average position for queries the target pages already rank for on page 2 – 4. These move first.
  • Internal PR delta in your own graph — sanity check that the intervention actually did what you modelled.

We rarely see meaningful ranking movement inside 14 days. Six to eight weeks is realistic for a site with steady crawl coverage. If nothing moves at all after ten weeks, the problem probably isn't internal linking — it's content quality or intent match.

Where we'd start

If you inherit a programmatic site tomorrow and want to fix internal linking without a six-month project:

  1. Crawl the site and dump edges to a table. One afternoon.
  2. Run PageRank on the graph and join against GSC. One afternoon.
  3. Pull the top 50 mismatch pages — high impressions, low internal PR — and hand-review them. One day.
  4. Build a minimal phrase-to-URL index for those 50 targets, wire it into your render pipeline behind a feature flag, and roll out to 10% of pages. One week.
  5. Measure at four and eight weeks against a holdout cohort before rolling to 100%.

That's the whole loop. It's not glamorous, and it doesn't need an LLM. It needs a graph, some SQL, and the discipline to measure. If you want a hand wiring this into an existing content engine, that's the kind of work our growth engineering team does day-to-day.

#Programmatic SEO#Internal Linking#Site Architecture#Growth Engineering

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project