AdvancedLinkTraining.com logo — a free link building course by Bill HartzerAdvanced Link TrainingA resource by Hartzer.com

Reading the link graph

What the link graph is

Your backlink report is a flattened one-hop view of a network. Learning to see the network is the skill.

Lesson 11 of 50Module 3 · Reading the link graph6 min read

After this lesson you should be able to

  • Describe a link profile as a network of nodes and edges rather than a list of rows
  • Number the tiers of a profile correctly and say what each one can and cannot tell you
  • Distinguish the Fresh index from the Historic index and explain why the gap is a crawl window
  • Report referring domains and distinct links instead of raw link counts

A graph, not a list

Two pieces of vocabulary and then we can work. A node is a page or a domain. An edge is a link between two nodes, and it has a direction — the link points one way. A link graph is the whole set of nodes and edges, which for the web as a whole is a structure of trillions of edges that no single crawler sees completely.

Your backlink report is a query against that structure: give me every edge whose head is this domain. It comes back as a table because tables are easy to render, but the answer it returns is one hop deep. Everything upstream of the linking page — which is where the value comes from — has been discarded before you saw the screen.

That discarding is the source of nearly every bad link purchase I have investigated. Consider two links, both from domains that a tool scores identically. The first sits in a post on a site's news section, three clicks from the homepage, linked from a category page and a monthly archive, on a domain where dozens of other pages link internally to that post. The second sits on a page created for the purpose, reachable only by its URL, linked from nothing. The domain metric cannot tell these apart, because the domain metric is not about the page. Only the graph can.

There is one honest simplification worth keeping. Search engines compute link value at page level and roll it up in ways nobody outside them can reproduce exactly. Every third-party tool is approximating with its own crawl and its own arithmetic. So when I say a link is worth something, I mean the structural argument for it is sound, not that I have measured the transfer. Nobody has.

Tiers, and why they start at zero

Tier numbering counts hops away from the site you are looking at:

  • Tier 0 — the root. The domain under analysis.
  • Tier 1 — the domains that link directly to it. This is your backlink report.
  • Tier 2 — the domains that link to your tier 1 domains.
  • Tier 3, 4, 5 — one further hop each time, outward into the graph.

Two warnings about the word. First, tiered link analysis and tiered link building are different subjects that share a vocabulary. Analysis is reading upstream to judge whether a link is real. Tiered link building is pointing links at your own links to inflate their value, and it is a link scheme. This module is entirely about the first. The tactic reference covers the second so that you recognize it when a vendor describes it in flattering language.

Second, tier membership is not exclusive. A domain can be in your tier 1 and your tier 3 at once, because there are multiple paths through a graph and the shortest one determines the label you see. When a tool tells you a domain is at tier 4, it is telling you the shortest path it found, not the only one.

The practical consequence is that tier depth is a rough measure of distance, not a hierarchy of importance. A tier 3 domain is not one third as relevant as a tier 1 domain. It is three hops away, and what matters is what the population at three hops looks like collectively.

Fresh and Historic are two indexes, not two eras

This is the single most misread pair of numbers in link analysis, and it is misread deliberately by people selling audits.

Majestic maintains two indexes. The Fresh index holds what the crawler has seen in a recent window — a rolling picture of the live web. The Historic index holds everything ever recorded, going back years. They are different datasets with different collection windows, not a before-and-after.

Here is my own site in both, captured the same day:

billhartzer.comFreshHistoric
External inbound links11,8151,382,639
Referring domains1,13022,260
Referring IPs6809,359
Referring subnets5286,684

A hundred and seventeen times as many links in one column as the other. Put those two figures side by side with no explanation and you can tell a frightening story: this site has lost 99 percent of its links. It is a story I have been shown, in a pitch, about a site that was fine.

What the gap mostly represents is a crawl window. Twenty-two thousand domains have linked to that site at some point in twenty years; 1,130 of them were seen linking during the recent window. Some of the rest genuinely stopped. Many were crawled years ago, are still live, and simply have not come round again — a site with modest traffic and few inbound links of its own gets revisited rarely, so its outbound links spend most of their existence outside the Fresh window.

Now the same pair for the abandoned domain: Fresh 182 links from 30 referring domains, Historic 31,004 links from 4,180. That profile really is dead — over 99 percent of it. The way you tell the two situations apart is the ratio, not the gap. Roughly 5 percent of my site's historic referring domains appear in Fresh; under 1 percent of the abandoned domain's do. One is a site being crawled at a normal rate. The other is a site almost nothing links to any more.

So: quote Fresh against Fresh and Historic against Historic, always. When someone shows you a decline, ask which index each number came from before you ask anything else.

Links, distinct links, and referring domains

The second way link counts mislead is duplication. Of my site's 10,727 live links in the Fresh index, Majestic classified 2,377 as distinct and 7,395 as duplicates, and its noise reduction cut the total by 56 percent. So the honest description of an 11,815-link profile is roughly 2,400 real links.

The mechanism is the sitewide placement. One person adds your link to a blogroll or a footer; the template renders it on every page; the crawler records one link per page. In the Historic index of that same profile, a single domain accounts for 219,159 links — 16.8 percent of everything ever recorded — and 218,030 of them carry an identical anchor pointing at an identical URL. One editorial decision, made once, by one person.

Which is why the reporting rule is simple and nearly universal among people who do this seriously: report referring domains. If a client wants link counts, give them distinct links and say why. A report whose headline number can be moved by one webmaster changing one template is not measuring your work.

The abandoned domain shows the same effect at small scale: 182 links, 31 distinct, 150 duplicates. And in its anchor text, one referring domain produces 123 links on the phrase search engine optimization. One footer.

What the graph cannot tell you

A reference that only flatters its tools is not a reference. Four limits, all of which I have watched people ignore:

  • Every index is one crawler's view. Two tools disagree about the same site constantly, and neither is lying. They crawled different things.
  • Missing is not absent. A link absent from an index may never have been crawled. My own audit found 20,768 live links pointing at targets the crawler could not fetch — most likely a host or firewall intermittently blocking it, which means every tool evaluating that site sees a degraded picture. That remains a hypothesis until server logs confirm it, and I state it as one.
  • Lost means no longer observed, which is not the same as removed. Crawl scheduling contributes to every loss figure you will ever read, including mine.
  • Metrics are measured now, not then. A domain that was strong in 2011 and is Trust Flow 0 today shows as Trust Flow 0, which biases any historical analysis in a direction you cannot easily quantify.

None of that makes the graph unusable. It makes it evidence rather than proof, and the difference matters most when you are writing a recommendation someone will act on.

Questions

Is the Fresh index better than the Historic index?

Neither is better; they answer different questions. Use Fresh to ask what the profile looks like now — for competitive comparison, for judging whether a link is live, for spotting recent change. Use Historic to ask what a site has ever accumulated, which matters for reclamation work and for understanding a domain's full history. The mistake is comparing a number from one against a number from the other.

How many tiers deep is it worth going?

Tier 2 is essential and tier 3 is usually informative. Tier 4 and tier 5 are mostly noise for a normal profile, though they are where a manufactured one tends to give itself away, because the fabricated layer sits further out. My habit is to record all four readings and spend the analysis time on tiers 2 and 3.

Why does my tool show far fewer links than another tool?

Because each tool has its own crawler, its own index size, its own refresh schedule and its own rules about what counts as a link. Differences of several times over are normal on the same domain. Pick one tool for tracking a profile over time, and never compare a number from one tool against a number from another as if the difference meant something happened.

Should I count links or referring domains?

Referring domains, almost always. Link counts are dominated by sitewide placements, and a sitewide link is one editorial decision repeated across a template. On my own profile a single referring domain contributed 16.8 percent of every link ever recorded. Report domains, and if you must report links, report the distinct count and explain the difference.