AdvancedLinkTraining.com logo — a free link building course by Bill HartzerAdvanced Link TrainingA resource by Hartzer.com

Reading the link graph

Detecting link networks and footprints

Seven domains with identical metrics are not seven decisions. They are one operation that forgot to vary something.

Lesson 14 of 50Module 3 · Reading the link graph7 min read

After this lesson you should be able to

  • Identify a link network from a shared metric signature across supposedly independent domains
  • Test infrastructure, registration and publishing signals without over-reading any single one
  • Judge when a neighborhood or subnet check is meaningful and when it is not
  • Decide what to do about a network depending on which tier it sits in

A footprint is a repeated decision

Running a link network means running many sites while pretending they are unrelated, and that is expensive in attention. Every site needs hosting, a registration, a template, a plausible name, content and a publishing rhythm — six decisions that must be made differently every time for the pretense to hold.

Nobody sustains that. Operators automate, and automation repeats. A footprint is simply the place where the repetition became visible. The forensic task is not to find one damning fact; it is to notice that a set of supposedly independent domains agree about something they had no reason to agree about.

This lesson leans on my investigations work as much as on SEO, and the standard of proof is the same: no single indicator is conclusive, and several weak indicators pointing the same way, across a population rather than an individual, is how you establish anything.

Two habits before the tests below. Always evaluate a set, because a footprint is a property of a group and every test here produces false positives applied to one domain. And hold onto the innocent explanation for as long as it survives — shared hosting is normal, common templates are normal, and agencies register domains in bulk for clients quite legitimately.

The metric signature: the test that works best

The strongest single test I know is also the cheapest. Scan a tier for domains whose Trust Flow and Citation Flow are almost identical to each other.

Here is a run from my own site's tier 2:

DomainTrust FlowCitation FlowCF minus TF
aligow.com365014
zomatt.com354813
mantaw.com344814
zlutag.com344814
dotpim.com334815
yelpad.com334815
ylutag.com334815

Seven domains. Trust Flow spans four points, Citation Flow spans two, and the gap between them is 13 to 15 in every single row. The names are nonsense five-and-six-letter brandables of the kind a script generates in batches.

Why does this work? Flow metrics are a function of a domain's inbound links, so for seven domains to land in the same narrow band they must have been linked in the same way, at the same time, from the same places. Independent sites with different histories do not converge on the same two numbers. Identical metrics mean a shared link source, and a shared link source means one operation.

Then the finding that made me write this lesson: the same domains appear in the manufactured profile's tier 4zlutag.com, zomatt.com, mantaw.com, aligow.com, yelpad.com, dotpim.com, ylutag.com, plus yellii.com, gigbic.com and fduty.com in the same band. Two sites with no relationship to each other, an active professional site and a long-abandoned SEO domain, and the same operation is attached to both at different depths. That is what a network at scale looks like from the outside: it is not aimed at you, it is aimed at everything.

The manufactured profile's tier 2 shows the same test on a different band — a Dutch airport taxi service, an NFL merchandise store, a pharmacy spam site, stand mixers, dolls and acrylic products, all clustered at Trust Flow 32 to 33. Six unrelated subjects, one number.

Infrastructure, registration and rhythm

The metric signature tells you a network exists. These tests help you bound it, and each carries a caveat that matters more than the test.

  • Shared subnets. Sites in one network often sit on adjacent IPs because they were provisioned together. Read the ratio, not the list: 1,130 referring domains across 680 IPs and 528 Class C subnets is roughly 1.7 domains per subnet, which is unremarkable. Twenty supposedly independent referrers on one subnet is not. Caveat: shared hosting puts thousands of unrelated sites on one IP.
  • Shared registration details. Same registrar, same nameservers, same creation date to the day, same renewal cycle, same privacy service. Caveat: privacy services are now the default and prove nothing on their own; a bulk creation date across a themed set is the stronger version.
  • Shared analytics and advertising identifiers. The same tracking or monetization ID in the source of several sites is one of the few near-conclusive links between properties, because an operator who wants aggregate revenue reporting has a reason to keep them together. Caveat: agencies legitimately share containers across client sites.
  • CMS and template fingerprints. Same platform, same theme, same plugin set, same image sizes, same category structure, same boilerplate in the privacy page. Caveat: popular themes are popular.
  • Publication cadence. Post dates clustered into a short burst then nothing, or a metronomic every-third-day rhythm with no seasonality, is machine output. Real publishing has holidays in it.
  • Outbound link behavior. A page with 606 outbound links to 525 external domains and no internal links is not a website; it is a links page. Real editorial content links inward as well as outward.

Weight these by how expensive they are to fake. Metrics and outbound-link behavior are hard to disguise because they are consequences of how the network operates. Registration privacy is trivial to fix, and any operator worth worrying about fixed it years ago.

The neighborhood check, and why it usually fails now

Conventional advice says to check your hosting neighborhood: look up the sites on your IP address and worry if they are disreputable. Majestic's Neighbourhood Checker does exactly that. Behind a content delivery network, the result is close to meaningless, and most sites are now behind one.

My own site resolves to 104.26.7.89, a Cloudflare shared address. Its listed neighbors include an adult site and a gambling site. A naive audit writes that up as a bad neighborhood. What it actually reflects is that a large number of unrelated sites use the same CDN, which says nothing about hosting quality, editorial standards or anybody's association with anybody.

It gets worse, and the tool is honest about it. On a second site, Majestic warned that the domain had another IP worth checking — CDN-fronted domains usually do. The first address did not list the site at all; the second, 172.67.68.181, showed it at position 11 with Trust Flow 19, Citation Flow 39 and 114 referring domains. One lookup would have given a confidently wrong answer.

So the working rules are: resolve every IP the domain answers on and check them all; if any of them belongs to a CDN or a large shared host, discard the neighbor list entirely rather than interpreting it; and note that a small site may not appear in its own IP's top-ranked list at all, which means nothing except that it is small.

Where the underlying idea still holds is in aggregate, on the referrers rather than on the target. Referring domains per subnet is a real measure of concentration, because it is computed across your whole profile rather than from one lookup. A profile spread thinly across many subnets was accumulated from many sources. One concentrated into a handful was not.

Anchor text as a footprint

Anchor text is where a manufactured profile most often gives itself away, because the whole point of buying a link is usually the anchor.

An earned profile's top anchors are boring. On mine the leaders are an empty anchor — image and bare links, 419 referring domains — then billhartzer, bill hartzer, billhartzer.com, several article titles, a visit site and a read more. No commercial exact-match phrase appears in the top fifteen. Even the misspelling billhatzer.com shows up, which is the kind of untidiness that only occurs naturally.

The manufactured profile's top anchors are a raw URL, then #1 website promotion and seo, then #1 website promotion & seo consultant, then search engine optimization. No brand anchors at all, because there is no brand — only the phrase somebody wanted to rank for, attached to a domain that is also that phrase.

One row in that table deserves its own note: the search engine optimization anchor comes from a single referring domain producing 123 links. That is one footer, counted 123 times, and it is why Lesson 3.1 insists you report referring domains. In an anchor table, always read the referring-domain column before the link column.

What to do when you find one

The answer depends entirely on where the network sits, and this is where most audits go wrong.

In your own tier 1 — the network links directly to you. Either you or somebody working for you placed those links, or someone pointed them at you unprompted; establish which. If they were placed, they are a liability whether or not they currently work, because the footprint that let you find them in ten minutes is available to a search engine with vastly more data. If they arrived unsolicited, this happens to everyone, and search engines have spent twenty years learning to discount rather than punish.

In your tier 2 or deeper — nothing. You did not place them, you cannot remove them, and there is nothing to disavow because they do not link to you. The finding is still worth recording, for one reason: it changes how you value your tier 1. A referring domain whose own inbound links come mostly from a network is a domain whose metrics are borrowed rather than earned, and you should discount it accordingly when scoring the profile.

In a prospect's profile, before you buy. This is where the lesson pays for itself. Ten minutes with the tier slider and the metric-signature scan tells you whether a site's authority is real or manufactured, and manufactured authority is what you are being sold when a marketplace listing has a good number and a bad story.

In a competitor's profile. Interesting, occasionally satisfying, rarely actionable. Understanding that their headline numbers are assembled rather than earned, and adjusting your own targets accordingly, is the useful part.

Questions

How many matching signals do I need before calling something a network?

My working standard is two independent signal types across a population of at least five domains. A metric signature plus shared infrastructure, or a metric signature plus identical publication cadence, is enough to act on commercially. One signal on one domain is never enough, because every test in this lesson produces false positives when applied to an individual.

Is a shared IP address proof that sites are related?

No, and it is weaker evidence now than it has ever been. Shared hosting puts thousands of unrelated sites on one address, and a CDN puts millions behind one. A real site of mine shares its Cloudflare IP with an adult site and a gambling site. Use IP data in aggregate — referring domains per subnet across a whole profile — rather than as a verdict on one lookup.

What if my own site is linked from a network I did not build?

It almost certainly is, and so is everyone's. There is an identifiable network in my own site's tier 2, on a profile where nothing was ever bought. Unsolicited links from networks are the normal condition of having existed online. Take no action, disavow nothing, and record it only as context for how much weight the affected referrers deserve.

Can I use these techniques to check a link vendor before buying?

Yes, and it is the highest-value use of the lesson. Ask the vendor for example placements, load each domain in the Link Graph, walk the tiers, and scan for a shared metric signature across the sample. If several supposedly unrelated sites share a band, you are being offered one network sold as a portfolio, and its footprint is as visible to a search engine as it was to you.