After this lesson you should be able to
- Check registrant, hosting and DNS records for shared ownership across candidate sites
- Find shared analytics, advertising and affiliate IDs in page source
- Read publication cadence, author bylines and about pages as editorial signals
- Recognize a cluster of sites with implausibly similar metrics as a single network
What a footprint is and why they exist
A footprint is a detail repeated across sites that are supposed to be unrelated. Networks exist because someone wants to control many linking sites at once. Controlling many sites at once requires standardization — the same registrar, the same host, the same theme, the same publishing routine, the same analytics account, the same content supplier — and every standardization is a footprint.
This is the part of link analysis I do most often in investigative and expert witness work, and the principle transfers directly from that domain: people are consistent, and consistency is evidence. Nobody running forty sites configures each one differently by hand. They cannot afford to. So the forty sites share something, and your job is to find what.
Two framing notes before the checks. First, a single shared attribute proves nothing — plenty of unrelated sites share a host or a WordPress theme. What matters is convergence: several independent signals agreeing across the same set of domains. Second, you are not trying to prove a case to a court standard. You are deciding whether to spend money. A well-founded suspicion is a sufficient reason to decline.
Ownership and registration signals
Start with who owns the domain, because ownership is the hardest thing to disguise completely.
- Registrant details. Most WHOIS records are privacy-protected now, which reduces but does not eliminate the value of the check. Look at the registrar, the privacy service used, and the creation date. Twelve sites registered at the same unusual registrar behind the same privacy provider within the same six-week window is a pattern.
- Historical WHOIS. Privacy is often added later, or lapses at renewal. Historical records frequently expose the registrant who set the domain up.
- Creation and expiry dates. Network domains are often expired domains re-registered to inherit an existing backlink profile. A domain created in 2009, with a gap in its archive between 2014 and 2022, and content that resumed on a completely different topic, was bought for its links.
- Name server configuration. Custom name servers on an obscure host, shared across a set of sites, are a strong link.
- Mail records. MX records pointing at the same self-hosted mail server across supposedly independent publications are hard to explain innocently.
Historical archives are worth the time. Look at what the domain published five years ago. A site presenting itself as an established home improvement publication that was a dentist's practice site until eighteen months ago is not what it claims.
Infrastructure and code signals
Now read the page source. This is where operators are least careful, because these details are invisible in a browser.
- Analytics IDs. A Google Analytics or Tag Manager identifier is account-specific. Two sites sharing one are administered by the same account. This is the single most conclusive footprint available, and it is visible in the raw HTML.
- Advertising and affiliate IDs. AdSense publisher IDs, affiliate network IDs, and ad partner tags all identify an account rather than a site.
- Verification tags. Search Console and Bing verification meta tags occasionally remain in the template.
- Hosting and IP. Shared IP addresses mean little on their own — shared hosting is normal. Shared IP on a small or unusual provider, combined with anything else on this page, means a great deal.
- CMS and template fingerprints. The same theme is common. The same theme with the same custom CSS file, the same modified footer markup, the same plugin set, and the same unusual widget arrangement is one person's build, deployed repeatedly.
- Structural tells. Identical URL patterns, identical image naming conventions, identical sitemap structures, identical category taxonomies across sites in different niches.
Editorial and human signals
Infrastructure signals prove common administration. Editorial signals tell you whether a real publication exists at all, and they are what you check when the infrastructure is clean.
Publication cadence
Real publications publish irregularly. People take holidays, get busy, respond to news. A site that has published exactly four posts a month, on the same days, for two years, is running a content schedule, not a publication. The opposite pattern is equally telling: eighty posts in one week eighteen months ago and nothing since, which is a domain that was stocked with content and left to age.
Authors
Look at the bylines. Are there names? Do those names have bios? Do the bios link to anything that exists — a LinkedIn profile, a personal site, a portfolio, other publications? Can you find the person anywhere else on the internet? Sites with no bylines, or with bylines like "Admin" and "Editor", or with a stock-photo author who appears on several unrelated sites, are not staffed. A single author name covering posts on tax law, dog grooming and cryptocurrency is not a polymath.
The about page
A thin About page is one of the most reliable tells in this list. Real organizations have a history, a location, staff, and something to say about why they exist. Network sites have three sentences about a passion for sharing quality information. Check whether there is a physical address, whether the address is real, whether the contact form goes anywhere, and whether the company is registered anywhere it claims to be.
Outbound link pattern
Read who else the site links to. A genuine publication links out to sources, competitors, and references, mostly to well-known sites, mostly with descriptive anchors. A network site's outbound commercial links go to unrelated businesses with exact-match anchors, one per article, always in the third paragraph.
A worked example: seven domains, one operator
Here is a case that illustrates convergence better than any list of checks. Seven domains came up as link prospects, presented as separate, independently owned sites in a handful of adjacent niches. Individually, each looked passable. Placed side by side, they did not survive five minutes.
The first thing that stood out was the metrics. All seven sat in a band of Trust Flow 33 to 36 and Citation Flow 48 to 50. That is not a coincidence; that is a target. Metrics on genuinely independent sites scatter, because they are the residue of unrelated histories. Seven unrelated sites landing inside a three-point Trust Flow band and a two-point Citation Flow band means one person was building to a specification — enough links to clear a buyer's filter, no more, because more costs money.
The second thing was the names. All seven were invented five-letter brandable domains — pronounceable nonsense of the kind sold in bulk by brandable-name marketplaces. Real businesses in these niches are named after their founders, their location, or their trade. Seven independent publishers do not all arrive at the same naming fashion at the same time.
The third was the ratio itself. Citation Flow near 50 against Trust Flow in the mid-30s is a profile with substantially more citation volume than trust behind it — consistent with links acquired in quantity rather than earned.
From there the rest fell out quickly: overlapping registration windows, the same theme with the same footer structure, a metronomic posting schedule, About pages that could have been swapped without anyone noticing, and no author who existed anywhere else on the web. One operator, seven domains, and a price list.
The transferable lesson is the metric band. When several prospects cluster tightly on metrics, stop evaluating them individually and start comparing them to each other. Independence produces scatter. Uniformity is a manufacturing signature.
What footprints prove, and what to do about it
Be precise about what you have found, because overreach here leads to bad decisions in both directions.
A shared analytics ID proves common administration. A shared host proves nothing by itself. A tight metric band across supposedly unrelated sites proves nothing on its own but is very hard to explain innocently. Convergence across ownership, infrastructure, and editorial signals is as close to conclusive as this work gets.
What you do with the finding depends on which side of the transaction you are on. As a buyer, decline — and decline the vendor, not just the site, because a vendor who offered you one network site has more. As someone auditing an inherited profile, document what you found before you touch anything; a disavow decision needs a record behind it. As someone assessing a competitor, understand that identifying a network does not entitle you to conclude anything about what Google will do with it.
Finally, keep your standards honest. A small independent blog with one author, an ugly theme, and irregular posting is not a network. It is a small independent blog, and it is often exactly the kind of relevant link this module has been arguing for. The signals here identify manufactured scale, not amateurism.
Questions
Is shared hosting on its own enough to reject a site?
No. Millions of legitimate sites share IP addresses on mainstream hosts, and rejecting on that basis alone would eliminate most small publishers. Shared hosting becomes meaningful only in combination — the same small provider plus overlapping registration dates plus an identical template plus no traceable authors.
How do I find shared analytics or AdSense IDs?
View the page source and search it for the identifier patterns used by analytics and ad platforms, then compare across sites. Several reverse-lookup services index these identifiers and will show other sites carrying the same one. It is the most conclusive single footprint because the ID belongs to an account, not a domain.
What if the network sites are genuinely well written?
Content quality is not the issue. A network is defined by common control of supposedly independent sites for the purpose of passing links, and good writing does not change the pattern the network leaves in the link graph. It only means the operator invested more, which usually means the price is higher.
Can I use footprint analysis on my own existing backlinks?
Yes, and it is one of the more useful audits available. Sort your referring domains by metric similarity, look for clusters, then check ownership and infrastructure signals within each cluster. Inherited profiles and past agency work frequently contain network links nobody currently on the team knew about.
Do private blog networks still work?
They work until they are found, and the economics of finding them have moved steadily in the search engines' favor. Every cost saving that makes a network profitable — shared hosting, shared templates, shared content sourcing, shared ownership — is also the signal that identifies it. If I can find the pattern from the outside in an afternoon, so can a system with the full crawl.