After this lesson you should be able to
- Define co-citation and distinguish it from a direct link
- Read a Related Sites list and filter string-similarity false positives out of it
- Use Clique Hunter to test whether two sites share a referrer population
- Interpret a null result as evidence rather than as a failed query
Co-citation: being mentioned in the same company
Co-citation is being cited alongside something else. If a page links to you and to three competitors, the four of you have been co-cited, whether or not you link to each other. Co-occurrence is the weaker relative: appearing near the same brands in text without any link at all.
The idea is borrowed from bibliometrics, where two papers cited together repeatedly are understood to belong to the same field even if neither cites the other. The web version is intuitively similar. If the pages that reference you also reference the recognized names in your subject, you are in that subject's neighborhood. If they reference article directories and unrelated e-commerce, you are in a different one.
How much weight search engines put on this is not something anyone outside them can state. What I can say is that co-citation is a good diagnostic regardless of how it is weighted: it describes the company a domain keeps, and the company a domain keeps is an accurate summary of how it acquired its links. It is also the part of link analysis most obviously relevant to how generative systems assemble an answer, since those systems work from associations between entities rather than from a link graph alone. That connection is real but not quantified, and I will not pretend otherwise.
Reading a Related Sites list
Majestic's Related Sites report returns the domains that most often appear near links to your target, with a context count for each. It is a co-citation report by another name, and the contrast between two of them is the clearest single diagnostic in this module.
For my own site, the top rows are ycombinator.com, seroundtable.com, hartzer.com, domainnamewire.com, forbes.com, searchengineland.com, x.com, prnewswire.com, searchenginejournal.com, medium.com, crn.com and techcrunch.com. Trade press, general technology press, wire services and my own second property. Whatever else you conclude, that is the neighborhood of a site in the search and domain-name industry.
For the abandoned domain, the list reads 1st-in-articles.com, 1stoparticles.com, 25000articles.com, 2articles.com, 365articles.com, momarticles.com, morganarticlearchive.com, affiliatesdirectory.com, adswise.com and web-visibility.net. Several have no home page title at all and several are Trust Flow 0. It is an article-directory farm, and it is a complete account of how that profile was built without needing to look at a single individual link.
The reading habit: look at the naming pattern first, the Trust Flow column second, and the context count third. A neighborhood of near-identical names built from one keyword is a neighborhood of one operation.
The false positives nobody warns you about
Rows 11 to 13 and 16 to 17 of my own Related Sites list are billhartzell.com, billhatton.ca, billhau.co.uk, billhauck.com and billhattier.com. Trust Flow 8 to 10, Citation Flow 1 to 3, classified under Home and Family, and utterly unconnected to anything I do.
They are string-similarity matches on the billha… prefix. Every relatedness algorithm blends several signals, and lexical similarity is one of them because it is cheap and usually helpful — sites with similar names often are related. Here it produces five confident wrong answers ranked among genuine industry peers.
I have never seen this documented, and it matters for two reasons. If you are auditing your own neighborhood, five junk rows in a top-twenty list distort your impression of it. And if you are using Related Sites for prospecting — which is a legitimate use — you will waste outreach on sites that were never relevant.
The filter is manual and takes a minute. Read every row and ask whether the name resembles the target's name more than the site resembles the target's subject. If it does, and the metrics are low, and the topical classification is unrelated, strike it. Anything you would not have recognized as a peer before running the report deserves the check.
Clique Hunter, and what a null result teaches
Clique Hunter takes several domains and returns the referring domains they share. It answers a question you cannot answer from any single profile: who links to all of these? The standard use is competitive — feed it three or four competitors and the output is a prospect list of sites already demonstrably willing to link within your subject.
The second use is forensic, and it is what I want to show you here. Run it on two sites you suspect are connected. Genuine independence produces a small, boring overlap; a real relationship produces a large one, or a small one made of specific sites nobody would land on twice by chance.
I ran the abandoned SEO domain against hartzer.com. The result was one shared referring domain, and it was blogspot.com. That is a null result, in its purest form. Blogspot is a free blogging platform with millions of hosts; sharing it as a referrer is like sharing a postcode with a stranger. Two sites in the same industry, sharing exactly one referrer, and it is the one that everybody shares.
Against my other site the overlap was four: blogspot.com again, plus a free article directory, a free-content site and a decision-support aggregator. Also nothing — legacy links of the kind that land on any site that has been discussing search for two decades. And a detail worth catching: blogspot sends 222 links to my site and 9 to the other one. Same source, wildly different footprint, which is a reminder that shared-referrer counts are as important as shared-referrer identity.
Null results are worth publishing, and almost nobody does, because a tool that only shows you positives teaches you to see patterns everywhere. Before you conclude that two sites are connected, run the query on two you know are not, and look at what nothing looks like.
Mutual links and link context
Two smaller reports finish the picture at link level.
Mutual Links lists the links running between two specific domains, in both directions, with page-level metrics on each end. Running it between two properties I own turned up something useful and slightly embarrassing: of the top links from one to the other, two were flagged deleted, with last-seen dates a couple of months apart — a buying brand mentions anchor last seen in early June and a seo expert witness anchor last seen at the end of April. I own both ends of both links. Nobody noticed until the report said so.
Another row is an image link, where the anchor is derived from alt text. Image links are easy to forget in an anchor audit and they can be a substantial share of a profile — the largest single anchor group on my site is the empty anchor, 419 referring domains, which is mostly images and bare URLs.
Link Context shows the surrounding text, the anchor in place, the count of outbound links on the page split internal and external, and a rough indication of where on the page the link sits. This is where you answer the question that started the module. A link inside a paragraph on a page with seven external links is editorial. A link on a page with 606 outbound links to 525 external domains is a directory entry. Both are one row in a backlink report.
Putting co-citation to work
Four uses, in the order I actually reach for them:
- Neighborhood diagnosis. Run Related Sites on your own domain first. If the list does not look like your industry, the profile has been assembled from sources that have nothing to do with your subject, and that is a strategy finding rather than a link finding.
- Prospecting. Clique Hunter across three or four competitors returns sites that link to all of them. They link within your subject, they link to more than one player, and they have therefore already answered the hardest question in outreach.
- Relationship testing. Use it to check whether two properties are connected — in due diligence, in an investigation, or when a vendor's portfolio looks suspiciously coherent.
- Reclamation. Sites that co-cite you but do not link are the warmest prospects available, and Module 8 shows why: the overwhelming majority of lost links were lost while the page stayed live and maintained.
One caution to close on. Co-citation describes association, not endorsement. Appearing alongside strong brands on a scraped list of company names is not the same as appearing alongside them in an editor's roundup, and no tool distinguishes the two for you. Look at the pages.
Questions
What is the difference between co-citation and a backlink?
A backlink is a link pointing at you. Co-citation is a page referencing you and something else, whether or not links are involved. You can be co-cited with a competitor by a page that links to neither of you. The practical difference is that backlinks are about value transfer, while co-citation is about association — what subject and what company a search engine, or a language model, sees you in.
Why does Related Sites return domains with similar names to mine?
Because lexical similarity is one input to the relatedness calculation and it misfires on distinctive name prefixes. My own report returned five unrelated sites matched on the first six letters of my surname, ranked among genuine industry peers. Filter the list by hand: if a row resembles your name more than it resembles your subject, and its metrics are low, strike it.
What does it mean if Clique Hunter returns almost nothing?
It means the sites genuinely do not share a referrer population, which is often exactly what you wanted to know. Two sites in the same industry sharing one referring domain, and that domain being a free blogging platform, is strong evidence of no relationship. Run the tool on two sites you know are unconnected first, so you recognize a null result when you see one.
Can I use Related Sites as a prospecting list?
Yes, with filtering. The genuine rows are sites that already publish near your subject, which makes them better qualified than most cold lists. Strip the string-similarity false positives, strip anything with no home page title, and check what the remaining domains actually publish before pitching. Clique Hunter across several competitors usually produces a better list for outreach specifically.