AdvancedLinkTraining.com logo — a free link building course by Bill HartzerAdvanced Link TrainingA resource by Hartzer.com

EarnedSafe

Original research and statistics

Publishing a number nobody else has is the most dependable way to earn links at scale - and the easiest way to get publicly corrected.

Verdict: The most reliable link-earning tactic there is, provided the methodology would survive a statistician reading it; a weak study earns criticism instead of coverage.

What original research is, in link terms

Original research means producing a number that did not exist before you produced it, and publishing it in a form other people can cite. That is the whole mechanism. Writers need figures; figures need a source; the source gets the link. Everything else in this tactic is execution detail.

It takes several forms, and they differ enormously in cost and defensibility:

  • Proprietary data analysis - counting something inside a dataset only you hold. Usually the strongest option, because it is genuinely unrepeatable.
  • Survey research - asking a defined population a set of questions. The most common form and by far the most frequently botched.
  • Public dataset analysis - taking government or regulator data and producing a cut nobody has published. Cheap, credible, and underused.
  • Experimental or observational studies - measuring what happens under defined conditions and reporting it.
  • Freedom-of-information requests - obtaining records that were not previously public. Slow, but journalists find the provenance compelling.

The reason this tactic sits at the top of the reliability ranking is that a good statistic keeps earning. A campaign asset that ran three years ago is dead; a cited figure gets picked up by every writer who searches for it, for as long as it is the best available number. That long tail is the actual return, and it is why the tactic justifies its cost.

How to run a study, step by step

  1. Write the headline finding first, as a hypothesis. If the sentence would not interest a journalist when true, do not run the study. Note that you must be equally willing to publish the boring result.
  2. Decide what would make the finding wrong. Doing this before you collect data is what separates research from marketing. Write down the confounders and the alternative explanations.
  3. Choose the method that fits the question. A survey measures what people say. Behavioral data measures what people do. Confusing the two is the single most common error in published marketing research.
  4. Define the population and the sample honestly. Who exactly are you generalizing to, how did you reach them, and who could not possibly have been reached? Write this down before fieldwork, not after.
  5. Collect the data. Neutral question wording, no leading options, no forced choices that manufacture the answer you wanted.
  6. Analyze conservatively. Report the effect size, not just the direction. Do not slice the sample until a subgroup produces something interesting; a subgroup of 40 people is not a finding.
  7. Publish the methodology in full, on your own site. Sample size, sampling method, fieldwork dates, exact question wording, panel or data source, and known limitations. Make the raw or aggregate data downloadable if you can.
  8. Package for citation. A canonical URL, a clear title, the key figures as plain text near the top, one or two charts that read correctly at small size, and a suggested citation line.
  9. Pitch it as a story, not as a report. The journalist needs the finding, the number, the caveat, and a quotable human.

What it costs in time and effort

Cost is driven by four things, in descending order of importance. Data acquisition comes first: analyzing data you already own costs analyst time only, while commissioning survey fieldwork costs whatever a research panel charges, and that scales steeply as the target population narrows. Reaching a general adult population is routine; reaching qualified anesthetists or fleet managers is a different order of expense.

Second is analytical skill. Somebody has to know what a confidence interval means, why a self-selected sample cannot be generalized, and when a difference is noise. Buying this in is far cheaper than publishing a study that a specialist takes apart in public.

Third is packaging - the write-up, the charts, the methodology page. Modest, but not zero, and skimping here wastes the money spent on the first two.

Fourth is promotion. A study nobody pitches earns almost nothing. Budget roughly as much attention to distribution as to production; the most common way this tactic fails is a genuinely good study published quietly and never sent to anyone.

On the calendar, eight to sixteen weeks from decision to first coverage is normal: two to four for design and fieldwork, two for analysis and write-up, and the rest for outreach and the slow accumulation of citations afterward.

When it works and when it does not

It works when you have access to data other people want and cannot get, when the subject has journalists covering it, and when you are willing to publish a result that does not flatter your product. It works particularly well for sites whose sector generates numbers as a by-product of operating - marketplaces, payment processors, job boards, logistics firms, anyone sitting on transaction data.

It does not work when:

  • The finding is already known. Confirming the obvious earns nothing, however rigorous the method.
  • The result must be favorable. If the study is only publishable when it supports your marketing claim, you are not doing research, and it will read that way.
  • Nobody covers the subject. No press means no citations, however good the data.
  • You cannot defend the method. A weak study in a technical field will be picked apart by practitioners, and that criticism outranks your study.
  • You need speed. This is a quarter-long project at minimum.

The failure mode worth naming explicitly: a badly built survey does not get ignored, it gets criticized. A poll of 200 self-selected newsletter subscribers, presented as "what consumers think", invites a rebuttal post from someone with a statistics background - and that rebuttal earns the links your study was supposed to earn, while permanently attaching your brand to it. The downside of this tactic is not zero return; it is negative return.

Common mistakes

  • Hiding the sample size. If the number is small, say so plainly and narrow the claim to match. Journalists who cannot find the sample size assume the worst, correctly.
  • Surveying your own customers and calling them consumers. Your customer base is not a random sample of anything, and someone will point that out.
  • Leading questions. Wording that produces the answer you wanted invalidates the result, and publishing the question wording is what makes it visible.
  • Torturing subgroups. Cutting a 1,000-person sample into twenty segments guarantees a spurious "finding" in one of them.
  • Confusing correlation with cause. "Firms that do X grow faster" is not evidence that X causes growth, and writing the headline as if it were is how a study gets discredited.
  • No methodology page. The single strongest signal that a study was built to be cited rather than to be true is a full methodology, published before anyone asks.
  • Publishing and stopping. The citations that matter accumulate over years, which means the study needs to stay reachable at a stable URL, get refreshed with new fieldwork, and get re-pitched whenever the subject is in the news.
  • Rounding in your own favor. Precision that flatters is the tell that the whole thing was reverse-engineered.

A worked example

The link decay study behind this site is a straightforward example of the tactic, and it is worth walking through because it shows what makes a study citable.

The data. A complete Majestic export for one domain - 1,301,839 links from 22,260 referring domains, on a profile where no link was ever bought. It is proprietary in the only sense that matters: nobody else can produce it, because nobody else has that history.

The finding. Links decay far faster than the industry assumes. The median referring domain stops linking after 1,080 days, and 49.6% of all links ever recorded are gone. The surprising part - the part that makes it a story rather than a statistic - is that 97.4% of those losses happened while the source page was still perfectly reachable. Only 4 of 645,202 lost links were lost to a 404. Pages are not disappearing; editors are removing links from pages that still exist.

The secondary findings. Survival scales with Trust Flow: 859 days at Trust Flow 0 against 3,353 days at Trust Flow 61 and above, roughly four times the lifetime. And 69.3% of the referring domains are Trust Flow 0 on a profile where nothing was bought, which is a useful corrective for anyone auditing a link profile for quality.

Why it is citable. The population is defined, the size is stated, the source of the data is named, the limitation is obvious and disclosed - it is one site, in one sector, and therefore not automatically generalizable. Anyone writing about link decay needs a number, and this is one with its working shown. That is the entire bar.

How to measure it

Judge a study on a longer horizon than a campaign, because its return arrives differently. Track referring domains to the study URL as the headline number, segmented by Trust Flow so you can tell genuine coverage from aggregator noise. Then track citation velocity - new referring domains per quarter - which for a good study should decline slowly rather than stop, because writers keep finding it.

Track unlinked citations separately. Statistics get quoted without attribution constantly, and each instance is a reclamation opportunity with an unusually high success rate, because the writer has a professional obligation to source a number.

Two further measures are worth having. Query capture: whether the study page ranks for the statistic itself, since that is what drives the long tail. And secondary reuse: whether other people's articles, decks and reports repeat your figure, which is the clearest evidence the number has entered general circulation.

Set the denominator honestly. On a twenty-year natural profile, only 446 referring domains reached Trust Flow 41 or above - about 22 a year. A single study that adds ten or fifteen strong domains has outperformed a normal year of organic accumulation, and comparing it against a target of "200 links" invented in a planning meeting will make a genuine success look like a failure.

The verdict

If you can only run one earned tactic, run this one. It is the most dependable way to earn links at scale, the only tactic whose output keeps earning years after the work is done, and the natural raw material for digital PR, expert commentary and everything else in this category.

The condition is that it has to be real research. The version of this tactic that fails is the version built backward: a conclusion chosen for marketing reasons, a survey engineered to produce it, a sample size buried, and a methodology page that does not exist. That version does not merely underperform - it invites public correction from people who know the subject, and their correction earns the links.

Do the smallest honest study you can afford rather than the largest impressive-sounding one you cannot defend. Publish the method in full. State the limitations before anyone else does. Then keep the page alive, refresh the data, and let it accumulate citations for the next five years, which is where the real return has always been.

Questions

How large does a survey sample need to be?

Large enough to support the claim you want to make, which depends on how narrow the claim is. A national consumer statement needs a properly sampled population in the high hundreds or above; a statement about a specialist profession may be defensible with far fewer if the sampling is sound. The rule that matters is simpler: publish the number and narrow the claim to fit it.

Can I use publicly available data instead of running my own study?

Yes, and it is badly underused. Government, regulator and open-data releases are full of cuts nobody has published. The original contribution is the analysis and the framing, not the collection. It costs a fraction of survey fieldwork and journalists trust the underlying source immediately.

What if the results do not support our marketing message?

Publish them anyway, or do not run the study. A finding that contradicts your own commercial interest is the most credible thing you can publish and tends to get more coverage, not less. If your organization cannot accept that possibility, choose a different tactic rather than producing research you will have to bend.

How long does a study keep earning links?

Years, if it remains the best available number and the URL stays stable. That long tail is the main return. It ends when someone publishes a newer, better study, which is a good argument for refreshing the fieldwork periodically rather than treating the study as finished.

Do I need a statistician?

You need someone who understands sampling, effect sizes and the limits of what your data can support - whether that is a hired statistician, an analyst, or an academic reviewer. Buying a few hours of that review is far cheaper than publishing a study that a specialist dismantles publicly.