After this lesson you should be able to
- Choose between survey data, public datasets and your own aggregate data
- Write a methodology statement that survives a specialist's scrutiny
- Judge whether a sample supports the claim you want to make
- Newsjack without attaching your brand to the wrong story
Four sources of data you can legitimately produce
Every credible data story comes from one of four places. Knowing which one you are in tells you what the story can claim and what it cannot.
- An original survey. You commission a panel or poll your own audience and report the answers. Strong on novelty, weak on authority, and the easiest of the four to build badly.
- Analysis of public data. Government statistics, regulator filings, court records, open datasets, published research. Cheap, defensible, and the route most people underuse because it takes analytical skill rather than budget.
- Your own aggregate operational data. What your systems already record, stripped of anything identifying and reported in aggregate. Genuinely proprietary — nobody else can replicate it — which is exactly why it works.
- Newsjacking. Attaching your data, or your expert's analysis, to a story that is already running. No production cost, very short window, higher risk.
What is not on that list: making numbers up, rounding a small sample into a national claim, or presenting a directional finding as a precise one. That is not a gray area. It is the thing that ends careers and, when a specialist notices, becomes the story instead of your campaign.
Running a survey that survives scrutiny
Surveys are popular in digital PR because they are fast and you can decide the topic. They are also where most campaigns fall apart, because a survey is a measuring instrument and a badly built instrument produces confident nonsense.
Sample size and what it buys you. The relevant question is not how big your sample is, but how precise a claim it supports. Precision improves with the square root of the sample size, which means going from a few hundred to a few thousand responses helps far less than people expect, and going from a few dozen to a few hundred helps enormously. The practical consequence: small samples can support a broad directional statement, and cannot support a headline that reports a difference of a couple of percentage points as a finding.
Who you asked matters more than how many. A sample of two thousand drawn entirely from your own newsletter tells you about your newsletter, not about the country. That is fine, if you say so. The failure is describing an audience survey as a national one. If you want to describe a population, you need a sample drawn to represent that population, which is what a reputable panel provider sells and why it costs what it does.
Question wording is not neutral. Leading questions produce the answer you wanted, and any journalist who covers polling can spot one immediately. Ask about behavior rather than intention where you can, avoid double-barrelled questions, and never write a question whose only purpose is to generate a statistic you have already drafted the headline for.
Subgroups are where surveys go to die. If you asked a thousand people and then report a finding about left-handed managers in one region, the figure may rest on a dozen responses. Report subgroup findings only when the subgroup is large enough to carry them, and say what the base size was.
Disclosing methodology, and why it wins you coverage
Publish the methodology on a page you control, alongside the story, before you pitch it. It should state, at minimum: who conducted the fieldwork, when it ran, how respondents were recruited, the total sample size, the base size for every figure you quote, the exact wording of the questions you report, and any weighting applied.
Two things happen when you do this. First, serious outlets can actually use the story — many have standing rules against reporting research without a methodology, and a missing one is a silent rejection you never hear about. Second, the methodology page is itself a linkable asset. It is the page other researchers cite, and it is the page that survives long after the news cycle, which matters when you remember that most links disappear over time.
Publishing the underlying tables, in a format someone can open and check, is better still. It signals that you expect scrutiny and are not worried about it. Very few campaigns do this, which is precisely why it works.
Analyzing public data and your own product data
Public data is the most underrated route in this discipline. Statistical agencies, regulators, courts, land registries, transport authorities and health bodies publish enormous quantities of material that almost nobody reads carefully. The story is usually in the joins: two datasets nobody has combined, a series nobody has indexed to inflation or population, a national figure nobody has broken down by area.
Three cautions. Check the licensing terms before you republish anything. Read the source's own notes on revisions and definitions, because a series that changed definition mid-way will produce a spectacular finding that is entirely an artifact. And expect the agency's own press office to be a competitor for the story, because they publish too.
Your own operational data is the strongest position of all when you have enough of it, because it cannot be replicated by a competitor and it describes behavior no survey can reach. The constraints are real and non-negotiable: aggregate to the point where no individual or client is identifiable, check what your privacy policy and terms actually permit, get whatever internal approvals your legal team requires, and never publish anything a client would be upset to recognize. If in doubt, raise the aggregation level until the doubt goes away. A slightly blunter finding that is unquestionably safe is worth more than a sharp one that generates a complaint.
Newsjacking without burning the brand
Newsjacking means inserting your data or your expert into a story that is already running. Done well it is the highest return-per-hour activity in digital PR, because the reporter is already writing and already needs exactly what you have.
The mechanics are simple and the execution is hard. You need to be monitoring the beat continuously, you need pre-prepared material you can reshape within the hour, and you need a spokesperson who can be reached and quoted quickly. The window on a fast-moving story is measured in hours; by the second day the reporters have their sources and you are too late.
The risk is judgment. Attaching your brand to a tragedy, a disaster or a bereavement to sell software is the reliable way to become the story in the worst sense, and the reputational cost outlasts any link. My rule is simple: if a reasonable person could describe your intervention as profiting from someone's misfortune, do not send it. Commentary on regulation, market movements, technology changes, weather patterns, sporting events and policy debate is fair game. Commentary on somebody's worst day is not.
The other failure is having nothing to add. A quote that restates the news is not commentary. If your expert cannot say something the reporter did not already know, stay out of it.
How bad research gets you criticized rather than covered
There is a whole world of people — statisticians, academics, specialist correspondents, methodology bloggers — who read PR surveys for sport. When they find one built badly, they publish. The pattern is consistent enough to describe.
It usually starts with a headline claim that sounds too neat. Someone asks for the sample size and the question wording, and either finds no methodology page or finds one that does not support the claim. Then the base sizes turn out to be tiny, or the sample turns out to be self-selected, or the question turns out to be leading, and the write-up that follows names the brand, names the agency, and ranks well for both.
That article does not decay the way a news placement does. It sits there. When someone researches your company later, they find it. The asymmetry is stark: a well-built study earns coverage that fades from the front page within days, while a badly built one earns criticism that persists for years. Given that trade, the extra week spent getting the methodology right is the cheapest insurance in the discipline.
Questions
How large does a survey sample need to be?
It depends on the claim, not on a magic number. Precision improves with the square root of the sample, so a few hundred responses supports a broad directional statement while a small difference between two groups needs far more. The more important question is whether the sample represents the population you are describing — a large self-selected sample is weaker than a smaller representative one.
Can I survey my own customers and call it research?
Yes, provided you describe it accurately as a survey of your customers or audience. That framing is honest and often more interesting, because your customers may be a group nobody else can poll. What you cannot do is present it as representative of a wider population. Journalists who cover polling check this, and misdescribing the sample is the fastest way to lose their trust permanently.
Is newsjacking risky?
It carries reputational risk that a planned campaign does not, because you are commenting on events you do not control. Commentary on policy, markets, regulation, technology and weather is generally safe. Commentary that reads as profiting from a tragedy is not, and the damage from getting that wrong outlasts any coverage you would have earned.
Do I have to publish the raw data?
You do not have to, but publishing the tables and the full methodology makes the story usable by outlets that have rules about reporting research, and gives you a durable page that other people cite. It also signals that you expect to be checked. Very few campaigns do it, which is part of why it works.
What if my data shows something unflattering about my own industry?
That is usually the better story. Findings that cut against the sponsor's commercial interest are more credible to a journalist precisely because they are unexpected, and they are much harder to dismiss as marketing. Suppressing the inconvenient finding and publishing only the flattering one is also the pattern that specialists look for.