AdvancedLinkTraining.com logo — a free link building course by Bill HartzerAdvanced Link TrainingA resource by Hartzer.com

Foundations: how links actually work

PageRank, Penguin, and what changed

Both things are true at once: links remain a core ranking input, and building them the wrong way still gets sites hurt.

Lesson 2 of 50Module 1 · Foundations: how links actually work5 min read

After this lesson you should be able to

  • Describe PageRank in plain language without the mathematics
  • Explain what Penguin changed about the economics of link buying
  • Reconcile the claim that links still work with the claim that schemes get caught
  • Identify which pre-Penguin habits are still circulating as advice

The original idea

PageRank was described publicly in the late 1990s by Larry Page and Sergey Brin while they were at Stanford, in the paper that also described the search engine that became Google. It is named after Page. The idea is simple enough to state in one sentence: a page is important if important pages link to it.

The definition is circular, and that is the clever part. You cannot compute a page's importance without already knowing the importance of the pages linking to it, so you compute the whole thing iteratively — assign every page an equal starting value, distribute each page's value across its outbound links, repeat, and the numbers converge. The model is often described as a random surfer: someone clicking links at random forever, occasionally jumping to a random page instead. A page's PageRank is the probability of finding that surfer there.

Two properties made this a genuine advance over what came before. It used information from outside the page being ranked, so it was harder to fake than keyword stuffing. And it was recursive, so a link from a page that was itself well-cited counted for more than a link from an unknown page. Every subsequent link-based system inherits both properties.

What PageRank was not, even at the start, is a complete ranking system. It was one component, computed independently of any query, combined with relevance signals at search time. Anyone describing PageRank as "how Google ranks pages" is describing 1998 badly.

The economy that grew around a number

For years Google published a coarse PageRank score in its browser toolbar — a 0-to-10 integer, updated occasionally. It was intended as a curiosity. It became a currency.

Once a visible number existed, a market formed around raising it. Link buying became an industry with rate cards keyed to toolbar PageRank. Article directories, paid blog networks, sitewide footer links, reciprocal link pages, forum signature spam and automated link software all existed to accumulate a metric. Google eventually stopped updating the public toolbar score and later removed it entirely, but by then the habits were set — and a great deal of link building advice still in circulation was written for that era.

It is worth being honest about why it worked: it did work, for a while, and spectacularly. That is exactly why the correction, when it came, was severe. Anyone who tells you link schemes never worked was not there.

What Penguin changed in practice

Penguin launched in 2012 as an algorithmic response to link spam and other manipulation. Panda, the year before, had targeted thin and low-value content; Penguin targeted the link graph. Google later folded Penguin into its core algorithm and made it run continuously rather than as periodic refreshes, and said it became more granular — devaluing spam rather than only demoting whole sites.

I will not give you a version-by-version history, because most of what circulates about specific refreshes is reconstructed from ranking trackers rather than from anything Google said. What matters is what changed in practice.

Before

  • Volume was the strategy. More links, more or less regardless of source, tended to help.
  • Exact-match anchor text at high ratios worked and was widely recommended.
  • Cheap networks and directories were a normal line item in an SEO budget.
  • The downside risk of a bad link was close to zero.

After

  • Links from patterns the system recognizes as manipulative are, at best, ignored.
  • Aggressive anchor text distributions became a liability instead of a lever.
  • Recovery from an algorithmic suppression became slow and uncertain — the work is cleanup, then waiting, with no ticket to file.
  • Google introduced a disavow tool, which is the clearest possible statement that some links are considered harmful rather than merely useless.

The two mechanisms are worth separating. An algorithmic effect happens automatically and silently; nobody tells you. A manual action is a human reviewer at Google applying a penalty, and it appears in Search Console with a reconsideration process attached. Most sites damaged by links were never manually penalised. They simply stopped getting credit for the thing that had been holding them up, which feels identical from the outside and is diagnosed completely differently.

Why both claims are true

"Links still work" and "link schemes get caught" are usually presented as opposing positions. They are not. They describe different things.

Links still work because the underlying idea has not been replaced. Google representatives have repeatedly described links as an important signal, and every competitive commercial search result is still populated by sites with substantial link profiles. Nothing has emerged that does the job a citation graph does — an independent, external, hard-to-self-declare indicator that other people find a document worth pointing at.

Schemes get caught because the systems that assess links got much better at recognizing patterns. Manipulated links tend to leave footprints: the same handful of hosting arrangements, the same templates, the same anchor text, the same link velocity, the same neighborhoods of sites linking to each other in ways no organic ecosystem produces. You are not trying to fool a formula. You are trying to fool a classifier that has seen every network anyone has ever built, alongside a spam team that buys the same services you do in order to map them.

The honest synthesis: the returns on genuine citations have held up, and the returns on manufactured ones have collapsed while their downside has grown. That is not a moral position, it is an economic one. Manufactured links now cost more, last less time, and carry a tail risk that can remove a business's main acquisition channel.

What this means for how you work now

Five practical consequences.

  • Stop optimizing for a number nobody publishes. There is no public PageRank. Third-party scores are models built by tool vendors from their own crawls. They are useful for comparison and useless as targets in themselves.
  • Assume every pattern you can execute at scale is a pattern that can be detected at scale. If a tactic is cheap and repeatable for you, it is cheap and repeatable for ten thousand other people, and that volume is exactly what makes a footprint visible.
  • Judge a link by whether a person would defend the placement. Not "would it pass a filter" but "if a Google engineer read this page, would the link look like an editor's decision?"
  • Treat anchor text as something that happens to you, not something you dictate. Real citations produce messy anchors: brand names, bare URLs, "this article", "according to". Clean, keyword-rich distributions look designed because they are.
  • Discount advice that predates the correction. A surprising amount of published link building material is a decade old and was written for an environment that no longer exists. Check when the technique was invented before you invest a quarter in it.

Questions

Does Google still use PageRank?

Google has said publicly that it still uses PageRank as one of many signals, while making clear that the modern implementation is not the 1998 formula and that the public toolbar score is gone for good. Treat PageRank as a family of link-based computations inside a much larger ranking system rather than as a score you can see, buy, or directly target.

Can I still get a penalty for bad links in 2026?

Yes. Manual actions for unnatural links still exist and still appear in Search Console. The more common outcome, though, is silent devaluation: the links simply stop counting and rankings drift down with no notification. That is harder to diagnose than a penalty because there is nothing to appeal and no confirmation that you have fixed the problem.

Should I disavow links I did not build?

Usually not. Google has said for years that its systems ignore the vast majority of spam links automatically, and the disavow tool can do harm if you feed it domains that were actually helping. Reserve it for cases where you have a manual action, or where you know a previous owner or agency bought links at scale. Module 8 covers the decision properly.

Was Penguin a one-time event?

No. It launched as a periodic filter, then Google incorporated it into the core algorithm so that it runs continuously. In practice that means there is no longer a refresh date to wait for. Cleanup work is reassessed as pages are recrawled, which is why recovery timelines vary from weeks to many months depending on how quickly the affected links get revisited.