← All articles

Do clicks rank pages? What the DOJ trial and the 2024 leak do and do not show

An evidence review of click data in Google ranking: what Google says, what the US v. Google court found from sworn testimony about Navboost and Glue, what the 2024 API documentation leak only names, why none of it supports CTR manipulation, and what a site owner can legitimately do about clicks.

LIPAI WANG ·

"Google uses clicks to rank pages" became received wisdom in SEO after two events: testimony in the US antitrust trial against Google, and the leak of Google's internal API documentation in May 2024. A whole sales pitch grew from it. If clicks rank pages, the argument goes, then buying clicks ranks pages.

This article reviews the evidence. It does not teach click manipulation and does not describe how anyone does it. It separates four kinds of claims that get blurred together:

  1. what Google documents about its own systems,
  2. what a federal court found, based on sworn testimony and trial exhibits,
  3. what the leaked documents name without explaining,
  4. what practitioners infer from all three.

Every claim below carries one of four labels: official guidance (Google documenting its own product), research finding (here, mostly the court record), our observation, or hypothesis. I treat the leaked documentation as a research finding about which fields existed, and as nothing more. All sources were checked on 2026-10-06. The court case is still on appeal, so check its status before you rely on this after early 2027.

The short answer

  • Clicks feed some Google ranking systems, in aggregate. Google says so in plain terms (official guidance). The court record describes one such system, Navboost, in some detail (research finding).
  • Nothing public gives weights. Not the court opinions, not Google's pages, not the leak. Nobody outside Google can say how much click data moves any given result.
  • Nothing public says that inflating your own clicks raises your rankings. That is an inference, and the record gives good reasons to doubt it.
  • Google's spam policies treat attempts to manipulate its systems, and automated queries to Google, as spam (official guidance). There is no separately named "click manipulation" policy. The general definitions already cover it.
  • The legitimate lever is the same as it always was: a title and snippet that describe the page accurately for the query, and a page that actually settles the searcher's question.

Evidence at a glance

Claim Google documents it Court found it (testimony and exhibits) Leak names it Status
Google uses aggregated, anonymized interaction data to judge relevance Yes Yes n/a Established
A system called Navboost memorizes which results users click for a query No Yes Yes (module names) Established, via testimony
Navboost is trained on about 13 months of data (18 months before 2017) No Yes No Established, via testimony
Glue is a wider log of queries, results shown and interactions on the results page No Yes Yes (names) Established, via testimony
Specific click fields such as "goodClicks" and "lastLongestClicks" exist No No Yes, names only Fields existed; meaning and use unknown
The weight of click data in ranking No No No Unknown
"Pogo-sticking" or dwell time on your page is a ranking factor No Not as such (see below) No definition Unproven
Buying or faking clicks durably raises a page's rank No No No Unproven; against spam policies

What Google says itself

Google's "How Search works" page on ranking results says it uses aggregated and anonymized interaction data to assess whether results are relevant to queries (official guidance, How Search works, register M1-19). It does not name the data, the systems that use it, or how much weight it carries. It does not mention dwell time, bounce rate or "pogo-sticking".

The more technical ranking systems guide (page dated 2025-12-10) describes many named systems and does not mention click or interaction data at all (official guidance, CL-11). Silence is not denial. It just means that page is no help either way.

Google staff have also spoken off the documentation. In a 2019 Reddit AMA, Gary Illyes of Google dismissed theories that dwell time and CTR are ranking signals as largely made up, according to Search Engine Roundtable's report (official guidance, but secondhand and informal, CL-12). In the same thread he described RankBrain as using historical data about what happened on the results page, not on the landing page. Read together, his comments push back on the idea that Google measures how long someone stays on your site. They don't say Google ignores what people click on the results page. The court record, discussed next, is about the results page.

What the court found

In August 2024, Judge Amit Mehta ruled that Google had illegally maintained a monopoly in general search (memorandum opinion, Doc. 1033, 2024-08-05). Several findings of fact deal with user data, because the case turned partly on whether Google's scale gave it an advantage rivals could not match. Findings of fact are a court's conclusions from testimony and exhibits under oath and cross-examination. They are the strongest public evidence we have, but they describe what witnesses said about systems at the time of trial (2023). They are not technical specifications (research finding throughout this section).

What "click data" means in the record. The court lists the kinds of click data a search engine can collect: which results a user clicks, whether the user returns to the results page and how quickly, how long the user hovers over results, and scrolling on the results page. It says a search engine learns from this about relevance and about the quality of pages users visit (finding 88, citing a Google exhibit and testimony by Eric Lehman, described in the opinion as a former Google distinguished software engineer; CL-01). Note what this is: a list of data types, not a description of a formula. "Returns to the results page and how quickly" is the closest the record comes to the idea practitioners call pogo-sticking. The opinion never says how, or whether, any ranking system uses that particular measure.

Navboost. The court describes Navboost as a signal that pairs queries and documents by memorizing user click data. It lets Google remember which documents users clicked for a query and spot when one document gets clicks across several queries. Google trained it on 18 months of data before 2017 and on 13 months since (finding 96, citing testimony by Lehman and by Pandu Nayak, Google's VP of Search; register M1-20). John Giannandrea, a former head of Google Search, testified that Navboost was considered very important.

Newer systems. RankBrain, DeepRank, RankEmbed and similar "generalization" systems rely less on user data and are meant to fill gaps where click data is thin. They were still designed with user data and trained on it, at smaller volumes. Newer language models did not replace Navboost (findings 97–98 and 102; CL-02). Sridhar Ramaswamy, a former Google Search executive who went on to found the rival engine Neeva, testified that working out which pages are most relevant still benefits a great deal from query and click information (finding 115; CL-02).

Abuse detection. In a passage about privacy, the court notes testimony that Google logs IP addresses partly to detect and fight botnets and fraudulent clicks (finding 123; CL-03). The context there is security and advertising, not organic ranking. It tells you Google watches for fake clicks. It doesn't tell you how organic systems treat them.

The 2025 remedies opinion: Glue

In September 2025 the same court decided what Google must do about it (remedies memorandum opinion, Doc. 1436, 2025-09-02; CL-04). Part of the remedy requires Google to share some user-interaction data with qualified competitors. To decide that, the court described the data:

  • Glue is a "super query log" (an expert witness's phrase). It records the query (text, language, location, device), what was shown (the ten blue links and other features such as maps, images and People also ask), interactions on the results page (clicks, hovers, time on the results page), and query interpretation such as spelling corrections.
  • Navboost data is part of Glue. Nayak testified at trial that Glue is essentially Navboost extended to everything else on the page. Navboost aggregates click-and-query data about web results and can be thought of as a giant table.
  • Some signals are simple counts. The court mentions raw signals such as how many times a page was clicked for a particular query, with Navboost as the example.
  • The remedy covers the data, not the models. Google does not have to hand over the signals or models it builds from Glue.

The final judgment followed in December 2025. Google appealed in January 2026 and the government cross-appealed in February 2026, according to Alphabet's Q2 2026 quarterly report (CL-05). The appeal could change the remedies. It is unlikely to change what witnesses said about how Navboost and Glue worked.

What the 2024 leak shows

In spring 2024, internal reference documentation for Google's Content Warehouse API was, according to the analysts who first reported it, published to a public code repository by mistake. Rand Fishkin and Mike King published the first analyses on 27–28 May 2024 (SparkToro, iPullRank). On 29 May, Google confirmed to The Verge that the documents were authentic. A spokesperson cautioned against assumptions based on "out-of-context, outdated, or incomplete information" (The Verge, CL-06). The Verge's own report noted that the documents don't show which data is actually used in ranking or how elements are weighted.

What the documents are matters. API reference documentation lists data structures: modules and the named fields inside them. You can see this in a public mirror of the package. One Navboost-related module lists fields named goodClicks, badClicks, lastLongestClicks, unsquashedClicks and impressions. Most of them have no description at all, only a type (number) (research finding, mirror of the leaked package, CL-07).

So the leak tells you that Google stored click counts under those names. It does not tell you:

  • what counts as a "good" or "bad" click,
  • whether a field is filled, used in ranking, used only for evaluation, or abandoned,
  • any weight, threshold or formula,
  • whether a field applies to organic results, other features or experiments.

Both lead analysts said as much. Fishkin warned readers not to treat any single field as proof of a ranking factor. King noted that the documents contain no scoring functions (CL-08). Their readings of what fields mean are informed guesses. I label them hypothesis, which is how they framed them.

The leak's value is that it is consistent with the trial record. Navboost exists, it handles click data, and Google records more than raw clicks. It adds field names. It adds no mechanism.

Where the sources agree, and where they stop

Taken together, the documented and testified record supports these statements:

  • Google collects interaction data on its results pages, aggregates it across users, and has used it in ranking for many years (M1-19, M1-20, CL-01, CL-04).
  • One important system, Navboost, works by memorizing which results get clicked for which queries, over roughly 13 months (M1-20, CL-04).
  • Newer machine-learned systems were trained on click data too, at smaller volumes (CL-02).
  • Google watches for fraudulent clicks, at least in its security and ads systems (CL-03).

And it does not support these:

  • "CTR is a ranking factor you can optimize like a dial." No source gives a weight or a response curve.
  • "Dwell time on your site is measured and rewarded." The court's list of click data is about behavior on the results page. "Duration on the SERP" is in the Glue description. Time on your own page is not described anywhere in the record. The leak's lastLongestClicks name suggests some notion of a long click, but its definition isn't public.
  • "Pogo-sticking triggers a penalty." The record names returns to the results page as a type of data. It never describes a penalty.
  • "Chrome data ranks your site." The leak contains Chrome-related field names. Neither court opinion I read says Chrome browsing data feeds organic ranking, and I found no Google statement that it does (hypothesis at best).

Why "CTR manipulation" claims don't follow

Some practitioners argue that because clicks are an input, buying clicks, running click bots or paying crowds to search and click must raise rankings. Here is why that inference fails, even before the policy question.

1. "X is an input" does not mean "more X wins". Any system that learns from behavior also has to handle noise, bias and abuse. Inputs can be normalized, capped, compared to what's expected at that position, or discarded when they look unnatural. The leak's "squashed" and "unsquashed" field names hint that some normalization happens, but the public record doesn't define it, so I won't claim more than that (hypothesis).

2. The data is aggregated over a lot of real traffic. The court describes Navboost as trained on 13 months of queries and clicks from all users, and Google's daily query volume as about nine times that of all its rivals combined (findings 87 and 96, M1-20; the remedies opinion adds that the 13 months cover queries and clicks from all users, CL-04). A manipulation campaign has to stand out against that baseline without looking unnatural. The record does not say how Google treats such patterns. It does say Google logs IP addresses partly to catch botnets and fraudulent clicks (CL-03). Whether a given campaign "works" is unknowable from public evidence (hypothesis).

3. The demonstrations aren't controlled. Public "click test" stories are usually single before/after observations on one query, with no control query, no repeated runs and no account of concurrent updates. A before/after change alone does not show cause. Reported effects are often described as short-lived, which is also what you'd expect from normal ranking fluctuation (hypothesis; register M1-24 on test design).

4. It is against Google's rules. Google's spam policies define spam as techniques that deceive users or manipulate Search systems into featuring content prominently, including attempts to manipulate AI responses in Search. They separately list machine-generated traffic, meaning automated queries to Google sent without permission, as a violation of both the spam policies and the Google Terms of Service (official guidance, spam policies, page dated 2026-08-28, CL-09). Google's Terms of Service also forbid using its services in fraudulent or deceptive ways (official guidance, Terms of Service, effective 2026-07-30, CL-10). Violations can lead to lower rankings or removal (register AG-04).

5. It damages your own measurement. Fake visits pollute Search Console clicks, analytics sessions and conversion rates. You lose the one dataset that tells you whether your pages work for real buyers. For a SaaS or B2B site the goal is qualified demand, activation and retention. Fake clicks produce none of those.

This article stops there. I won't describe methods, vendors or "safe" volumes, because the honest answer to "how do I do this safely" is that the evidence doesn't show it works and the policy says don't.

What a site owner can legitimately do

If real users' choices on the results page feed some ranking systems, the legitimate implication is modest and familiar: earn the click honestly, then deserve it.

Make the title and snippet match the query's intent. Google builds title links automatically from the title element, headings and other page text, and may replace a title that is vague, stuffed or inaccurate (official guidance, title links, M1-08). Snippets come mainly from page content, and the meta description is used when it describes the page better (official guidance, snippets, M1-07). Write one accurate, specific title per page. Make the first paragraph say what the page answers. Don't promise something the page doesn't deliver: a misleading title might win a click, but it loses the visitor.

Check the live results before you write. If the query returns comparison lists and your page is a product page, or the reverse, your page may be the wrong kind of answer for that query (our method, hypothesis). Our BOFU pages article covers this for SaaS buyer queries.

Settle the searcher's question on the page. Google's people-first guidance asks whether a visitor leaves feeling they learned enough to reach their goal (official guidance, helpful content, M1-09). Put the answer near the top, give specifics a reader can check, and make the next step obvious.

Measure CTR by query and position, against your own baseline. In Search Console, CTR depends heavily on position and on what else is on the results page. Compare a page with its own history at a similar average position, not with an industry curve. Filtered totals undercount because rare queries are anonymized (official guidance, Performance report, M1-15). Our page-2 CTR case study shows one way to read a collapse.

Test title and snippet changes like an experiment. Change a set of pages, hold a similar set back as a control, and compare CTR at matched positions over equal windows. Note any announced Google updates. Small sites often can't detect small effects; if yours can't, report what you saw as an observation (our method, hypothesis, CL-14).

Build real demand. People who already know your brand search for it and pick it on the results page. That is a result of product and marketing work, not a ranking trick. Treat branded search as an outcome to measure, not a lever to pull (our reading, hypothesis).

How to read the next "clicks rank pages" claim

When someone cites the trial or the leak, ask:

  1. Which source, exactly? A court finding (with its number), a Google page, a leaked field name, or someone's interpretation?
  2. Is it about the results page or about your site? The record is about results-page behavior.
  3. Does it give a weight or a mechanism? If it does, where does the weight come from? No public source has one.
  4. Is the evidence for "it works" controlled? A control set, repeated runs, a stated effect size and a duration.
  5. Does the advice require deceiving Google or users? If yes, it is outside Google's spam policies, whatever the evidence says.

If you want a second opinion on what is actually moving your search traffic, RankPropel's audit and baseline sprint starts from your Search Console and conversion data, not from ranking-factor lists. The method above works the same if you run it yourself.

Sources

US District Court for the District of Columbia, United States v. Google LLC, No. 1:20-cv-03010-APM: memorandum opinion on liability (Doc. 1033, 2024-08-05; findings of fact 86–98, 102, 115, 123) and memorandum opinion on remedies (Doc. 1436, 2025-09-02; data-sharing remedies section), both via CourtListener. Alphabet Inc., Form 10-Q for the quarter ended 2026-06-30 (appeal status). Google: How Search works: ranking results, ranking systems guide, spam policies (page dated 2026-08-28), Terms of Service, title links, snippets, creating helpful, people-first content, Search Console Performance report. Google's statement on the leak: The Verge, 2024-05-29. Leaked documentation: public mirror of the Content Warehouse API package. First analyses of the leak: SparkToro and iPullRank (both May 2024). Gary Illyes's 2019 AMA comments as reported by Search Engine Roundtable. All retrieved 2026-10-06.