Google Ranking Factors: What Site Owners Aren't Told

Головні фактори

Google has never published a complete list of its ranking factors. What people usually call “the 200 factors” is Brian Dean’s compilation at Backlinko: some items are confirmed by Google, some are SEO observations, and some are guesswork. For a site owner, a list like that makes it hard to tell what actually matters.

That changed in 2023–2024. First, Google executives explained under oath in a US antitrust trial how some of its ranking systems work. Then Google’s internal Search documentation leaked publicly. For the first time, it became possible to compare what Google says in public with what is written in its own documents.

This article was first published in 2019 as an addition to the Backlinko list. We have rewritten it from scratch: what Google says officially, what came out of the trial and the leak, which 2019 advice no longer works, and what all of this means for a site owner.


1. What Google says officially

In its How Search Works explainer, Google names five groups of factors: the meaning of the query, the relevance of the page, the quality of the content, the usability of the page, and the user’s context (location, language, settings). It does not list specific signals or their weights.

More detail is in the guide to ranking systems. Google describes systems rather than factors — algorithms that each handle part of the evaluation. There are 17 of them today. These are the ones most relevant to a typical business website:

System What it does
BERT, MUM, neural matching Understand what the query and the page mean, not just whether the words match
RankBrain Connects words to concepts and helps with queries the system has not seen before
Link analysis and PageRank Evaluate who links to a page and how
Freshness systems Surface newer content where recency matters for the query
Original content system Favors the original source over rewrites of it
Reviews system Rewards reviews written from first-hand experience
Passage ranking Finds the specific passage on a long page that answers the query
Site diversity system Limits how many results from one domain appear at the top
Spam detection systems (including SpamBrain) Demote or remove pages that break the rules

 

The same guide has a section on retired systems. Panda became part of the core algorithm in 2015, Penguin in 2016, and Hummingbird evolved into later systems. The helpful content system, launched as a separate update in 2022, became part of the core ranking systems in March 2024. There is no separate "helpfulness filter" anymore: content quality is assessed continuously, in every core update.


2. What came out of the US v. Google trial

In the US Department of Justice case against Google (filed in 2020), the company was accused of monopolizing the search market. To explain why scale gives Google an advantage, the court had to look at how Search uses data about its users.

The key testimony came in October 2023 from Pandu Nayak, Google’s Vice President of Search. He confirmed the existence of a system called NavBoost and described it as one of the important signals. NavBoost has been running since around 2005 and looks at how people interact with search results: what they click, whether they come back to the results page, and which result they settle on. It uses the last 13 months of click data and helps narrow tens of thousands of candidate documents down to a few hundred early in the ranking process.

This matters for SEO because Google representatives had spent years publicly downplaying clicks as too "noisy" a signal. The trial showed that clicks do count — not as a CTR counter for a single page, but as months of accumulated data on what satisfies people for a given query.

In August 2024 the court ruled that Google is a monopolist, and on 2 September 2025 Judge Amit Mehta issued the remedies decision: Google must share some of its search index and user interaction data with qualified competitors. The ruling does not change ranking itself, but it confirms once more how valuable user behavior data is to search.


3. The 2024 Search documentation leak

3.1. What these documents are and how to read them

On 13 March 2024, internal documentation for the Google Content Warehouse API — a reference of the data structures Search systems work with — ended up in a public GitHub repository. In May 2024 it was analyzed and published by Rand Fishkin (SparkToro) and Mike King (iPullRank). It covers 2,596 modules and 14,014 attributes.

On 29 May 2024, Google confirmed that the documents were authentic, but warned against drawing conclusions from information that is out of context, outdated or incomplete.

That caveat deserves to be taken seriously:

  • the documentation describes what data Google stores, but not whether each attribute is used in ranking or how much weight it carries;
  • some attributes may already have been outdated or experimental at the time of the leak;
  • there are no formulas or coefficients, so no one can build a "ranking" of factors from it.

Still, the documents do show what Google is able to measure. And several things the company had publicly denied for years are in there.

3.2. What was found in the documentation

Attribute What it means What Google said publicly
siteAuthority An authority score for the site as a whole, not just for an individual page Google representatives repeatedly said they have no "overall domain authority"
goodClicks, badClicks, lastLongestClicks NavBoost data: good and bad clicks, and which result the user settled on Clicks and time on page were called unreliable signals not used for ranking
hostAge The age of the host; its description says it is used to sandbox fresh spam The existence of a "sandbox" for new sites was denied
chromeInTotal The number of site views according to Chrome browser data Google had previously said Chrome data is not used for organic ranking
siteFocusScore, siteRadius How focused a site is on one topic, and how far an individual page strays from it Not commented on publicly
titlematchScore How well a page’s title matches the query Consistent with official guidance
isAuthor and other author fields Google stores information about who wrote a document E-E-A-T was described as a concept from the rater guidelines, not a separate factor
bylineDate, syntacticDate, semanticDate Three different dates: the one shown on the page, the one in the URL or title, and the one inferred from the content Consistent with the advice not to change a date without really updating the content

 

The main takeaway from the leak is not any single attribute. It is that Google evaluates the site as a whole and how people behave in the search results far more than it had admitted in public. A page does not rank in a vacuum: it is affected by the domain’s authority, the site’s topical focus, and how users have responded to that site over months.


4. 2019 advice that no longer works

The first version of this article included points that almost every SEO blog repeated at the time. Some are now outdated, and some were wrong from the start.

What was advised What is actually true
AMP pages help mobile rankings and get a ⚡ badge in results Since 2021, AMP is not required even for Top Stories, and the badge is gone. Page speed matters, not the technology
The Mobilegeddon update boosts mobile-friendly sites The move to mobile-first indexing is complete: since 5 July 2024, Google crawls sites with its mobile crawler only. A site that does not work on a smartphone can drop out of the index
Content hidden in tabs on mobile is not counted With mobile-first indexing, content in tabs and accordions is fully counted — it is a normal way to save screen space
Speed: under 3 s on desktop, under 2 s on mobile Google looks at Core Web Vitals: LCP under 2.5 s, INP under 200 ms, CLS under 0.1. INP replaced FID in 2024
Sites with social media profiles rank higher; likes and shares are a factor Social signals are not a direct factor. Social media can bring traffic and brand mentions, but not rankings on their own
Articles should be at least 500 words Word count is not a factor. A page needs to fully answer the query — sometimes that takes 200 words, sometimes 3,000
The more pages, the more Google likes the site Mass-producing pages to rank is a policy violation (scaled content abuse). Weak pages drag down the whole site
Bounce rate and time on site from analytics affect rankings Google does not take data from your Google Analytics. It sees behavior in its own results — through NavBoost
E-A-T Since December 2022 it is E-E-A-T: Experience, the author’s first-hand experience, was added
The Fred, Payday Loans and Hummingbird algorithms Names of individual updates have lost practical meaning: Google assesses quality continuously, and what counts as spam is set out in official policies

 

Google Doodles and search "easter eggs", which were also in the first version, never affected rankings.


5. What this means for a site owner

Put the official documentation, the trial testimony and the leak together, and a few practical conclusions emerge.

5.1. A satisfied user matters more than a position

NavBoost has been collecting data for years on whether people found what they were looking for. A page that users keep bouncing back from to the results will gradually lose rankings, even if it is technically perfect. That is why the title and meta description should describe the page honestly, and the page itself should answer the query straight away rather than after three screens of introduction.

Faking clicks or behavior with bots is a bad idea: 13 months of data smooth out short spikes, and artificial traffic falls under the policies on machine-generated traffic.

5.2. Authority is earned by the site, not the page

Scoring the site as a whole means that a new page on an authoritative domain starts from a better position than a perfect page on a new one. Authority comes from quality links from relevant sources, brand mentions and branded searches. Weak, outdated or duplicate pages lower the overall score, so pruning content sometimes does more than publishing new content.

5.3. Topical focus

A site that writes about everything looks less expert to Google than a site focused on its niche. If a furniture store’s blog publishes horoscopes to chase traffic, it dilutes the site’s topic. A logical site structure, where sections support each other, is one of the simplest ways to show focus.

5.4. Authors and first-hand experience

Google stores author data, and the reviews and helpful content systems reward material written from experience. Author pages, real case studies, and your own photos and data instead of stock images are not a formality. Google states plainly that E-E-A-T is not a ranking factor in itself, but its systems look for signals that align with it.

5.5. Technical foundation

Technical parameters rarely lift a site on their own, but technical problems can stop it completely. The minimum to keep under control:

  • pages are indexed, and service pages and duplicates are kept out of the index;
  • the mobile version has the same content as the desktop version;
  • Core Web Vitals are in the green at least on key page templates;
  • crawl errors and manual actions are monitored in Google Search Console.

At the same time, Google says explicitly that good Core Web Vitals scores do not guarantee top rankings, and a relevant page with an average experience can outrank a fast but irrelevant one.

5.6. Dates, freshness and patience

Google distinguishes between several dates on a page and can notice when the date has changed but the content has not. Update the date when you have genuinely updated the material, as we did with this article.

Don’t judge a new site by its first few months: host age is taken into account, and Google treats a sudden start from a new domain with caution. It is one more reason to be skeptical of promises of quick top rankings.


6. What Google considers spam

Instead of naming individual algorithms, Google now sets out its spam policies. In March 2024 it added three new policies that affect many sites:

  • scaled content abuse — many pages created mainly to manipulate rankings, whether a person or AI wrote them;
  • expired domain abuse — buying an old domain with a history to host low-value content on it;
  • site reputation abuse — third-party content published on an authoritative site to benefit from its rankings, not to serve its audience.

Among the types of link spam, Google specifically names links from low-quality directories and bookmark sites, links embedded in widgets distributed across other sites, and paid articles that pass ranking credit. So some points from the 2019 list are still relevant — they are just officially documented now.


7. In short

  • No one has a complete list of ranking factors. Officially, Google describes 17 ranking systems, not factors with weights.
  • The US trial confirmed that clicks in search results count: the NavBoost system uses 13 months of data.
  • The 2024 documentation leak showed that Google evaluates the authority of the site as a whole, its topical focus, host age and Chrome data, even though this had been publicly denied. The documents contain no attribute weights.
  • AMP, social signals, “at least 500 words” and “more pages” are not ranking factors.
  • The practical conclusion is simple: a page should answer the query better than competitors, the site should be focused and technically accessible, and authority should be earned, not faked.

If you need to find out what is holding your site back, start with an SEO audit. For ongoing work, see our SEO services.

Get in touch!

Get More Value

You will get the best content from us to help your business grow.
Your request has been sent.

    Subscribe

    Advance with us — apply now

    Get a free consultation and business promotion assessment
    Your request has been sent.

      Send