Google Patents On Website Design, Layout & User Behaviour

There's a question that comes up a lot in SEO circles and it rarely gets answered with anything other than opinion, i.e. does Google actually look at how your page is designed, how it's laid out, the order you present information in, and does it use real user behaviour to work out whether that presentation is any good or not?

In the SEO space, you'll see one crowd confidently claiming that UX is a ranking factor, usually with nothing to back it up beyond a vague nod towards Core Web Vitals, and then you'll see the opposite crowd saying Google can't see your design at all so stop worrying about it. Neither position is particularly well evidenced, so rather than add another opinion to the pile I went and read the actual patent filings on Google Patents to see what's genuinely on record, and the picture that emerges is far more interesting than either camp is letting on because the capability is very clearly all there, it's been sat there for over twenty years, but the way the pieces fit together isn't what most people assume.

Now, before I go into this further, it's my belief that Google uses website design, layout & behaviour to assess:

  1. Is the website design and layout more condusive to a positive user experience

  2. Is the utilised design & colour scheme aligned with other sites where performance data is positive or negative?

  3. Is the information organisation optimal based on aggregated user behaviour signals over time?

  4. Is the coverage of design, layout and information poor, fair, good, excellent?

  5. Is the document likely to be useful and of benefit to the end user?

What's REALLY key to understand is that design, layout/ux, content and behaviour are ALL independent BUT, it's likely with RankBrain & Helpful Content that Google correlates aggregated data when weighting a document. There's a lot more that goes along with this such as Navboost, Click performance, Site quality score, authority and various trust indicators.

If you think about it - website design, layouts and UX all have a direct correlation on how they impact user behaviour - so it's not that a badly designed or badly laid out site can't rank, but it's more about the "aggregation" of positive signals together.

We know Google uses behavioural data in document ranking, and we know that Google has given some weighting to CRuX / Page speed for ranking as these impact user experience - the same can be said for what your content looks like, how its presented, your use of design and visual elements.

Thinking Outside the Box

Think of it like this:

  • You could write the worlds most amazing article, but if it's presented in a page as a giant slab of text with no real formatting, segregation, use of media etc. it's far more likely to have POORER engagement metrics / negative click signals - so irrespective of HOW good the content is, the visual presentation of it alone (design, layout) could become a "negative factor"

  • Now, imagine Google has lots of aggregated ranking data for URLS that compete with your article where it has behavioural data mappable to interaction by design - that gives Google an incredible amount of data to work with - which content actually gets more engagement and why

  • Google could then ultimately rank something by not only the quality of the information but the presentation of that information


It could look something like this ?

Now - the reason I bring this up is because back in 2019 I was telling people in the SEO community that Google was 100% using user behaviour as a ranking factor, I got called conspiracy theorist, told that Google wouldn't / couldn't do that etc.

But I knew Google was - how? well firstly I looked at Chrome in more detail and I ran my own test with my wholesale bulk food website that doesn't actually sell any products yet I got it to rank.

You see to get the data, Google has the best way to do it that didn't require websites to have google search console or google analytics installed, it had CHROME! Which had a majority browser market share (along with other browsers which sit atop chromium i.e. Opera).

Google could collect user behaviour data and telemetry information via Chrome:

Google called it "making the web better" by providing data via the browser - we see this now as "make searches and browsing better". And this is how Google could easily get lots of useful data to help them build a picture of what people found useful / helpful.

According to the DOJ trial - Google admitted - we don't understand documents, we fake it

And written clearly there - we watch how people react to documents and memorize their responses, well how do they do that?

Simple, via Chrome.

Now, I don't believe Google when it says we do not understand documents - a lot has changed in the last few years, give Gemini a document and it'll be able to understand what the document is, albeit not in true intelligence form (As LLMS work on prediction architecture).

But - the very thing I had harped on about for years was confirmed!

So, if Google is using user behaviour data, we know clearly that there's going to be sub-sets of things that tie in with this thus DESIGN, LAYOUT, UX, Content Information Priority etc.

What Google has always been open about

The Chrome User Experience Report is the thing that was never hidden, it's a public dataset that reflects how real-world Chrome users experience popular destinations on the web, it's the Google dataset of the Web Vitals programme, and all the user-centric Core Web Vitals metrics are represented in it.

The data is collected from users who have opted into syncing their browsing history and have usage statistics enabled, and Google uses CrUX site speed data internally to inform search ranking signals alongside other factors like page quality, relevance and mobile friendliness.

Now here's the part that's directly relevant to my analysis on the patents, because one of the three Core Web Vitals is Cumulative Layout Shift, which is literally a measurement of how much the layout of the page moves around during loading, i.e. the unwanted jumps that frustrate and confuse users. So there is already a layout stability metric being collected from real people's machines, aggregated, and fed into ranking, and Google has never denied any of that.

How much weighting is given here - we'll never know, but what you SHOULD know is that:

  1. Google has a massive reason to incentivise people making their websites lighter and faster - because it's CHEAPER for Google to crawl, analyse, process - therefore if every webpage could lose a few KB in speed optimisation, that would amount to thousands of petabytes of data savings for Google

  2. User experience is impacted not only by design and layout but how that design and layout is rendered - with CLS and slow loading pages being the thing that frustrates a lot of users browsing the web - missed tap targets because of shifting design can lead to mis-clicking or sub-navigating without meaning to

What Google denied, and what the leak showed

The behavioural side is where it gets muddier, because Google spent years publicly downplaying it. John Mueller & others at Google have stated repeatedly in 2021 that he didn't think Google used anything from Chrome for ranking, acknowledging only Chrome User Experience Report data for page experience signals.

Then in 2024 the Content Warehouse API documentation went up on GitHub by accident, over 2,500 pages describing more than 14,000 attributes, and Google confirmed the documents were authentic on the 29th of May 2024.

I wrote an article on NavBoost Unpacked - what the google content warehouse leak actually tells us about click based ranking and I also put together an entire copy of that data warehouse leak here.

The leak was considered significant precisely because parts of it appeared to contradict public statements Google representatives had made about the algorithm, particularly around the use of click data and Chrome browser data in rankings. The specific field everyone latched onto was chromeInTotal, described as total views or visits to the site from Chrome, and it sits in a module related to page quality score, assessing site-level Chrome views.

That doesn't tell you how heavily it's weighted or whether it's even live, and anyone claiming otherwise is going beyond the evidence, but it does confirm that Chrome-derived data exists inside the ranking infrastructure rather than being confined to a public performance report.

As I wrote above, I knew Google would use user behaviour data, my tests ratified my beliefs. In 2019 > 2020 I built a wholesale food website on Wordpress and basically created a full dummy eCommerce website which you can see at bulkco.co.uk >

Using a mix of SEO techniques I was able to get the website ranking for HUNDREDS of wholesale food keywords from wholesale chocolate to wholesale alcohol.

I did it via a process of product stacking + behavioural manipulation using CTR Booster + VPS + residential proxies to basically simulate website traffic, browsing and engagement. Because all the requests were routed through residential proxies, to Google it saw what looked like genuine user behaviour from lots of different IP's with lots of varied and emulated behaviour.

I tested this for MONTHS until the rankings came - and we ranked for a lot!

By 2021/2022 we were getting 800-1000 visitors per day, I then switched off CTR booster as rankings had climbed.

Then - to test the user behaviour aspect, I allowed users to spend time browsing the site adding to the cart, but, unbeknownst to them they could never check out (it wasn't until they'd spent 5, 10 sometimes 15 minutes going through the website adding to their cart that they realised they couldn't actually checkout).

But, by then it was too late, the user had browsed and spent time on the site.

So, I then disabled the ability for people to add any products to the cart:

Engagement performance absoluely tanked with average session metrics going down from 7 minutes 50 seconds with 7 pages per session down to 40 seconds and 3 pages per session.

And following that, the rankings absolutely tanked.

I restarted CTR Booster and traffic manipulation and rankings came back again, I tested this on other sites too.

Testing gave me confidence that Google was using the behavioural data in ranking assessments.

WHAT A PATENT ACTUALLY TELLS US, AND WHAT IT DOESN'T

Before getting into any of this it's worth being clear about what we're looking at, because the SEO industry has a genuinely bad habit of finding a patent abstract, screenshotting it and declaring the thing a confirmed ranking factor, which isn't how any of this works and it damages the credibility of anyone doing this properly.

A patent tells you what a company thought was worth protecting at a given moment in time, it doesn't tell you whether the thing was ever built, whether it deployed, whether it deployed and then got switched off in 2011, or whether it's running right now in some heavily modified form.

Google itself puts a disclaimer on every single patent page stating that the legal status shown is an assumption rather than a legal conclusion and that the assignee list may be inaccurate, so if Google won't vouch for the metadata on its own patent search product then you probably shouldn't be building a client seo strategy off the back of an abstract.

Where patents genuinely are useful is in understanding capability and intent, because if a company spent money on filing something then engineers spent time working on the problem and somebody in that business considered the problem real enough to protect, and that's worth knowing even if you can't map it directly to a live system.

DOES GOOGLE ANALYSE PAGE LAYOUT AT ALL?

The patent to start with here is US7676745B2, "Document segmentation based on visual gaps", which was invented by Daniel Egnor, filed in December 2004 and assigned to Google, and it's one of the most overlooked filings in the entire portfolio as far as SEO is concerned.

The core idea in it is that a document can be segmented based on a visual model of that document, and the visual model is determined according to the amount of visual white space or gaps that appear in the document, with that visual model then being used to identify a hierarchical structure of the page which is then used to segment the page into parts. In other words Google filed a patent on understanding the structure of a page by measuring the space between things, which is a fairly significant thing to have on file given how often you'll hear that Google is only reading text.

The detail underneath is better still, because the filing is explicit that this isn't simply DOM parsing, i.e. different HTML elements may be assigned various weights, numerical values, that attempt to quantify the magnitude of the visual gap that element introduces into the rendered document, so a break tag and a div and a table cell boundary aren't just structural markers sitting in a tree, they carry a value that represents how much visual separation a human being would actually perceive when looking at the page.

The worked examples in that patent are mostly focused on pulling apart individual business listings and reviews that happen to share a single page, as it came out of the local search side of things, but the filing itself states that the general hierarchical segmentation technique could be applied to any type of signal in a document rather than just geographic ones, so the scope was deliberately left wide open from the start.

IT WASN'T JUST GOOGLE DOING THIS

The thing that turns this from an interesting patent into an actual argument is that Google wasn't working on this in isolation, all three of the major search engines of that era arrived at the same conclusion independently and within roughly the same window, which is a much stronger signal than any single filing on its own.

Microsoft has US7428700B2, "Vision-based document segmentation", which is the VIPS work, and it takes the same fundamental approach of identifying visual blocks in a document and detecting the separators between those blocks in order to build a content structure, but it goes a step further than Google's public claims because in the Microsoft filing the blocks themselves get ranked according to how well they match the query criteria, with document level rankings then being derived from those block level rankings.

Yahoo has US8849725B2, "Automatic classification of segmented portions of web pages", filed in 2009, and this is the most explicit of the three when it comes to design and layout because the machine learning feature space includes layout features taken on render, i.e. the absolute size and position of each segmented portion of the page and the size of a segmented portion relative to the page as a whole, and those features feed into segment classification where segment scores relate to segment content quality scores, with the filing stating that this information can be provided to the ranking function of a search engine.

So you've got three separate search engines, all working on this at roughly the same time, all concluding that the rendered layout of a page carries meaning and that the meaning is worth feeding into how documents get ranked, which is a much more compelling case than pointing at one Google patent and hoping nobody asks follow up questions.

THE REASONABLE SURFER PATENT IS REALLY A DESIGN PATENT

The big one for this whole article is US7716225B1, "Ranking documents based on user behavior and/or feature data", filed by Google in June 2004 and granted in May 2010, with a continuation sitting at US8117209B1, and the SEO industry generally refers to this as the Reasonable Surfer patent and files it away mentally under link building, which I think is a mistake because when you actually sit down and read it what you've got is a design and layout patent wearing a link patent's coat.

The mechanism described is that the system builds a model out of two things, feature data relating to the different features of a link, and user behaviour data relating to the navigational actions people took around those links, and it then uses that model to assign weights to links which subsequently feed into document ranking, so it's learning from real behaviour rather than assuming that every link is equal.

And this is the thing that I want you to keep in mind!

Where it gets interesting for anyone thinking about design is in what counts as feature data, because the examples given in the filing include the font size of the anchor text, the position of the link measured in terms of whether it sits in an HTML list or in running text, whether it appears above or below the first screenful viewed on an 800 x 600 browser display, which side of the document it sits on i.e. top, bottom, left or right, and whether it's in a footer. A link sitting in the main content area in a font and colour that make it stand out, positioned near the top of the page and pointing somewhere topically related, may carry a considerably higher probability of being clicked and therefore pass a fair amount of value, whereas that same link in the footer in the same colour and font as everything around it, pointing at something unrelated, may pass very little.

That's Google patenting the exact mechanism people keep speculating about, i.e. observing how real people behave, correlating that behaviour against visual placement, typography and screen position, learning the relationship and then applying it as a weight in ranking, and it's been sat on file since 2004 which is long before most of the current debate around UX signals even started.

IMPORTANT CAVEAT HERE!

Google doesn't actually need Chrome in order to analyse your layout and design, because it renders your pages itself, i.e. the same leak that revealed chromeInTotal also revealed rendering systems like HtmlrenderWebkitHeadless which deal with JavaScript pages, and the visual gaps patent from 2004 predates Chrome entirely since Chrome didn't launch until 2008. All the layout segmentation work we went through in the article was being done long before Google had a browser, and it's done on Google's own infrastructure against pages it has crawled, so Chrome adds nothing there.

It's scoped to links rather than whole page layouts, which is an important caveat and I'll come back to it, but the underlying mechanism is precisely the one being asked about. The 800 x 600 reference is worth pausing on too because it dates the thinking with total precision and it tells you that the fold was being modelled as a real measurable pixel boundary rather than as a metaphor, and whilst what that means in 2026 across phones, tablets, laptops and ultrawide monitors is a genuinely open question, the intent behind it was never remotely ambiguous.

WHAT ABOUT USER BEHAVIOUR ON ITS OWN?

Separately from all the layout work you've got the click and dwell patents, which are much better known in SEO circles so I'll keep this part relatively short and to the point.

The "Modifying search result ranking based on implicit user feedback" family covers US8661029B1, US10229166B1, US11188544 and various others, and the mechanism throughout is a measure of relevance built from the ratio of longer views of a document result to shorter views of that same result, with that measure then being output to a ranking engine to influence future searches for the same query, which is the long clicks good and short clicks bad idea expressed with far more mathematical care than the usual SEO summary of it.

This is the family sitting underneath what we now know as Navboost, and it's worth stressing that pretty much everything we know about Navboost as a live production system comes from Pandu Nayak's testimony during the 2023 US antitrust proceedings rather than from any official Google documentation, so whilst sworn testimony carries a lot more weight than a blog post it still isn't a specification document.

Alongside that you've got the Panda lineage in US9767157B2, "Predicting site quality" from Navneet Panda and Yun Zhou, which builds phrase models from previously scored sites and then uses relative phrase frequencies to predict a quality score for a site the system hasn't seen before. That one is worth mentioning mainly because it's entirely phrase based, i.e. there is no layout or design component in it whatsoever, which is a useful corrective to the idea that everything Google does is secretly about presentation.

THE PATENT RARELY MENTIIONED

Here's the thing that complicates the whole general user behaviour argument in SEO and it's a filing that gets almost no coverage in SEO despite being probably the most relevant one to this article topic. I've seen people in SEO discuss patents, but this one, nadda, almost nothing - yet to me I think it's key to understand it.

US8938463B1 is titled "Modifying search result ranking based on implicit user feedback and a model of presentation bias", and what it describes is taking a feature that indicates presentation bias, building a prior model that represents the background probability of a result being selected given the values of those presentation features, and then outputting that prior model to the ranking engine specifically in order to reduce the influence of presentation bias on the final ranking - again this will tie in with so many other things around NavBoost, User Engagement, Twiddlers etc.

This patent is worth reading carefully because it cuts directly against the naive version of the "design affects rankings" argument, i.e. Google didn't just patent noticing that presentation influences behaviour, Google patented a method for subtracting that influence back out again so that what's left over is closer to genuine user satisfaction rather than being an artefact of how something happened to be displayed. So if your pages are generating good behavioural signals then part of the patent stack exists specifically to work out how much of that was your layout doing the heavy lifting versus your content actually answering the question, which is exactly what you'd want to be doing if you were running a search engine and exactly what you'd hate if you were hoping to win purely on presentation.

This is absolutely KEY, why?

BECAUSE GOOD CONTENT ISN'T ALWAYS HELPFUL CONTENT AND HELPFUL CONTENT ISN'T ALWAYS GOOD CONTENT!!!!!

Why do I shout this?

Because if you think back to the initial role out of Google's Helpful Content algo it absolutely DECIMATED millions of sites (many of amazing quality).

I saw SO many content sites, publishers and affiliates get wiped out overnight, September 2023, it was absolutely brutal, but, to me, I got it.

Those in the SEO community who have followed me over the years know I've been massively vocal on Google, SEO testing, findings etc. Many know that I run seo-audits.io and during that time, I was auditing websites back to back that got truly wiped out by Google's HCU algo.

Why?

  • HCU wasn't about content quality - it was about HELPFUL CONTENT - and why did that impact sites with GOOD CONTENT? it was because the weighting attribution that Google used went more towards user behaviour derrived from behavioural data - this is why SO many UGC sites like Reddit saw growth

  • HCU was more about behaviour, and we know this because the sites that often saw HCU wipeouts had good content but often had issues with content bias (because affiliate sites tend to be biased to get you to buy products) and because people wanted real feedback and insights on pre-purchases, so they'd go to UGC communities where they'd engage longer with the content because it was often raw, unfiltered, unbiased

We saw SO many legitimate sites get wiped out, remember Housefresh?

They got ABSOLUTELY obliterated despite having some of the best home air purifier reviews on the web!

It wasn't about content quality, it was about BEHAVIOUR!.

SO COULD ANY OF THIS FORM PART OF AN ALGORITHM?

The honest answer, and I've looked pretty hard for a reason to say otherwise, is that the capability is unambiguously there across the portfolio but the loop is never explicitly closed in any filing I can find.

Google can segment a rendered page using visual gaps, it can classify what each segment of that page actually is, it can tell whether something sits above or below the first screenful, it knows font sizes and positions and footer placement, it holds a twenty year old patent on learning from real user behaviour which of those visual arrangements people actually engage with, and it has an entirely separate patent family covering how to turn long clicks and short clicks into a usable relevance measure. Every single component you'd need exists out there, but to what degree its used, we don't know.

What doesn't exist, as far as I can see, is a patent that says generate a layout quality score and then rank documents on that layout quality score, nothing closes it that neatly, and anyone telling you otherwise is guessing rather than citing. The closest Google ever came publicly was the Page Layout algorithm, i.e. the Top Heavy update from January 2012 where sites with excessive ads above the fold were demoted, and even then Google's public line was that a variety of signals algorithmically determine what type of ad or content appears above the fold with no further detail to share, and no patent was ever pointed at it.

We also had the pop up interstitials amongst other things (oh, and the making it difficult to exit or click back is thrown in here for good measure).

Where that leaves us is that design and layout almost certainly aren't a DIRECT ranking factor with a score attached to them, but the components required to infer layout quality from behaviour have been sat in Google's portfolio since 2004, they've been refined and re-filed repeatedly since then, and there's an entire patent dedicated to being careful about how presentation contaminates behavioural data, which is a very different picture from "Google can't see your design" and an equally different picture from "make it look nice and you'll rank".

IMO for what it's worth, I honestly believe that Google can and does use aggregated data, so for example:

  1. Google could look at common layouts for high performance sites - the ones with the best click behaviour / positive click signals

  2. It could use that data to assess "document quality" by the presentation and layout of that content - it could use extrapolated data as part of re-ranking or initial document rank asssessment

  3. Google could go as far as looking at common click behaviours by element type, content position, media types, sub-links and more

WHAT SHOULD YOU ACTUALLY DO WITH ANY OF THIS?

None of the above is worth much if it just makes you feel clever at a conference, so here's how I'd translate it into things that are actually worth doing - and I'm not just saying this, I do everything as part of SEO testing to see what works for me when I'm doing SEO.

Some of the key things I woulod generally advise include:

Get your answer above the first screenful, not because there's some fold penalty waiting to be applied but because every patent in this stack treats screen position as a real and measurable property of a page, and because burying the answer produces short clicks and short clicks are the one thing that the implicit feedback family is unambiguously designed to punish.

Use genuine visual separation between the sections of your pages, because the visual gaps patent means whitespace is doing structural work whether you intended it to or not, and sections that run into each other are harder to segment cleanly which means they're harder to understand cleanly.

Stop treating footer links as though they're equivalent to body links, because the Reasonable Surfer filing is completely explicit on this point and position, font size, colour and surrounding context all feed into the weight assigned, so a hundred footer links were never the same as a hundred contextual body links and the patent has been saying so since 2004.

Make your main content visually dominant relative to everything else on the page, because the Yahoo filing weights segment size relative to the overall page, and if your navigation, sidebar, cookie banner and newsletter interstitial are collectively occupying more visual real estate than the thing people actually came to the page for then you're making an implicit statement about information priority that machines can measure.

And don't over-engineer specifically for behavioural signals, because US8938463B1 is Google effectively telling you that they're actively working to strip presentation effects back out of the data, so anything you do purely to create the appearance of engagement is being modelled as noise by design.

THE FULL LIST OF PATENTS

All of these are free to read on Google Patents and I'd strongly recommend reading the claims rather than the abstracts, because the abstracts oversell every single one of them. If you want to search the portfolio yourself then use the assignee syntax, i.e. inassignee:"Google LLC" combined with terms like rendered, layout, viewport, segment or visual, and filter down to granted patents only because a published application that never went on to grant is a much weaker thing to hang an argument on.

  • US7676745B2 - Document segmentation based on visual gaps - Google

  • US7716225B1 and US8117209B1 - Ranking documents based on user behavior and/or feature data - Google

  • US8661029B1, US10229166B1 and US11188544 - Modifying search result ranking based on implicit user feedback - Google

  • US8938463B1 - Modifying search result ranking based on implicit user feedback and a model of presentation bias - Google

  • US9767157B2 - Predicting site quality - Google

  • US7428700B2 - Vision-based document segmentation - Microsoft

  • US8849725B2 - Automatic classification of segmented portions of web pages - Yahoo