Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

1–10 of 365 posts

Re: Only Google is really allowed to crawl the web

#2
I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt.

If I actually wanted to restrict bots, it would be much easier to restrict googlebot because they actually follow the rules.

I don't disagree in principle that there should be an open index of the web, but for once I don't see Google as a bad actor here.

Re: Only Google is really allowed to crawl the web

#3
https://knuckleheads.club/the-googlebot-monopoly/ has actual details.

> Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Microsoft, Yahoo and two other non-search engines are not allowed to crawl certain pages on census.gov, but are otherwise allowed to crawl whatever else they can find on the website. This tells us that there are two different classes of crawlers in the eyes of the operators of census.gov: those given wide access, and those that are totally denied.

> And, broadly speaking, when we examine the robots.txt files for many websites, we find two classes of crawlers. There is Google, Microsoft, and other major search engine providers who have a good level of access and then there is anyone besides the major crawlers or crawlers that have behaved badly in the past that are given much less access. Among the privileged, Google clearly stands out as the preferred crawler of choice. Google is typically given at least as much access as every other crawler, and sometimes significantly more access than any other crawler.

Re: Only Google is really allowed to crawl the web

#5
The idea of a public cache available to anyone who wishes to index it is ... kind of compelling.

If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines.

I don't think it'll ever happen, but it's interesting to think about.

Re: Only Google is really allowed to crawl the web

#6

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

A lot of news websites restrict any crawler other than Google. And this does not happen only via robots.txt.

Re: Only Google is really allowed to crawl the web

#7

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

LinkedIn profile/Quora answer are accessible by Google bot without signin

Re: Only Google is really allowed to crawl the web

#9

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

What do you think this is used for?

https://developers.google.com/search/docs/advanced/crawling/...

Re: Only Google is really allowed to crawl the web

#10
Even that won't change much. There is no way Google can be out-googled by other search engines because of its market dominance: more traffic means more clicks, more clicks mean better search results, better search results will drive more traffic.

I try bing and DDG for a week or so every 6 months. I always switch back to google eventually because the results are so much better.

Google can only be disrupted if something new is invented, something different than search but delivering way better results. I have no clue what that might be. But I hope someone is working on it.

Post reply on HN