Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

31–40 of 365 posts

Re: Only Google is really allowed to crawl the web

#31
post #6

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

A lot of news websites restrict any crawler other than Google. And this does not happen only via robots.txt.

Indeed, years ago I had scripts to automatically fetch URLs from IRC and I quickly realized that if I didn't spoof the user agent of a proper web browser many websites would reject the query. Googlebot's UA worked just fine however.

Re: Only Google is really allowed to crawl the web

#32
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

> If it was the only indexer allowed, and it was publically governed

Which would put it under government regulation and be forever mired in politics over what was moral, immoral, ethical or unethical and all other kerfuffle. To an extent, it’s already that way, but that would make it worse than it is currently.

Re: Only Google is really allowed to crawl the web

#33
post #17

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

It's hilarious to think there exists people who think googlebot does not get special treatment from website operators. Here is an experiment you can do in a jiffy, write a script that crawls any major website and see how many URL fetches it takes before your IP gets blocked. Googlebot has a range of IP addresses that it publicly announces so websites can whitelist them.

[deleted]

Re: Only Google is really allowed to crawl the web

#35
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

So out law web scrapping entirely?

Re: Only Google is really allowed to crawl the web

#36

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

See figure I.4 on page 24 of this UK government report: https://assets.publishing.service.gov.uk/media/5efb1db6e90e0...

Additional evidence here: https://knuckleheads.club/the-evidence-we-found-so-far/

Re: Only Google is really allowed to crawl the web

#37
post #24

Earlier quoted context omitted.

I don't understand that. The crawling access is mostly the same as it ever was. Google's SERP pages are not. A mutually beneficial search engine that respects it's sources would still crawl the same. Google used to be that. The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search…

That's incorrect. Before the search oligopolies formed, new search engines could start up. There was excite, hotbot, altavista, and more. Now they don't have access. Search these comments for census.gov.

There are companies that do pretty well in this space, like ahrefs, for example. They do resort to trickery, like proxy clients that look like home computers or cell phones. But, if a small entity like ahrefs can do it, anyone can do it.

In a nutshell, though, I don't see equal access for all crawlers changing anything. Maybe that's the first barrier they hit, but it isn't the biggest or hardest one by far. Bing has good crawler access, but shit market share.

Re: Only Google is really allowed to crawl the web

#38
post #24

Earlier quoted context omitted.

That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.

I don't understand that. The crawling access is mostly the same as it ever was. Google's SERP pages are not. A mutually beneficial search engine that respects it's sources would still crawl the same. Google used to be that. The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search…

[deleted]

Re: Only Google is really allowed to crawl the web

#39
post #28
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

Imagine if Donald Trump decided that indexing Joe Biden's campaign site was unacceptable. A mandated singular public cache has potential slippery slopes.

Imagine if Donald Trump decided to tax campaign donations to Joe Biden's campaign at 100%.

I am unconvinced by the "slippery slope" argument being deployed by default to any governmental attempt to combat tech monopolies.

Re: Only Google is really allowed to crawl the web

#40

Even that won't change much. There is no way Google can be out-googled by other search engines because of its market dominance: more traffic means more clicks, more clicks mean better search results, better search results will drive more traffic. I try bing and DDG for a week or so every 6 months. I always switch back to google eventually because the results are so much better. Google can only be disrupted if somethi…

Yup. My opinion has long been that the only thing that will take down google is a massive increase in NLP, such that the historical click data can be outperformed by a straight up really good NLP model
Post reply on HN