Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

161–170 of 365 posts

Re: Only Google is really allowed to crawl the web

#161

This is not really about Google. Websites block crawlers because they get abused / crashed by Crawlers. In the early days (2000-2010) Google not only got banned by some websites, it even got DNS-banned for abusing some DNS domains. You see, Google already has already built the "megacrawlers" described in this article, it can melt any website on the Internet, even Facebook - the largest, and they paid a high price for…

Thank you for those insights, it's a topic I'm interested in. Agree with what you're saying about naive bots hitting websites/hosts/subnets too hard, in the context of site owners being hit by multiple bots for multiple reasons and them questioning the return they'll get.

I'd be interested to know more info wrt DNS lookups. Did you apply a blanket rate limit on the number of DNS requests you'd make to any particular server?

From past experience I know the .uk Nominet servers would temp-ban if you were doing more than a few hundred lookups per second. At the next host level down, was there a blanket limit or was it dependent on the number of domains that nameserver was responsible for?

Re: Only Google is really allowed to crawl the web

#162

This has been submitted to HN quite a few times. https://news.ycombinator.com/item?id=25426662 (Most comments; 11 comments) https://news.ycombinator.com/item?id=25417067 (3 comments) https://news.ycombinator.com/item?id=25546867 (Most recent; 89 days ago) https://news.ycombinator.com/item?id=25543859 https://news.ycombinator.com/item?id=25424852

Wasn't aware of that. Resubmitting interesting content that hasn't got traction earlier on is however explicitly allowed in the guidelines IIRC.

And linking past threads on the same subject is helpful.

Re: Only Google is really allowed to crawl the web

#163
Maybe it would be nice if some sort of simple central index of "URLs + their last updated timestamp/version/eTag/whatever" would exist, updated by the site owners themselves? ("push"-notification)

Meaning that whenever a page of a website would be created or updated, that website itself would actively update that central index, basically saying "I just created/deleted page X" or "I just updated the contents of page X".

The consequence would be that...

1) ...crawlers would not have anymore to actively (re)scan the whole Internet to find out if anything has changed, but they would only have to query that central index against their own list URLs & timestamps to find out what needs to be (re)scanned.

2) ...websites would not have to just wait&hope that some bot would decide to come by to have a look at their sites, nor they would have to answer over and over again requests that are just meant to check if some content has changed.

Re: Only Google is really allowed to crawl the web

#164
post #127
post #105

Earlier quoted context omitted.

Is it legal for a government entity to issue a robots.txt like that? Maybe the line between use and abuse hasn't been delinated as well as it needs to be.

Is failure to honor a robots.txt a crime? Or rather, would it be unlawful to spoof a user agent to access this publicly available data? After the linkedin [0] case it seems reasonable to think not. [0]: https://www.eff.org/deeplinks/2019/09/victory-ruling-hiq-v-l...

Spoofing user-agents hasn't worked in a long time for anything but small operations because search engines publish specific IP ranges their scrapers use.

Re: Only Google is really allowed to crawl the web

#165
On a related note, Cloudflare just introduced "Super Bot Fight Mode" (https://blog.cloudflare.com/super-bot-fight-mode/) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower until pages won't load at all), presented with CAPTCHAs or outright blocked. In my opinion this will turn the part of the web that Cloudflare controls into a walled garden not unlike Twitter or Facebook: In theory the content is "public", but if you want to interact with it you have to do it on Cloudflare's terms. Quite sad really to see this happen to the web.

Re: Only Google is really allowed to crawl the web

#166
post #93
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

The census data is available for bulk download, mostly as CSV (for example [1]). Scraping census.gov is worse for both the Census Bureau (which might have to do an expensive database query for each page) and for the scraper (who has to parse the page).

Blocking scrapers in robots.txt is more of a way of saying, "hey, you're doing it wrong."

It's also worth noting that the original article is out of date. The current robots.txt at census.gov is basically wide-open [2].

[1] https://www.census.gov/programs-surveys/acs/data/data-via-ft...

[2] https://www.census.gov/robots.txt

Re: Only Google is really allowed to crawl the web

#167

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

The current anti-trust doctrine in the US has a goal of protecting consumers - not competition. What Google is doing is arguably great for consumers but awful to their competitors/other organizations. Technically, companies can simply block Google using robots.txt - but in reality that will lose them more money than the current partial disintermediation by Google is costing them - and Google knows this.

It's a tall order to convince the courts that Google's actions consumers, or is illegal: after all, being innovative in ways that may end up hurting the competition is a key feature of a capitalist society - proving that a line has been crossed is really hard, by design.

Re: Only Google is really allowed to crawl the web

#168

Can we take a moment to talk about this club's business model? There's not even any information to see what the "private forum access" that you have to pay for is about, what kind of people are in it...or even to know about what happens with the money. For me, this sounds like a scam. I mean, no information about any company. No imprint. No privacy policy. No non-profit organization. And just a copy/paste wordpress i…

They want you to pay them to "research" google's web crawling monopoly. It's really just a donation, but they don't frame it like that. Probably more credible than using a crowd funding website, because it sounds like their pushing for actual legislation.

> Meet with legislators and regulators to present our findings as well as the mock legislation and regulations. We can’t expect that we can publish this website or a PDF and then sit back while governments just all start moving ahead on their own. Part of the process is meeting with legislators and regulators and taking the time helping them understand why regulating Google in this way is so important. Showing up and answering legislators’ questions is how we got cited in the Congressional Antitrust report and we intend to keep doing what’s worked so far.

Re: Only Google is really allowed to crawl the web

#169

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

On the other hand, I do not want my site to go down thanks to a few bad 'crawlers' that fork() a thousand http requests every second and take down my site, forcing me to do manual blocking or pay for a bigger server/scale-out my infrastructure. Why should I have to serve them?

Re: Only Google is really allowed to crawl the web

#170

Earlier quoted context omitted.

It sounds like it - and third-party companies will often show you flights that involve different companies on the different legs - which can leave you in a pickle because technically each airline's job is to get you to the end of THIER flight, not the entire journey.

And sometimes with a change of airport!

I remember when in Germany some budget airlines used to say they'd fly to "Frankfurt" (FRA) but actually flew to "Frankfurt-Hahn" (HHN) - 115km away. After arrival in HHN they put you on a bus to FRA that took about 2 hours.
Post reply on HN