Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

81–90 of 365 posts

Re: Only Google is really allowed to crawl the web

#81

This has been submitted to HN quite a few times. https://news.ycombinator.com/item?id=25426662 (Most comments; 11 comments) https://news.ycombinator.com/item?id=25417067 (3 comments) https://news.ycombinator.com/item?id=25546867 (Most recent; 89 days ago) https://news.ycombinator.com/item?id=25543859 https://news.ycombinator.com/item?id=25424852

Hooray! Looks like I'm one of today's lucky 10,000. :)

https://xkcd.com/1053/

Re: Only Google is really allowed to crawl the web

#82

Earlier quoted context omitted.

Perhaps there could be some kind of 'Crawler consortium'? Under this consortium, website owners would be allowed to either allow all crawlers (approved by the consortium) or none at all (that is, none that is in the consortium, i.e. you could allow a specific researcher or something to crawl your website on a case-by-case basis). This consortium would be composed of the search engines (Google, MS, other industry memb…

> Perhaps there could be some kind of 'Crawler consortium'? An industry-wide agreement not to compete for commercially valuable access to suppliers of data? Comprised of companies that are current (and in some cases perennial) focusses of antitrust attention? I think there might be a problem with that plan.

Well, yes, and a common solution to anti-trust cases, that I know of, is some kind of industry self-regulation. In this case I wouldn't trust the industry only to self-regulate; hence, they should at invite (while keeping a minority but not insignificant position) governments and civil society (ngos and other organizations) to participate.

Could you better describe your objections?

Re: Only Google is really allowed to crawl the web

#84
post #70
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

You're wrong on a lot of facts here. Google Flights doesn't get its data just by crawling, they get it from Sabre, the FAA, Eurocontrol, etc. Airlines are, obviously, extremely pleased to disseminate this information. Google Flights "gives back" in the exact same way as any other travel outlet: they book passengers. As for Wikipedia, the WMF is quite happy that most of their traffic is now served by Google. WMF is in…

They don't get individual flight status (what I was talking about) from Sabre or the FAA or Eurocontrol. I didn't get into fares and planned schedules and Google Flights, that's a different topic. I was talking about the big widget you get for queries on status for a particular flight, which is not Google Flights.

They have relented in some ways, rolling out stuff in the widget like: "The airline has issued a change fee waiver for this flight. See what options are available on American's website"

But obviously, that kind of stuff isn't shown on Google for quite some time after it exists on the source site. And the widget pushes the organics off the fold unless you have a huge monitor.

As for Wikipedia, I was referring to this: https://news.ycombinator.com/item?id=26487993

"Airlines are, obviously, extremely pleased to disseminate this information"

In the same way that publishers love AMP, yes. They don't actually like it, but they are forced to make the best of it.

Re: Only Google is really allowed to crawl the web

#85
post #47

Earlier quoted context omitted.

> run [by] a private company and accessed for a small fee That is exactly the opposite of a public cache.

Not really. It serves the same function. Either you pay this hypothetical company or ??? pays to keep up the public one.

Just because it serves the same function does not mean the implementation is the same. Private military contractors and a US infantry squad serve the same function, but the implementation completely changes their context.

That being said what I think you're arguing for would be the implementation of a public utility or private-public business. If that's the case then yes, what you're saying is correct.

Re: Only Google is really allowed to crawl the web

#87
post #18

I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…

A company I worked for ~7 years ago ran its own focused web crawler (fetching ~10-100m pages per month, targeting certain sections of the web). There were a surprising number of sites out there that explicitly blocked access to anyone but Google/Bing at the time. We'd also get a dozen complaints or so a month from sites we'd crawled. Mostly upset about us using up their bandwidth, and telling us that only Google was…

I usually recommend setting only Google/Bing/Yandex/Baidu etc to Allow and everything else to Disallow.

Yes, the bad bots don't give a fuck, but even the non-malicious bots (ahrefs, moz, some university's search engine etc) don't bring any value to me as a site owner, take up band width and resources and fill up logs. If you can remove them with three lines in your robots.txt, that's less noise. Especially universities do, in my opinion, often behave badly and are uncooperative when you point out their throttling does not work and they're hammering your server. Giving them a "Go Away, You Are Not Wanted Here" in a robots.txt works for most, and the rest just gets blocked.

Re: Only Google is really allowed to crawl the web

#88

Earlier quoted context omitted.

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.

Can you elaborate on this? Do you mean shorter layovers?

Re: Only Google is really allowed to crawl the web

#89
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Are there any actual repercussions for just ignoring robots.txt?

Your crawler's IP might get banned, eventually.

Re: Only Google is really allowed to crawl the web

#90
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla.

2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned.

The instant Google becomes confident they can supplant Wikipedia, they will.

Post reply on HN