This has been submitted to HN quite a few times. https://news.ycombinator.com/item?id=25426662 (Most comments; 11 comments) https://news.ycombinator.com/item?id=25417067 (3 comments) https://news.ycombinator.com/item?id=25546867 (Most recent; 89 days ago) https://news.ycombinator.com/item?id=25543859 https://news.ycombinator.com/item?id=25424852
Only Google is really allowed to crawl the web
81–90 of 365 posts
Re: Only Google is really allowed to crawl the web
#82Earlier quoted context omitted.
Perhaps there could be some kind of 'Crawler consortium'? Under this consortium, website owners would be allowed to either allow all crawlers (approved by the consortium) or none at all (that is, none that is in the consortium, i.e. you could allow a specific researcher or something to crawl your website on a case-by-case basis). This consortium would be composed of the search engines (Google, MS, other industry memb…
> Perhaps there could be some kind of 'Crawler consortium'? An industry-wide agreement not to compete for commercially valuable access to suppliers of data? Comprised of companies that are current (and in some cases perennial) focusses of antitrust attention? I think there might be a problem with that plan.
Could you better describe your objections?
Re: Only Google is really allowed to crawl the web
#83Re: Only Google is really allowed to crawl the web
#84The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
You're wrong on a lot of facts here. Google Flights doesn't get its data just by crawling, they get it from Sabre, the FAA, Eurocontrol, etc. Airlines are, obviously, extremely pleased to disseminate this information. Google Flights "gives back" in the exact same way as any other travel outlet: they book passengers. As for Wikipedia, the WMF is quite happy that most of their traffic is now served by Google. WMF is in…
They have relented in some ways, rolling out stuff in the widget like: "The airline has issued a change fee waiver for this flight. See what options are available on American's website"
But obviously, that kind of stuff isn't shown on Google for quite some time after it exists on the source site. And the widget pushes the organics off the fold unless you have a huge monitor.
As for Wikipedia, I was referring to this: https://news.ycombinator.com/item?id=26487993
"Airlines are, obviously, extremely pleased to disseminate this information"
In the same way that publishers love AMP, yes. They don't actually like it, but they are forced to make the best of it.
Re: Only Google is really allowed to crawl the web
#85Earlier quoted context omitted.
> run [by] a private company and accessed for a small fee That is exactly the opposite of a public cache.
Not really. It serves the same function. Either you pay this hypothetical company or ??? pays to keep up the public one.
That being said what I think you're arguing for would be the implementation of a public utility or private-public business. If that's the case then yes, what you're saying is correct.
Re: Only Google is really allowed to crawl the web
#86Re: Only Google is really allowed to crawl the web
#87I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…
A company I worked for ~7 years ago ran its own focused web crawler (fetching ~10-100m pages per month, targeting certain sections of the web). There were a surprising number of sites out there that explicitly blocked access to anyone but Google/Bing at the time. We'd also get a dozen complaints or so a month from sites we'd crawled. Mostly upset about us using up their bandwidth, and telling us that only Google was…
Yes, the bad bots don't give a fuck, but even the non-malicious bots (ahrefs, moz, some university's search engine etc) don't bring any value to me as a site owner, take up band width and resources and fill up logs. If you can remove them with three lines in your robots.txt, that's less noise. Especially universities do, in my opinion, often behave badly and are uncooperative when you point out their throttling does not work and they're hammering your server. Giving them a "Go Away, You Are Not Wanted Here" in a robots.txt works for most, and the rest just gets blocked.
Re: Only Google is really allowed to crawl the web
#88Earlier quoted context omitted.
They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…
Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.
Re: Only Google is really allowed to crawl the web
#89https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…
Are there any actual repercussions for just ignoring robots.txt?
Re: Only Google is really allowed to crawl the web
#90The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…
2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned.
The instant Google becomes confident they can supplant Wikipedia, they will.