Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

61–70 of 365 posts

Re: Only Google is really allowed to crawl the web

#61
post #50
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

>However, that sort of search used to (most times) lead to a visit to the airline web site. I don't think that's correct. In the old days you'd either call a travel agent or use an aggregator like expedia. Google muscles out intermediaries like Expedia, Yelp, and so on. It's not likely much better or worse for the end user or supplier. Just swapping one middleman for another.

Only if Google stays around long term. I wouldn't be surprised if each free product on its graveyard took down a dozen of competing products before it was killed of.

Re: Only Google is really allowed to crawl the web

#62
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Perhaps there could be some kind of 'Crawler consortium'?

Under this consortium, website owners would be allowed to either allow all crawlers (approved by the consortium) or none at all (that is, none that is in the consortium, i.e. you could allow a specific researcher or something to crawl your website on a case-by-case basis).

This consortium would be composed of the search engines (Google, MS, other industry members), as well as government appointed individuals and relevant NGOs (electronic frontier foundation, etc?). There would be an approval process that simply requires your crawl to be ethical and respect bandwidth usage. Violations of ethics or bandwidth limits could imply temporary or permanent suspension. The consortium could have some bargain or regulatory measures to prevent website owners from ignoring those competitive and fairness provisions.

Re: Only Google is really allowed to crawl the web

#63
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I swear something like 50% of those digests are totally incorrect as well. It's amazing they have kept the feature because it has never had a very high signal-to-noise ratio. I never trust what's presented in these digests without double-checking the source page.

Re: Only Google is really allowed to crawl the web

#64
post #54

Earlier quoted context omitted.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

I don't think you're quite clear on what the words "public" and "private" mean. "Public" is not a synonym for "run by the government" and "private" is not a synonym for "closed to everyone but the owner". Restaurants, for example, are generally open to the public, but they are not public. A restaurant owner is, with a few exceptions, free to refuse service to anyone at any time. If it's "exactly the same as a public…

Ok I am not clear then, but I’m less clear after your comment! In a public cache, who would you want to decide who has access? Is simply saying “anyone who pays has access” enough to qualify as public? if so, then I agree and this was my (possibly poorly phrased) intention in the original comment.

But imo the restaurant model is also fine; in most cases people have access and it works.

Re: Only Google is really allowed to crawl the web

#65
Around a decade ago, I was part of the team responsible for msnbot (a web crawler for bing). There used to be robot.txt (forgot the extension now). Most of the website was giving 10-20x higher limits to googlebot than rest other crawler.

Google definitely has unfair advantage there.

Bing and duckduckgo still provide very reasonable result with 10-20x less resources but not at par of google.

Re: Only Google is really allowed to crawl the web

#66
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Perhaps there could be some kind of 'Crawler consortium'? Under this consortium, website owners would be allowed to either allow all crawlers (approved by the consortium) or none at all (that is, none that is in the consortium, i.e. you could allow a specific researcher or something to crawl your website on a case-by-case basis). This consortium would be composed of the search engines (Google, MS, other industry memb…

> Perhaps there could be some kind of 'Crawler consortium'?

An industry-wide agreement not to compete for commercially valuable access to suppliers of data?

Comprised of companies that are current (and in some cases perennial) focusses of antitrust attention?

I think there might be a problem with that plan.

Re: Only Google is really allowed to crawl the web

#67
post #8

Earlier quoted context omitted.

Seriously? Google is a private cache of the web. That is the problem.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

[deleted]

Re: Only Google is really allowed to crawl the web

#68
post #61
post #50

Earlier quoted context omitted.

>However, that sort of search used to (most times) lead to a visit to the airline web site. I don't think that's correct. In the old days you'd either call a travel agent or use an aggregator like expedia. Google muscles out intermediaries like Expedia, Yelp, and so on. It's not likely much better or worse for the end user or supplier. Just swapping one middleman for another.

Only if Google stays around long term. I wouldn't be surprised if each free product on its graveyard took down a dozen of competing products before it was killed of.

Then someone can start a competitor up again, right? Assuming there's actually a market for it.

Re: Only Google is really allowed to crawl the web

#69
post #28

Earlier quoted context omitted.

Imagine if Donald Trump decided that indexing Joe Biden's campaign site was unacceptable. A mandated singular public cache has potential slippery slopes.

Imagine if Donald Trump decided to tax campaign donations to Joe Biden's campaign at 100%. I am unconvinced by the "slippery slope" argument being deployed by default to any governmental attempt to combat tech monopolies.

This is an argument against centralization more than it is against government.

"One index to rule them all" seems more fraught with difficulty than, "large cloud providers are unhappy that crawlers on the open web are crawling the open web".

Re: Only Google is really allowed to crawl the web

#70
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

You're wrong on a lot of facts here. Google Flights doesn't get its data just by crawling, they get it from Sabre, the FAA, Eurocontrol, etc. Airlines are, obviously, extremely pleased to disseminate this information. Google Flights "gives back" in the exact same way as any other travel outlet: they book passengers.

As for Wikipedia, the WMF is quite happy that most of their traffic is now served by Google. WMF is in the business of distributing knowledge, not in the eyeballs business. Serving traffic is just a cost for them. The main problem has been that the average cost for Wikipedia to serve a page has gone up, because many readers read it via Google, and more people who visit Wikipedia are logged-in authors, which costs them more to serve. I'm sure there's an easy solution to this problem (for example, beneficiaries of Wikipedia can donate compute facilities and services, or something along those lines).

Post reply on HN