Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

211–220 of 365 posts

Re: Only Google is really allowed to crawl the web

#211

Earlier quoted context omitted.

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. 2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned. The instant Google becomes confident they can supplant Wikipedia, they will.

Not sure why you're being downvoted; I completely agree with what you're saying (modulo questionable usage of "sponsor"). If Wikipedia were to try to charge for this use of their data, Google would likely make it a priority to drop the Wikipedia blurbs, either without replacement, or with data sourced elsewhere.

Re: Only Google is really allowed to crawl the web

#212

Earlier quoted context omitted.

Aren't there anti trust laws to prevent this kind of thing?

Antiturst laws are hard to enforce in the United States. Monopolies themselves aren't illegal. To be convicted of an antitrust violation, a firm needs to both have a monopoly and needs to be using anticompetitive means to maintain that monopoly. The recent "textbook" example was of Microsoft, which in the 90s used its dominant position to charge computer manufacturers for a Windows license for each computer sold, reg…

It's got a monopoly on "search ads" by far.

Re: Only Google is really allowed to crawl the web

#213

Earlier quoted context omitted.

Aren't there anti trust laws to prevent this kind of thing?

Antiturst laws are hard to enforce in the United States. Monopolies themselves aren't illegal. To be convicted of an antitrust violation, a firm needs to both have a monopoly and needs to be using anticompetitive means to maintain that monopoly. The recent "textbook" example was of Microsoft, which in the 90s used its dominant position to charge computer manufacturers for a Windows license for each computer sold, reg…

It's not that hard, we're just out of practice due to the absurd Borkist economic theories we've been operating under for 40+ years. The laws are all there if the head of the DOJ antitrust division has the gumption to go reverse some bad precedents.

> In the realm of travel planning (to pick an example from these comments), it's barely a blip.

They used their monopoly in web search to gain non-negligible marketshare in entirely unrelated industry. That's text book anti-competitive behavior.

Google can argue whatever they want, but the argument that they're enabling other businesses is a bad one. It casts Google as a private regulator of the economy, which is exactly what antitrust laws are intended to deal with.

Re: Only Google is really allowed to crawl the web

#214

Earlier quoted context omitted.

The current anti-trust doctrine in the US has a goal of protecting consumers - not competition. What Google is doing is arguably great for consumers but awful to their competitors/other organizations. Technically, companies can simply block Google using robots.txt - but in reality that will lose them more money than the current partial disintermediation by Google is costing them - and Google knows this. It's a tall o…

consumers are in this case the advertisers. google has a monopoly on search ads and does enforce it, being a drain on the economy since in many fields you only succeed if you spend on search ads

> consumers are in this case the advertisers.

If someone could convince the courts that this is correct, then I'm sure Google would lose. However, I bet dollars to donuts Google's counter-arguement would be that the people doing the searching and quickly finding information are also consumers, and they outnumber advertisers and may be harmed by any proposed remediation in favor of advertisers.

Re: Only Google is really allowed to crawl the web

#215

Earlier quoted context omitted.

Antiturst laws are hard to enforce in the United States. Monopolies themselves aren't illegal. To be convicted of an antitrust violation, a firm needs to both have a monopoly and needs to be using anticompetitive means to maintain that monopoly. The recent "textbook" example was of Microsoft, which in the 90s used its dominant position to charge computer manufacturers for a Windows license for each computer sold, reg…

Is web search even a "market" independent of ads?

yes

Re: Only Google is really allowed to crawl the web

#216
post #93
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

So from the website's point of view there is no difference between 'crawling' and 'scraping'. Census.gov I assume has a ton of very useful information which is in the public domain which a host of potential companies could monetize by regularly scraping census.gov. Census.gov's purpose to make this information available to people is served by google, yahoo and bing. On the other hand if I have a business which is bas…

In the case of Census.gov, they offer an API to get the data[0]. It's actually pretty nice. Stable, ton of data, fairly uniform data structure across the different products. Very high rate limits, considering most data only needs retrieved once a year. I think they understand the difference between crawling and scraping.

[1] https://www.census.gov/data/developers.html

Re: Only Google is really allowed to crawl the web

#218
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I've noticed that sometimes Google had updated flight information before the displays at the airport.

Re: Only Google is really allowed to crawl the web

#219

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

In the early 90s there were various nascent systems for essentially public database interfaces for searching

The idea was that instead of a centralized search, people could have fat clients that individually query these apis and then aggregate the results on the client machine.

Essentially every query would be a what/where or what/who pair. This would focus the results

I really think we need to reboot those core ideas.

We have a manual version today. There's quite a few large databases that the crawlers don't get.

The one place for everything approach has the same fundamental problems that were pointed out 30 years ago, they've just become obvious to everybody now.

Re: Only Google is really allowed to crawl the web

#220
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They're making it easier to search for flights and arrange a trip. It's UX and makes me not hate the airlines/travel process as much. And I end up buying the flight from the airline anyways, and in many cases doing the arranging on the airline site in the end once it's determined, so Google is giving that back. They're not taking stuff from the airlines, I mean what ads and stuff are on the airline sites anyways spec…

You're talking about Google Flights, which is completely unrelated to flight status.
Post reply on HN