Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

51–60 of 365 posts

Re: Only Google is really allowed to crawl the web

#51
post #3

https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…

Are there any actual repercussions for just ignoring robots.txt?

Re: Only Google is really allowed to crawl the web

#52
post #47

Earlier quoted context omitted.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

> run [by] a private company and accessed for a small fee That is exactly the opposite of a public cache.

Not really. It serves the same function. Either you pay this hypothetical company or ??? pays to keep up the public one.

Re: Only Google is really allowed to crawl the web

#53
post #50
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

>However, that sort of search used to (most times) lead to a visit to the airline web site. I don't think that's correct. In the old days you'd either call a travel agent or use an aggregator like expedia. Google muscles out intermediaries like Expedia, Yelp, and so on. It's not likely much better or worse for the end user or supplier. Just swapping one middleman for another.

I can't prove it was that way, but I spent a lot of time in the space. For a long time, the airline's site used to be the top organic result, and there was no widget. Similar for other travel related searches (not just airlines) over time. Google has been pushing down organic results in favor of ads and widgets for a long time...and slowly, one little thing at a time. Like no widgets -> small widget below first organic result -> move the widget up -> make it bigger -> etc.

Re: Only Google is really allowed to crawl the web

#54
post #8

Earlier quoted context omitted.

Seriously? Google is a private cache of the web. That is the problem.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

I don't think you're quite clear on what the words "public" and "private" mean. "Public" is not a synonym for "run by the government" and "private" is not a synonym for "closed to everyone but the owner". Restaurants, for example, are generally open to the public, but they are not public. A restaurant owner is, with a few exceptions, free to refuse service to anyone at any time.

If it's "exactly the same as a public cache" then it's public, even if it is managed by a private company. The difference is not in who has access, the difference is in who decides who has access.

Re: Only Google is really allowed to crawl the web

#55
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…

Aren't there anti trust laws to prevent this kind of thing?

Re: Only Google is really allowed to crawl the web

#56
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

The way to look at this from Google’s point of view is to realise that most websites are slow and bad[1], so if Google sent you there you would have a bad experience with a bad slow website trying to find the information you want. Google want to make it better for you. [1] it feels like Google have contributed a lot to websites being slow and bad with eg ads, amp, angular, and probably more things for the other 25 le…

[deleted]

Re: Only Google is really allowed to crawl the web

#57
post #8

Earlier quoted context omitted.

Seriously? Google is a private cache of the web. That is the problem.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

You can API google search results to make a meta-search engine if you want to but it's like $5 / 1k requests.

Re: Only Google is really allowed to crawl the web

#58

This has been submitted to HN quite a few times. https://news.ycombinator.com/item?id=25426662 (Most comments; 11 comments) https://news.ycombinator.com/item?id=25417067 (3 comments) https://news.ycombinator.com/item?id=25546867 (Most recent; 89 days ago) https://news.ycombinator.com/item?id=25543859 https://news.ycombinator.com/item?id=25424852

[deleted]

Re: Only Google is really allowed to crawl the web

#59

Even that won't change much. There is no way Google can be out-googled by other search engines because of its market dominance: more traffic means more clicks, more clicks mean better search results, better search results will drive more traffic. I try bing and DDG for a week or so every 6 months. I always switch back to google eventually because the results are so much better. Google can only be disrupted if somethi…

Yup. My opinion has long been that the only thing that will take down google is a massive increase in NLP, such that the historical click data can be outperformed by a straight up really good NLP model

That's interesting. Is anyone working on this already? SV startup? And: don't you think Google is in the best position to build such a thing?

Re: Only Google is really allowed to crawl the web

#60
post #37

Earlier quoted context omitted.

That's incorrect. Before the search oligopolies formed, new search engines could start up. There was excite, hotbot, altavista, and more. Now they don't have access. Search these comments for census.gov.

There are companies that do pretty well in this space, like ahrefs, for example. They do resort to trickery, like proxy clients that look like home computers or cell phones. But, if a small entity like ahrefs can do it, anyone can do it. In a nutshell, though, I don't see equal access for all crawlers changing anything. Maybe that's the first barrier they hit, but it isn't the biggest or hardest one by far. Bing has…

[deleted]
Post reply on HN