https://knuckleheads.club/the-googlebot-monopoly/ has actual details. > Let’s take a look at the robots.txt for census.gov from October of 2018 as a specific example to see how robots.txt files typically work. This document is a good example of a common pattern. The first two lines of the file specify that you cannot crawl census.gov unless given explicit permission. The rest of the file specifies that Google, Micros…
Only Google is really allowed to crawl the web
51–60 of 365 posts
Re: Only Google is really allowed to crawl the web
#52Earlier quoted context omitted.
Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee
> run [by] a private company and accessed for a small fee That is exactly the opposite of a public cache.
Re: Only Google is really allowed to crawl the web
#53The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
>However, that sort of search used to (most times) lead to a visit to the airline web site. I don't think that's correct. In the old days you'd either call a travel agent or use an aggregator like expedia. Google muscles out intermediaries like Expedia, Yelp, and so on. It's not likely much better or worse for the end user or supplier. Just swapping one middleman for another.
Re: Only Google is really allowed to crawl the web
#54Earlier quoted context omitted.
Seriously? Google is a private cache of the web. That is the problem.
Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee
If it's "exactly the same as a public cache" then it's public, even if it is managed by a private company. The difference is not in who has access, the difference is in who decides who has access.
Re: Only Google is really allowed to crawl the web
#55The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…
Re: Only Google is really allowed to crawl the web
#56The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
The way to look at this from Google’s point of view is to realise that most websites are slow and bad[1], so if Google sent you there you would have a bad experience with a bad slow website trying to find the information you want. Google want to make it better for you. [1] it feels like Google have contributed a lot to websites being slow and bad with eg ads, amp, angular, and probably more things for the other 25 le…
Re: Only Google is really allowed to crawl the web
#57Earlier quoted context omitted.
Seriously? Google is a private cache of the web. That is the problem.
Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee
Re: Only Google is really allowed to crawl the web
#58This has been submitted to HN quite a few times. https://news.ycombinator.com/item?id=25426662 (Most comments; 11 comments) https://news.ycombinator.com/item?id=25417067 (3 comments) https://news.ycombinator.com/item?id=25546867 (Most recent; 89 days ago) https://news.ycombinator.com/item?id=25543859 https://news.ycombinator.com/item?id=25424852
Re: Only Google is really allowed to crawl the web
#59Even that won't change much. There is no way Google can be out-googled by other search engines because of its market dominance: more traffic means more clicks, more clicks mean better search results, better search results will drive more traffic. I try bing and DDG for a week or so every 6 months. I always switch back to google eventually because the results are so much better. Google can only be disrupted if somethi…
Yup. My opinion has long been that the only thing that will take down google is a massive increase in NLP, such that the historical click data can be outperformed by a straight up really good NLP model
Re: Only Google is really allowed to crawl the web
#60Earlier quoted context omitted.
That's incorrect. Before the search oligopolies formed, new search engines could start up. There was excite, hotbot, altavista, and more. Now they don't have access. Search these comments for census.gov.
There are companies that do pretty well in this space, like ahrefs, for example. They do resort to trickery, like proxy clients that look like home computers or cell phones. But, if a small entity like ahrefs can do it, anyone can do it. In a nutshell, though, I don't see equal access for all crawlers changing anything. Maybe that's the first barrier they hit, but it isn't the biggest or hardest one by far. Bing has…