Maybe a naïve question but what prevents Knuckleheads’ from ignoring the robots.txt and crawl the side anyway? And if it's so easy to do, how does Google have a monopoly on crawling then?
Only Google is really allowed to crawl the web
331–340 of 365 posts
Re: Only Google is really allowed to crawl the web
#332Earlier quoted context omitted.
Is it legal for a government entity to issue a robots.txt like that? Maybe the line between use and abuse hasn't been delinated as well as it needs to be.
> Is it legal for a government entity to issue a robots.txt like that? I may be wrong (this isn't my area), but I was under the impression that robots.txt was just an unofficial convention? I'm not saying people should ignore robots.txt, but are there legal ramifications if ignored? I'm not asking about techniques sites use to discourage crawlers/scrapers, I'm specifically wondering if robots.txt has any legal weight…
Looks like Google is trying to turn it into an RFC though.
Re: Only Google is really allowed to crawl the web
#333Earlier quoted context omitted.
They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers. Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they…
Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.
Re: Only Google is really allowed to crawl the web
#334Earlier quoted context omitted.
Fine, the content is free. But if your crawlers want access to the content, then pay! Simple as that.
Will it be a flat fee, so that I, a lowly one-man crawler developer will not be able to afford it? Will it be that only Google can afford it, thus making their monopoly position even stronger? Is there a Wikipedia crawling "welfare" program if I'm not a trillion dollar mega company?
Re: Only Google is really allowed to crawl the web
#335On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…
Re: Only Google is really allowed to crawl the web
#336The way things work in practice, much of the web tries to prevent any type of programmatic access though a combination of edge-tech and robots.txt policies.
Content producers are addicted to Google and block competition, even though it is killing them. As Shakespeare put it: "Like rats that ravin down their proper bane, A thirsty evil; and when we drink we die."
Re: Only Google is really allowed to crawl the web
#337Earlier quoted context omitted.
Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.
You can screwed any time you book a connecting flight on two different airlines even if the times aren't tight. For instance if one is cancelled. If you use the same airline they will make sure you get to the destination.
Re: Only Google is really allowed to crawl the web
#338Earlier quoted context omitted.
Does the concergie of a hotel take anything away when he informs you that your flight has been delayed?
It's hard for me to make that an apt analogy. She's not well known as a portal to find websites, which is what Google had been for most of its existence. It's pretty difficult to come up with a non-computer analogy for how Google works now. Pick a different space, and the power imbalance is quite clear. If they wanted, they could destroy StackExchange very quickly with these widgets.
Re: Only Google is really allowed to crawl the web
#339- google will still have some "extra" crawling to keep his monopoly
- everyone else would be fighting for an access to said cache, which will not be able to carry everyone who wishes. So, there will be rationing and favors
- that will quickly become a bureaucracy which would envelop every aspect of internet activity
- "undesirable" sites will be easily ejected from cache and forgotten forever
- you will end up with a single entity paid by public, can't go bankrupt, in charge of whole internet. Google will go bankrupt if they don't satisfy people - these people don't even have to do a good job - there's nobody else on the market (remember, we started with the need to eject google? This will eventually happen).
Re: Only Google is really allowed to crawl the web
#340Earlier quoted context omitted.
> You see this recently with Wikipedia. Google's widgets have been reducing traffic to Wikipedia pretty dramatically. Enough so that Wikipedia is now pushing back with a product that the Googles of the world will have to pay for. Do you have a link of that product/service from Wikipedia?
https://diff.wikimedia.org/2021/03/16/introducing-the-wikime... Discussed not long ago: https://news.ycombinator.com/item?id=26484080