On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…
Only Google is really allowed to crawl the web
181–190 of 365 posts
Re: Only Google is really allowed to crawl the web
#182I'd like to see some data on their claim that website operators are giving googlebot special privileges. As far as I can tell it would be a huge pain in the ass to block crawler bots from my servers, not that I've tried. I have some weird pages that tend to get crawlers caught in infinite loops, and I try to give them hints with robots.txt but most of the bots don't even respect robots.txt. If I actually wanted to re…
As inflammatory as the headline of the page looks, they literally admit it's not google's fault in the smaller text lower down:
"This isn’t illegal and it isn’t Google’s fault, but"
Re: Only Google is really allowed to crawl the web
#183Earlier quoted context omitted.
Aren't there anti trust laws to prevent this kind of thing?
The current anti-trust doctrine in the US has a goal of protecting consumers - not competition. What Google is doing is arguably great for consumers but awful to their competitors/other organizations. Technically, companies can simply block Google using robots.txt - but in reality that will lose them more money than the current partial disintermediation by Google is costing them - and Google knows this. It's a tall o…
google has a monopoly on search ads and does enforce it, being a drain on the economy since in many fields you only succeed if you spend on search ads
Re: Only Google is really allowed to crawl the web
#184Earlier quoted context omitted.
I'm generally anti business. But I have to disagree. "The Public" that the government serves includes businesses. Businesses (ignoring corporate personhood bullshit) are owned and operated by people. I do not want the government deciding "what purposes" e.g. non-commercial, serve the public good. The public gets to decide that. (charging a license for commercial use is maybe ok (assuming supporting that use costs gov…
A specific case where this favorite-picking by government enables corruption: https://en.wikipedia.org/wiki/Nationally_recognized_statisti... And an example from the quickly-approaching future, when there will be Nationally Recognized Media Organizations who license "Fact-Checkers," through which posts to public-facing will have to be submitted for certification and correction.
Re: Only Google is really allowed to crawl the web
#185I tried to set up YaCy [1] at home to index a few of may favorite smaller websites, so I could quickly search just them. That turned out to be a bad idea. Some ended up blocking my home IP address and others reported me to my ISP. None of these sites were that large, and I wasn't continuously crawling them... [1] https://yacy.net/
Re: Only Google is really allowed to crawl the web
#186Earlier quoted context omitted.
Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…
Wikimedia recently announced Wikimedia Enterprise for "organizations that want to repurpose Wikimedia content in other contexts, providing data services at a large scale". So they're pretty clearly looking to monetize organizations which consume their data in a for-profit context.
You could e.g. just cover operational cost and/or improve the service quality from it.
Re: Only Google is really allowed to crawl the web
#187The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…
Similar to some of the concepts of "Linked Data", maybe - https://en.wikipedia.org/wiki/Linked_data.
The problem is getting to a standard, it would essentially need to be federated search so a standard would have to be established (de facto most likely).
Also, indexes and storage, distribution of processing load.. peer-to-peer search is already a thing, but it doesn't seem to be a core function of the Internet.
This is basically the same concept as making an "open" version of something that is "closed" in order to compete, I guess.
Re: Only Google is really allowed to crawl the web
#188Earlier quoted context omitted.
Aren't there anti trust laws to prevent this kind of thing?
Just tell people to stop using google. Go direct.
Enough people realizing Google is trapping and cannibalizing traffic to the other sites it feeds off of, and choosing to do other things EXCEPT touching Google properties, is THE ONLY way they'll be unseated.
No clear legal path to stop a bully means it's an ethical / habit path.
Not saying there's any easy way, just that this is it.
Re: Only Google is really allowed to crawl the web
#189Earlier quoted context omitted.
Aren't there anti trust laws to prevent this kind of thing?
Antiturst laws are hard to enforce in the United States. Monopolies themselves aren't illegal. To be convicted of an antitrust violation, a firm needs to both have a monopoly and needs to be using anticompetitive means to maintain that monopoly. The recent "textbook" example was of Microsoft, which in the 90s used its dominant position to charge computer manufacturers for a Windows license for each computer sold, reg…
Re: Only Google is really allowed to crawl the web
#190Earlier quoted context omitted.
Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. 2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned. The instant Google becomes confident they can supplant Wikipedia, they will.
> Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. Google isn't a sponsor of Mozilla, they're a customer. Do people think Google is "sponsoring" Apple with $1.5 billion a year too?
These are two very different companies with a very different relationship with Google. And very different influences on Google.
Google wants to be on iOS. It brings customers to Google. A lot of them. iOS is possibly more profitable to Google than Android even with all the payments Apple extracts from them.
Google needs Mozilla so that Google may pretend that there's competition in browser space and that they don't own standards committees. The latter already isn't really true, and Google increasingly doesn't care about the former.