Basically writing some code that tests whether "a website is currently online or offline" is much, much harder than you think, because, yep, the only company that can do that is Google.
Only Google is really allowed to crawl the web
21–30 of 365 posts
Re: Only Google is really allowed to crawl the web
#22The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
Re: Only Google is really allowed to crawl the web
#23I suppose this could be crowdsourced if I didn't know about politics and how any attempt at delegating the responsibility for blessing sites and their indexes would become a controversy. Google takes lots of heat about its behavior already, but Google is a private entity and can indulge its private prerogatives for the most part. Without that independence this couldn't function.
Re: Only Google is really allowed to crawl the web
#24The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.
The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."
Re: Only Google is really allowed to crawl the web
#25Seems like a private cache of the web would solve the problem? Why does it need to be public?
Seriously? Google is a private cache of the web. That is the problem.
Re: Only Google is really allowed to crawl the web
#26Google and Microsoft have never hammered any website I’ve run into the ground. Crawlers from other other, smaller, search engines have, to the point where it was easier to just block them entirely.
Part of the problem is that sites want search engine to index their site, but not allow random people just scrapping the entire site. So they do the best they can, and forget that Google isn’t the web. I doubt it’s shady deals with Google, it’s just small teams doing the best they can and sometimes they forget to think ideas through, because it’s good enough.
Re: Only Google is really allowed to crawl the web
#27The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…
Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they can jack up advertising fees on customers/competitors and unfairly build their own service into search above both ads and organic results.
Re: Only Google is really allowed to crawl the web
#28The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…
A mandated singular public cache has potential slippery slopes.
Re: Only Google is really allowed to crawl the web
#29Earlier quoted context omitted.
That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.
I don't understand that. The crawling access is mostly the same as it ever was. Google's SERP pages are not. A mutually beneficial search engine that respects it's sources would still crawl the same. Google used to be that. The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search…
Re: Only Google is really allowed to crawl the web
#30We're left with a monopoly that is Google, destroying it now could be foolish.