Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

21–30 of 365 posts

Re: Only Google is really allowed to crawl the web

#21
I've definitely scraped by this problem on several occasions. Recently I was writing a tool to check outgoing links from my site, to see which sites are offline (it's called notfoundbot). What I found was that many sites have "DDoS Protection" that makes such an effort impossible, other sites whitelist the CuRL headers, others like it when you pretend to be a search engine.

Basically writing some code that tests whether "a website is currently online or offline" is much, much harder than you think, because, yep, the only company that can do that is Google.

Re: Only Google is really allowed to crawl the web

#22
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.

Re: Only Google is really allowed to crawl the web

#23
If the shared cache ever became significant enough to matter it would be devastated by marketers, scammers and other abusers. Google employs the groomers that make their index at least tolerable, if still clearly imperfect. Without that cadre of well compensated expertise to win the arms race against such abusers the scheme is not feasible.

I suppose this could be crowdsourced if I didn't know about politics and how any attempt at delegating the responsibility for blessing sites and their indexes would become a controversy. Google takes lots of heat about its behavior already, but Google is a private entity and can indulge its private prerogatives for the most part. Without that independence this couldn't function.

Re: Only Google is really allowed to crawl the web

#24
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.

I don't understand that. The crawling access is mostly the same as it ever was. Google's SERP pages are not. A mutually beneficial search engine that respects it's sources would still crawl the same. Google used to be that.

The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."

Re: Only Google is really allowed to crawl the web

#25
post #8
post #4

Seems like a private cache of the web would solve the problem? Why does it need to be public?

Seriously? Google is a private cache of the web. That is the problem.

Google doesn’t give anyone access to said cache. I mean one crawler with a shared api among competitors. So exactly the same as the public cache, but run my a private company and accessed for a small fee

Re: Only Google is really allowed to crawl the web

#26
While I don’t disagree with the idea that all crawlers should have equal access, we also need to address the quality of many crawlers.

Google and Microsoft have never hammered any website I’ve run into the ground. Crawlers from other other, smaller, search engines have, to the point where it was easier to just block them entirely.

Part of the problem is that sites want search engine to index their site, but not allow random people just scrapping the entire site. So they do the best they can, and forget that Google isn’t the web. I doubt it’s shady deals with Google, it’s just small teams doing the best they can and sometimes they forget to think ideas through, because it’s good enough.

Re: Only Google is really allowed to crawl the web

#27
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

They are not just taking away internet traffic, but in the flights example, they actually acquired an aggregate flight/travel company and so they are actually entering markets and competing with their own ad customers.

Then it comes fully circle to Google unfairly using their market position vis-a-vis data, search and advertising. It’s a win-win Google lets the data dictate which markets to enter and on one hand they can jack up advertising fees on customers/competitors and unfairly build their own service into search above both ads and organic results.

Re: Only Google is really allowed to crawl the web

#28
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

Imagine if Donald Trump decided that indexing Joe Biden's campaign site was unacceptable.

A mandated singular public cache has potential slippery slopes.

Re: Only Google is really allowed to crawl the web

#29
post #24

Earlier quoted context omitted.

That's the result of the crawling, and it preventing competition. Google would much prefer that people complain about the details while ignoring the root cause.

I don't understand that. The crawling access is mostly the same as it ever was. Google's SERP pages are not. A mutually beneficial search engine that respects it's sources would still crawl the same. Google used to be that. The core problem is incentives: http://infolab.stanford.edu/~backrub/google.html "we believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search…

That's incorrect. Before the search oligopolies formed, new search engines could start up. There was excite, hotbot, altavista, and more. Now they don't have access. Search these comments for census.gov.

Re: Only Google is really allowed to crawl the web

#30
I'm not sure if it is a good thing if there is a public cache of everything that Google has. The issue is websites will simply stop serving content to Google to protect their content from being accessed by their competitors, this in turn will make search much worse and will force us back to the pre-search dark ages of the internet. The sites may even serve an even more crippled version of their content just to get hits but there is no doubt search quality will suffer.

We're left with a monopoly that is Google, destroying it now could be foolish.

Post reply on HN