Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

201–210 of 365 posts

Re: Only Google is really allowed to crawl the web

#202

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

So, one more reason to hate Cloudflare and every single website that uses it.

Or maybe don’t “hate” folks who are just trying to put some content online and don’t want to deal with botnets taking down their work? You know, like what the internet was intended for.

Re: Only Google is really allowed to crawl the web

#203
dupe/posted earlier etc

I also got confused about this page as there's another project of theirs around right now about RIP Google Reader that's on a seperate domain...

Funny a site that's all about google this and that doesn't have clear URL/pages for their articles that can be linked to easily geez

Original post/discussion from the source, 3 months ago: https://news.ycombinator.com/item?id=25417067

Re: Only Google is really allowed to crawl the web

#204

Earlier quoted context omitted.

That's true, but it can save you a ton of money. You just have to be aware of the risks and plan accordingly. I have typically used this strategy when flying back to the US from the EU. Take an EZJet or similar low cost airline from random small EU city to a larger EU city like Paris, London, Frankfurt, etc... and book the return trip to the US from the larger city. I've also been forced to do this from some EU citie…

Yeah this strategy is good, but you need to allow a long layover like 6 hours if you have to go through immigration and change airports for the connection which happens pretty often with ryanair and ezjet. It’s a big pain, but it does save money.

If you're booking each leg with different carrier, I find it best to pay the little extra with kiwi.com and they give you guarantee for the connection. I missed connection twice and they always got me on the next flight to the destination for free.

Re: Only Google is really allowed to crawl the web

#205

Earlier quoted context omitted.

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

Google was/is also the largest sponsor of Mozilla. This doesn't stop Google from sabotaging Mozilla. 2 mln is probably Google's hourly profit. For that they get one of the biggest knowledge bases in the world. It's basically free as far as Google is concerned. The instant Google becomes confident they can supplant Wikipedia, they will.

NOT a sponsor of Mozilla. Google buys web traffic (as default search engine) for ~$300M and turns it into several times that $ in ad revenue.

Re: Only Google is really allowed to crawl the web

#206
post #137

> Only a select few crawlers are allowed access to the entire web, and Google is given extra special privileges on top of that. Hmm, so set up a VPN on the Google Cloud so you have a Google IP address, use a Google User-Agent, and go!

https://developers.google.com/search/docs/advanced/crawling/...

describes the procedure for checking "is this source Google it". You couldn't fake it just by running on gcp

Re: Only Google is really allowed to crawl the web

#207
post #69

Earlier quoted context omitted.

Imagine if Donald Trump decided to tax campaign donations to Joe Biden's campaign at 100%. I am unconvinced by the "slippery slope" argument being deployed by default to any governmental attempt to combat tech monopolies.

This is an argument against centralization more than it is against government. "One index to rule them all" seems more fraught with difficulty than, "large cloud providers are unhappy that crawlers on the open web are crawling the open web".

If the impact stopped at "large cloud providers" being unhappy, I think that you're correct. But I think we've seen considerably downstream "difficulty" for the rest of society from search essentially being consolidated into one private actor.

Re: Only Google is really allowed to crawl the web

#208

Earlier quoted context omitted.

I don't really understand your comment. Marketers, scammers and other abusers already publish to the web with the intention to be included in a crawl. Postprocessing crawl data is already a thing. Assuming this hypothetical shared crawl cache were to exist, it does not preclude google (and all consumers of that cache) doing their own processing downstream of that cache. Does it? What's the new attack vector?

> I don't really understand your comment. If you don't then you fail to appreciate the amount of labor it takes to thwart bad actors from ruining indexes. Abusers do publish to the web, and we enjoy not wallowing in their crap because small army of experienced and expensive people at a select few Big Tech companies are actively shielding us from it. It's easy to anticipate the malcontent view; 'Google spends all its…

>Will the published caches be 99% crap

Yes. It will be exactly as crap as whatever's published on the web.

And the utility of google's search engine would be to perform their proprietary processing on top of the publicly-available crawl results. Analogous to how their search is already preforming proprietary processing on top of a crawl cache.

>If you don't then you fail to appreciate the amount of labor it takes to thwart bad actors from ruining indexes.

Did you miss the part where I said "Assuming this hypothetical shared crawl cache were to exist, it does not preclude google (and all consumers of that cache) doing their own processing downstream of that cache. Does it?"

Re: Only Google is really allowed to crawl the web

#209
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Wikipedia isn't monetized. Doesn't it benefit them if Google is serving their content for free and people are finding the information they want without having to hit Wikipedia?? And also, isn't Google the largest sponsor for Wikipedia already? In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2]. [1] https://techcrunch.com/2019/01/22/google-org-donates-2-milli... [2] https://en.wikipedia.org/wiki/Wi…

> Wikipedia isn't monetized.

No, but they often ask for donations when you visit the site, which people won't see if they just see the in-line blurb from Wikipedia on the Google results page.

> In 2019 - Google donated $2M [1]. In 2010, Google also donated $2m [2].

$2M is a pittance compared to what I expect Google believes is the value of their Wikipedia blurbs. If Wikipedia could charge for use of this data (which another commenter claims they are working on doing), they could easily make orders of magnitude more money from Google.

Of course, my expectation is that Google would rather drop the Wikipedia blurbs entirely, or source the data elsewhere, than pay significantly more.

Re: Only Google is really allowed to crawl the web

#210

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

On the other hand, I do not want my site to go down thanks to a few bad 'crawlers' that fork() a thousand http requests every second and take down my site, forcing me to do manual blocking or pay for a bigger server/scale-out my infrastructure. Why should I have to serve them?

You can use the same rate-limiting for all crawlers, Google or not.
Post reply on HN