Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

241–250 of 365 posts

Re: Only Google is really allowed to crawl the web

#241
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I'm not sure I agree with this. I think airline websites are so garbage filled that they've driven people to use the simple alternative of the google flights checkout. It's a bit of a vicious cycle, but In general most websites are so chock filled with crap that not having to click into them for real is a relief!

Yeah, the Google flights issue is difficult. On one hand, the business practice is problematic. On the other hand, Google flights is so much better than its competitors it's ridiculous.

If there was a way to split Google flights into a separate company and somehow ensure it wouldn't devolve into absolute trash like its competitors, that would be a good thing.

Re: Only Google is really allowed to crawl the web

#242
I disallow scanning on all my projects. After GDPR I also removed all analytics - I realised it is just a time sink - instead of focusing on content I would often focus on getting the bigger numbers. I am not a marketer, so it didn't have much value to me and it would just enlarge Google dataset without any payment. I get that you cannot find my projects in the search engine. I am okay with that :-)

Re: Only Google is really allowed to crawl the web

#243
post #236

Earlier quoted context omitted.

I'm not sure I agree with this. I think airline websites are so garbage filled that they've driven people to use the simple alternative of the google flights checkout. It's a bit of a vicious cycle, but In general most websites are so chock filled with crap that not having to click into them for real is a relief!

It's not Google's prerogative to scrape a website and display its content, no matter how awful the website.

If 1 airline let me view information in a friendly fashion and the other didn't I would do business with the first.

Lest we forget the money in that scenario is from butts in seats not clicks on a website. The particular example is ill chosen as google is actually taking on a cost, taking nothing, and gifting the airline a better ui.

Re: Only Google is really allowed to crawl the web

#244
post #236

Earlier quoted context omitted.

I'm not sure I agree with this. I think airline websites are so garbage filled that they've driven people to use the simple alternative of the google flights checkout. It's a bit of a vicious cycle, but In general most websites are so chock filled with crap that not having to click into them for real is a relief!

It's not Google's prerogative to scrape a website and display its content, no matter how awful the website.

If you make an awful website that can be scrapped it's a matter of when not if someone will take your data and give it to your consumers whether your trying to upsell them or not...

Re: Only Google is really allowed to crawl the web

#245
post #241

Earlier quoted context omitted.

I'm not sure I agree with this. I think airline websites are so garbage filled that they've driven people to use the simple alternative of the google flights checkout. It's a bit of a vicious cycle, but In general most websites are so chock filled with crap that not having to click into them for real is a relief!

Yeah, the Google flights issue is difficult. On one hand, the business practice is problematic. On the other hand, Google flights is so much better than its competitors it's ridiculous. If there was a way to split Google flights into a separate company and somehow ensure it wouldn't devolve into absolute trash like its competitors, that would be a good thing.

It was ITA and prior to Google buying them, did a pretty good business selling backend flight shopping services to aggregators and airlines.

Shopping for flights is a surprisingly technically difficult thing to do well.

Re: Only Google is really allowed to crawl the web

#246

Earlier quoted context omitted.

You can use the same rate-limiting for all crawlers, Google or not.

Googlebot is pretty careful and generally doesn’t cause these problems.

Right, then they shouldn't be effected by the rate-limiting, as long as its reasonable. If it was applied evenly to all clients/crawlers, it'd at least allow the possibility for a respectful, well designed crawler to compete.

Re: Only Google is really allowed to crawl the web

#248
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

I'm not sure I agree with this. I think airline websites are so garbage filled that they've driven people to use the simple alternative of the google flights checkout. It's a bit of a vicious cycle, but In general most websites are so chock filled with crap that not having to click into them for real is a relief!

BA had some tracking request inline on the “payment processing” page which when blocked by my pihole prevents me from ever getting to the confirmation page, just have to refresh your email and wait for the best.

I have no idea how these companies, which make quite a decent amount of money at least up until 2020, can have such utterly poor sites.

I once counted some 20+ redirects on a single request during this process heh..

Re: Only Google is really allowed to crawl the web

#249

On a related note, Cloudflare just introduced "Super Bot Fight Mode" ( https://blog.cloudflare.com/super-bot-fight-mode/ ) which is basically a whitelisting approach that will block any automated website crawling that doesn't originate from "good bots" (they cite Google & Paypal as examples of such bots). So basically everyone else is out of luck and will be tarpitted (i.e. connections will get slower and slower unti…

How hard is it to ask Cloudflare to let you crawl?

It's not Cloudflare who is deciding it. It's the website owners who request things like "Super Bot Fight Mode". I never enable such things on my CF properties. Mostly it's people who manage websites with "valuable" content, e.g. shops with prices who desperately want to stop scraping by competitors.
Post reply on HN