Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

281–290 of 365 posts

Re: Only Google is really allowed to crawl the web

#281

Earlier quoted context omitted.

Common Crawl is attempting to offer this as a non-profit: https://commoncrawl.org

o/t but what the hell are they doing to scroll on that page? I move my fingers a centimeter on my trackpad and the page is already scrolled all the way to the bottom. Hijacking scroll like this is one of the biggest turnoffs a website can have for me, up there with being plastered with ads and crap. It's ok imo in the context of doing some flashy branding stuff (think Google Pixel, Tesla splashes) but contentful page…

Add *##+js(aeld, scroll) to your uBO filters. That will stop scroll JS for all websites.

Re: Only Google is really allowed to crawl the web

#282

Earlier quoted context omitted.

Be careful when using Google Flight, last time I checked they use significantly less margins between flights so it’s shorter trips but much riskier.

You can screwed any time you book a connecting flight on two different airlines even if the times aren't tight. For instance if one is cancelled. If you use the same airline they will make sure you get to the destination.

> even if the times aren't tight

Depending on the definition of "tight" each of us have. I remember having 40mins in Munich, and that is a BIG airport. Especially if you disembark on one side of the terminal and your flight is on the far/opposite end. That's 25-30mins brisk walking. With 5000 people in-between you could as well miss your flight. No discussion about stopping to get a coffee or a snack.. you'll miss your flight.

Re: Only Google is really allowed to crawl the web

#283
post #202

Earlier quoted context omitted.

So, one more reason to hate Cloudflare and every single website that uses it.

Or maybe don’t “hate” folks who are just trying to put some content online and don’t want to deal with botnets taking down their work? You know, like what the internet was intended for.

> don’t want to deal with botnets taking down their work

Botnets and automated crawling are completely different things. This isn't about preventing service degradation (even if it gets presented that way). It's an attempt by content publishers to control who accesses their content and how.

Cloudflare is actively assisting their customers to do things I view as unethical. Worse, only Cloudflare (or someone in a similarly central position) is capable of doing those things in the first place.

Re: Only Google is really allowed to crawl the web

#285
post #18

Earlier quoted context omitted.

A company I worked for ~7 years ago ran its own focused web crawler (fetching ~10-100m pages per month, targeting certain sections of the web). There were a surprising number of sites out there that explicitly blocked access to anyone but Google/Bing at the time. We'd also get a dozen complaints or so a month from sites we'd crawled. Mostly upset about us using up their bandwidth, and telling us that only Google was…

Isn't that the website owners right though? I'm not sure I understand the problem here. If Google is taking traffic and reducing revenue, a company can deny in robots.txt. Google will actually follow those rules - unlike most others that are supposedly in this 2nd class.

> Isn't that the website owners right though?

No. The internet is public. Publishers shouldn't get any say in who accesses their content or how they do it. As far as I'm concerned, the fact that they do is a bug.

Re: Only Google is really allowed to crawl the web

#286
post #18

Earlier quoted context omitted.

A company I worked for ~7 years ago ran its own focused web crawler (fetching ~10-100m pages per month, targeting certain sections of the web). There were a surprising number of sites out there that explicitly blocked access to anyone but Google/Bing at the time. We'd also get a dozen complaints or so a month from sites we'd crawled. Mostly upset about us using up their bandwidth, and telling us that only Google was…

I usually recommend setting only Google/Bing/Yandex/Baidu etc to Allow and everything else to Disallow. Yes, the bad bots don't give a fuck, but even the non-malicious bots (ahrefs, moz, some university's search engine etc) don't bring any value to me as a site owner, take up band width and resources and fill up logs. If you can remove them with three lines in your robots.txt, that's less noise. Especially universiti…

> they're hammering your server

Why can't you just ratelimit IPs that are "too active" for your server to handle?

Re: Only Google is really allowed to crawl the web

#287
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

Standardized interoperability enables overall progress. Every airline doesn't need their own webpage. They could all provide a standard API.

"Every airline doesn't need their own webpage. They could all provide a standard API."

That's sort of how it works in the corporate booking tool world. It is decidedly not a better experience for end users, IMO.

There's quite a lot about each airline that is different, so any unified approach is a lowest common denominator. You'll notice things like loyalty points, for example, have more rich data on the airline's website. And that some fares are ONLY on the website. Or that seat maps have more useful detailed info, etc.

And that's all shopping/booking. Departure control, flight status, upgrade/downgrade, check-in, seat upgrades, standby, etc, are for the most part only on the airline's website.

Re: Only Google is really allowed to crawl the web

#288

Earlier quoted context omitted.

Right, then they shouldn't be effected by the rate-limiting, as long as its reasonable. If it was applied evenly to all clients/crawlers, it'd at least allow the possibility for a respectful, well designed crawler to compete.

The problem is, if you own a website, it takes the same amount of resources to handle the crawl from Google and FooCrawler even if both are behaving, but I'm going to get a lot more ROI out of letting Google crawl, so I'm incentivized to block FooCrawler but not Google. In fact, the ROI from Google is so high I'm incentivized to devote extra resources just for them to crawl faster.

We know that. No one claims websites are doing this for no reason. It's explicitly written in the article.

But this sub-thread is about misbehaved crawlers.

Re: Only Google is really allowed to crawl the web

#289
post #202

Earlier quoted context omitted.

Or maybe don’t “hate” folks who are just trying to put some content online and don’t want to deal with botnets taking down their work? You know, like what the internet was intended for.

Internet was certainly not intended for centralization. I hit Cloudflare captchas and error pages so often it's almost sickening. So many things are behind Cloudflare, things you least expect to be behind Cloudflare.

It's easy enough to bypass most Cloudflare “anti-bot” with an unusual refresh pattern or messing with a cookie. (It's easier to script this than solve the CAPTCHAs.)

Re: Only Google is really allowed to crawl the web

#290
post #5

The idea of a public cache available to anyone who wishes to index it is ... kind of compelling. If it was the only indexer allowed, and it was publically governed, then enforcing changes to regulation would be a lot more straightforward. Imagine if indexing public social media profiles was deemed unacceptable, and within days that content disappeared from all search engines. I don't think it'll ever happen, but it's…

Here's an idea... what if search became a peer-to-peer standardized protocol that is part of the stack to complement DNS? E.g. instead of using DNS as the primary entry point, you use a different protocol at that level to do "distributed search". DNS would still play a role too, but if "search" was a core protocol, the entry point for most people would be different. Similar to some of the concepts of "Linked Data", m…

https://en.wikipedia.org/wiki/YaCy
Post reply on HN