Live data from Hacker News

Only Google is really allowed to crawl the web

knuckleheads.club

361–365 of 365 posts

Re: Only Google is really allowed to crawl the web

#361
post #279
post #180

I tried to set up YaCy [1] at home to index a few of may favorite smaller websites, so I could quickly search just them. That turned out to be a bad idea. Some ended up blocking my home IP address and others reported me to my ISP. None of these sites were that large, and I wasn't continuously crawling them... [1] https://yacy.net/

I have been running my own Searx instance in AWS for a while and have not gotten blocked yet anywhere

Wow, I've never heard of Searx - https://searx.space/

It's privacy friendly proxied search results, a lot like the Nitter project for Twitter. https://github.com/zedeus/nitter/wiki/Instances

Re: Only Google is really allowed to crawl the web

#362
post #12

The bigger problem, to me, is not around crawling. It's the asymmetrical power Google has after crawling. Google is obviously on a mission to keep people on Google owned properties. So, they take what they crawl and find a way to present that to the end user without anyone needing to visit the place that data came from. Airlines are a good example. If you search for flight status for a particular flight, Google prese…

>Google's widgets have been reducing traffic to Wikipedia pretty dramatically.

But wouldn't this be a good thing? Since wikipedia is a nonprofit aiming to provide knowledge, google stealing & caching their content might help serve information to more while reducing wikipedia's server load, so IMO it might not be so bad for wikipedia.

Re: Only Google is really allowed to crawl the web

#363
If you could understand, that what you see in a Google search is a manipulated reality, an algorithm created to their oun profit and believes, you would never look back, or would, if you was agent Smith. But my friend, all US companies are in the same row, pointing to the same direction, guided by the oracle of Inteligent Board of Psycopath Agents, who thinks are defending US nation. If you want to see what is ouside the chains of FBIntranet, you need be smart, all the easy things (tor, vpn...) are maded to catch lasy people.

Re: Only Google is really allowed to crawl the web

#364

Earlier quoted context omitted.

> they're hammering your server Why can't you just ratelimit IPs that are "too active" for your server to handle?

CGNAT

Now that's a really good point. I wonder why there isn't a standard protocol for signalling upstream that a particular connection is abusive and to please rate limit the path at the source on your behalf? It would certainly add complexity, but the current situation is hardly better.

Re: Only Google is really allowed to crawl the web

#365

Earlier quoted context omitted.

> they're hammering your server Why can't you just ratelimit IPs that are "too active" for your server to handle?

From some I could, but why would I? If they're not adding value and they don't want to behave, I don't see a reason to spend money to adapt my systems to be "inclusive" towards their usage patterns.

In context, you're justifying blocking all automated traffic, even that which does behave, by pointing out that some of it doesn't. That attitude seems lazy at best, malicious at worst.
Post reply on HN