Earlier quoted context omitted.
And unfortunately, cloudflare is everywhere. This trend will make it even harder for projects like a new search engine to enter the game.
Because if you don't have it some a-hole will go and ddos your site or you want to prevent a hug-of-death because of reasons. It seems a lot of issues happen because bad players are continued to allowed to thrive, example: everybody uses a big provider because they're the only ones that solved the spam issue.
Facebook was used as a proxy by web scraping bots
21–30 of 125 posts
Re: Facebook was used as a proxy by web scraping bots
#22You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…
Long term, a new HTTP META method would be interesting. I wonder if something like that has ever been considered. Providers like Cloudflare would hopefully be more lenient with these requests.
I don't think FAAANG (or any other big players) would have much interest in making it happen in the standard though, since it would undercut their big-player advantage.
Re: Facebook was used as a proxy by web scraping bots
#23Earlier quoted context omitted.
Because if you don't have it some a-hole will go and ddos your site or you want to prevent a hug-of-death because of reasons. It seems a lot of issues happen because bad players are continued to allowed to thrive, example: everybody uses a big provider because they're the only ones that solved the spam issue.
I use Zoho.com and I rarely get spam, if ever.
Re: Facebook was used as a proxy by web scraping bots
#24Earlier quoted context omitted.
Because if you don't have it some a-hole will go and ddos your site or you want to prevent a hug-of-death because of reasons. It seems a lot of issues happen because bad players are continued to allowed to thrive, example: everybody uses a big provider because they're the only ones that solved the spam issue.
I use Zoho.com and I rarely get spam, if ever.
Re: Facebook was used as a proxy by web scraping bots
#25Can you real-time crawl twitter ? Pretty sure they have a special deal with google to instant ping on new tweets.
How many websites actually ping google on new content ?
Re: Facebook was used as a proxy by web scraping bots
#26You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…
Re: Facebook was used as a proxy by web scraping bots
#27That's pretty interesting, Facebook as a "web scale / hundreds of pages per second" batch web page summarizer. I imagine you could build a pretty decent general purpose search engine that way...free crawler.
Re: Facebook was used as a proxy by web scraping bots
#28You can't create your own link previewer, cloudflare will put a captcha in front of every website. All I want is a a freaking tag. They don't seem eager to fix it either, their proposed solution is to contact every website owner (seriously) to ask them to whitelist you[1]. Frankly, i wish facebook or cloudflare offered their previewer as a free service, since most websites have them whitelisted. 1. https://community.…
Long term, a new HTTP META method would be interesting. I wonder if something like that has ever been considered. Providers like Cloudflare would hopefully be more lenient with these requests.
Accept: application/json
Would be a reasonable alternative? Wasn't this supposed to be the point of content negotiation?Re: Facebook was used as a proxy by web scraping bots
#29Earlier quoted context omitted.
Long term, a new HTTP META method would be interesting. I wonder if something like that has ever been considered. Providers like Cloudflare would hopefully be more lenient with these requests.
I wonder if Accept: application/json Would be a reasonable alternative? Wasn't this supposed to be the point of content negotiation?
I guess what this is really about is, I hate to say it, but something in the direction of the semantic web, where web servers (and in this case, CloudFlare et al) actually gain a deeper understanding of the content they serve, and a web browser / crawler being able to query this content directly.
Re: Facebook was used as a proxy by web scraping bots
#30Earlier quoted context omitted.
And unfortunately, cloudflare is everywhere. This trend will make it even harder for projects like a new search engine to enter the game.
Because if you don't have it some a-hole will go and ddos your site or you want to prevent a hug-of-death because of reasons. It seems a lot of issues happen because bad players are continued to allowed to thrive, example: everybody uses a big provider because they're the only ones that solved the spam issue.