Earlier quoted context omitted.
Do you not realize the more fundamental problem with you, as a company, essentially being the one who gatekeeps crawler access to the web?
Customers pay for this as a feature. Why would they feel it's a fundamental problem? There's nothing that says admins need to let you crawl their site.
Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
121–130 of 294 posts
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#122Earlier quoted context omitted.
For instance there is no way for distributed search engines to work with CloudFlare. No, "contact me and we'll help" is not always a solution.
That response is just a way to move the discussion out of the public domain without actually addressing it. It’s a scam.
That said, I don't think it's a good situation that this is the solution rather than a proper, documented position that people can work to.
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#123Earlier quoted context omitted.
Customers pay for this as a feature. Why would they feel it's a fundamental problem? There's nothing that says admins need to let you crawl their site.
If people intend this to happen, sure . But how many people who put their sites behind Cloudflare is aware that this might be a side effect?
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#124Earlier quoted context omitted.
If you are building a search engine and getting blocked you can always contact me and I'll make sure that the teams that work on bot detection and DDoS are aware. We would like to know because we should not be blocking a legit crawler like this.
For every developer that sees this message, a few dozen will have given up.
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#125Cloudflare is both a great thing and a terrible thing that has happened to the internet in recent years. Great in that they have a fantastic UI to add your site in, basically shielding the average user from attacks. Bad from a standpoint of that now only Google, Bing, and maybe other big search engines have the capabilities to actually crawl the internet now. I don't see us getting a massive innovation in search on t…
Maybe I'm cynical, but I don't see any innovation to be done in search. Google results have become much less useful over the past few years. If they cannot solve search with basically unlimited resources, how is a tiny company going to? 1. Filtering ever increasing trillions of spam/clickbait pages 2. Figuring out which results are useful information vs corporates trying to sell something. Those problems are not solv…
Nice pasta, can't believe someone would use it unironically.
Maybe I'm cynical, but I don't see any progress to be done in government transparency. USA's government has become way more overreaching and much less transparent over the decades. If USA cannot solve this issue, how a small country going to?
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#126Of course this idea raises privacy and authority concerns, similar to certificate authorities.
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#127Earlier quoted context omitted.
Hi! Thanks for taking the time to reply. You mentioned about "legit" crawlers, what defines a "legit" crawler in the eyes of Cloudflare, and what happens when Cloudflare suddenly decides it does not want to honour that "agreement"? What happens if/when Cloudflare is sold, or the contact who greenlit these smaller "legit" crawlers moves on and decides that it no longer agrees with said website anymore? Is a price comp…
There are a lot of hypotheticals here. I think you'll convince CloudFlare, and their customers, if you could name names and mention specific examples? If you are a price comparison site getting blocked by cloudflare, site owners may be losing sales, and that's good feedback.
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#128Cloudflare is both a great thing and a terrible thing that has happened to the internet in recent years. Great in that they have a fantastic UI to add your site in, basically shielding the average user from attacks. Bad from a standpoint of that now only Google, Bing, and maybe other big search engines have the capabilities to actually crawl the internet now. I don't see us getting a massive innovation in search on t…
I don't see us getting a massive innovation in search on the internet now that Google has such a massive foothold, and companies like Cloudflare stop innovation from happening. How are we "stopping search innovation"?
No such filters should exist. Is it really that awfully bad without anything but basic filters (ban an IP for flooding)? Are there like, operations that try and spam every single Cloudflare-hosted website 24/7?
Legitimately curious if your anti-bot measures come from actual bad experience with the internet or is it just a liability limitation move? (Namely to reduce potential suing surface by angry data owners and/or three-letter agencies.)
---
Basically, if I am experimenting with a basic crawling program and I hit websites A and B 20 times each in a space of one hour, is that really deserving of a captcha or extra auth methods?
Not flaming but I am really curious. Do you have any data and rationale posted somewhere that go into deeper detail about why Cloudflare's bot detection is how it is?
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#129Earlier quoted context omitted.
I do not understand what the parent means by a "distributed search engine" and I do not know what problem they are facing.
A search engine which is not run centrally by one organization on infrastructure in a known network, but rather something like YaCy where individual users run crawler nodes on networks that vary over time. Which makes "contact us for an exception" a no-go, as the relevant source IPs will constantly be changing.
Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea
#130Earlier quoted context omitted.
Say I’m interested in building a small scale domain-specific search engine and only just started development. There’s no prototype yet and may never be. In this situation, how do you determine it’s a legit crawler? And what about crawlers with even more limited scopes (targeting only a handful of sites) that they can’t possibly be called search engines? Are they ever considered legit?
Be a good netizen? Respect robots.txt. Don't lie in your User-Agent. Don't crawl at a ridiculous rate. All those are a good starting point.