Live data from Hacker News

Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

herrjemand.medium.com

121–130 of 294 posts

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#121
post #118

Earlier quoted context omitted.

Do you not realize the more fundamental problem with you, as a company, essentially being the one who gatekeeps crawler access to the web?

Customers pay for this as a feature. Why would they feel it's a fundamental problem? There's nothing that says admins need to let you crawl their site.

If people intend this to happen, sure . But how many people who put their sites behind Cloudflare is aware that this might be a side effect?

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#122
post #29

Earlier quoted context omitted.

For instance there is no way for distributed search engines to work with CloudFlare. No, "contact me and we'll help" is not always a solution.

That response is just a way to move the discussion out of the public domain without actually addressing it. It’s a scam.

jgrahamc is one of the most active HN users. He's in the top 20 on the "leaderboard". He has a good reputation. I'd hesitate to call this offer a scam.

That said, I don't think it's a good situation that this is the solution rather than a proper, documented position that people can work to.

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#123
post #121
post #118

Earlier quoted context omitted.

Customers pay for this as a feature. Why would they feel it's a fundamental problem? There's nothing that says admins need to let you crawl their site.

If people intend this to happen, sure . But how many people who put their sites behind Cloudflare is aware that this might be a side effect?

I would wager that most people that purchase Cloudflare are probably aware of the features it offers

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#124

Earlier quoted context omitted.

If you are building a search engine and getting blocked you can always contact me and I'll make sure that the teams that work on bot detection and DDoS are aware. We would like to know because we should not be blocking a legit crawler like this.

For every developer that sees this message, a few dozen will have given up.

Exactly this. Cloudflare actively blocks legit crawlers. It shouldn't be dependent on seeing some random hn comment from some random at cloudflare to get that fixed.

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#125
post #12

Cloudflare is both a great thing and a terrible thing that has happened to the internet in recent years. Great in that they have a fantastic UI to add your site in, basically shielding the average user from attacks. Bad from a standpoint of that now only Google, Bing, and maybe other big search engines have the capabilities to actually crawl the internet now. I don't see us getting a massive innovation in search on t…

Maybe I'm cynical, but I don't see any innovation to be done in search. Google results have become much less useful over the past few years. If they cannot solve search with basically unlimited resources, how is a tiny company going to? 1. Filtering ever increasing trillions of spam/clickbait pages 2. Figuring out which results are useful information vs corporates trying to sell something. Those problems are not solv…

Maybe I'm cynical, but I don't see any innovation to be done in {AREANAME}. {PRODUCTNAME} has become much less useful over the past few years. If {MONOPOLISTNAME} cannot solve {AREANAME} with basically unlimited resources, how a tiny company going to?

Nice pasta, can't believe someone would use it unironically.

Maybe I'm cynical, but I don't see any progress to be done in government transparency. USA's government has become way more overreaching and much less transparent over the decades. If USA cannot solve this issue, how a small country going to?

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#126
Could this be solved (in large part) if key makers like YubiKey did I.D. verification on purchase? Then, to do the type of "farming" that's mentioned in this article, you'd need to organize a large group of people to all buy the keys rather than just submit a bulk order to Alibaba.

Of course this idea raises privacy and authority concerns, similar to certificate authorities.

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#127
post #98
post #73

Earlier quoted context omitted.

Hi! Thanks for taking the time to reply. You mentioned about "legit" crawlers, what defines a "legit" crawler in the eyes of Cloudflare, and what happens when Cloudflare suddenly decides it does not want to honour that "agreement"? What happens if/when Cloudflare is sold, or the contact who greenlit these smaller "legit" crawlers moves on and decides that it no longer agrees with said website anymore? Is a price comp…

There are a lot of hypotheticals here. I think you'll convince CloudFlare, and their customers, if you could name names and mention specific examples? If you are a price comparison site getting blocked by cloudflare, site owners may be losing sales, and that's good feedback.

Price comparison example: https://shucks.top/ sometimes gets blocked by cloudflare. Most recent was getting blocked from checking B&H.

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#128
post #12

Cloudflare is both a great thing and a terrible thing that has happened to the internet in recent years. Great in that they have a fantastic UI to add your site in, basically shielding the average user from attacks. Bad from a standpoint of that now only Google, Bing, and maybe other big search engines have the capabilities to actually crawl the internet now. I don't see us getting a massive innovation in search on t…

I don't see us getting a massive innovation in search on the internet now that Google has such a massive foothold, and companies like Cloudflare stop innovation from happening. How are we "stopping search innovation"?

By being gatekeepers on which website crawling is okay and which is not.

No such filters should exist. Is it really that awfully bad without anything but basic filters (ban an IP for flooding)? Are there like, operations that try and spam every single Cloudflare-hosted website 24/7?

Legitimately curious if your anti-bot measures come from actual bad experience with the internet or is it just a liability limitation move? (Namely to reduce potential suing surface by angry data owners and/or three-letter agencies.)

---

Basically, if I am experimenting with a basic crawling program and I hit websites A and B 20 times each in a space of one hour, is that really deserving of a captcha or extra auth methods?

Not flaming but I am really curious. Do you have any data and rationale posted somewhere that go into deeper detail about why Cloudflare's bot detection is how it is?

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#129

Earlier quoted context omitted.

I do not understand what the parent means by a "distributed search engine" and I do not know what problem they are facing.

A search engine which is not run centrally by one organization on infrastructure in a known network, but rather something like YaCy where individual users run crawler nodes on networks that vary over time. Which makes "contact us for an exception" a no-go, as the relevant source IPs will constantly be changing.

Yeah, I certainly don't want those crawlers anywhere near my servers. Block by default and allow site admins to unblock should they want to seems like the best way. It is also already how it works with Cloudflare.

Re: Cloudflare’s CAPTCHA replacement with FIDO2/WebAuthn is a bad idea

#130
post #38

Earlier quoted context omitted.

Say I’m interested in building a small scale domain-specific search engine and only just started development. There’s no prototype yet and may never be. In this situation, how do you determine it’s a legit crawler? And what about crawlers with even more limited scopes (targeting only a handful of sites) that they can’t possibly be called search engines? Are they ever considered legit?

Be a good netizen? Respect robots.txt. Don't lie in your User-Agent. Don't crawl at a ridiculous rate. All those are a good starting point.

Even then, I've run into issues with scraping several sites for the reverse image search engine I operate. Luckily, in most cases I have been able to get in touch with the people running those sites to get a rule added for my IPs to allow them through. That's not scalable though, and limits where I can scrape/crawl from. Even something as simple as checking a site for updates every hour or two tends to get blocked after a few times. TBH, one of the only things I have found which helps is lying in the user agent and copying CF cookies. Luckily, I haven't had to play with that for a few months due to whitelisting, so not sure it it would still help. Things change rapidly.
Post reply on HN