Live data from Hacker News

Cloudflare Introduces Default Blocking of A.I. Data Scrapers

nytimes.com

281–290 of 342 posts

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#281
post #201

Earlier quoted context omitted.

Not sure if you're joking, but if you're not: Congratulations on using a very "normal/safe" OS/browser/IP. I get captchas daily, without using any VPN and on several different IPs (work, home, mobile). The only crime I can think of is that I'm using Firefox instead of Chrome.

I use a VPN and firefox and I get some extra captchas but not enough to be annoying. And you don't have to do anything more than tap the checkbox. Meanwhile a bunch of "security" products other websites use just flat out block you if you're on a VPN. Other sites like youtube or reddit are in between where they block you unless you are logged in. Cloudflare is the least obtrusive of the options.

No, the least obtrusive option is the one you don't even notice because it actually works (or offers a non-painful secondary flow when it doesn't).

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#282
post #112

Earlier quoted context omitted.

It's cloudflare and parasites like them that will make the internet un-free. It's already happening, I'm either blocked or back to 1998 load times be cause of "checking your browser". They are destroying the internet and will make it so only people who do approved things on approved browsers (meaning let advertising companies monetize their online activity) will get real access. Cloudflare isn't solving a problem, th…

I use Firefox with adblocking and some fingerprinting anti-measurements and I rarely hit their challenges. Your IP reputation must be bad. They have an addon [1] that helps you bypass Cloudflare challenges anonymously somehow, but it feels wrong to install a plugin to your browser from the ones who make your web experience worse 1: https://developers.cloudflare.com/waf/tools/privacy-pass/

> Your IP reputation must be bad.

And for an extremely large number of honest users, they cannot realistically avoid this.

I live in India. Mobile data and fibre are all through tainted CGNAT, and I encounter Cloudflare challenges all the time. The two fibre providers I know about use CGNAT, and I expect others do too. I did (with difficulty!) ask my ISP about getting a static IP address (having in mind maybe ditching my small VPS in favour of hosting from home), but they said ₹500/month, which is way above market rate for leasing IPv4 addresses, more than I pay for my entire VPS in fact, so it definitely doesn’t make things cheaper. And I’m sceptical that it’d have good reputation with Cloudflare even then. It’ll probably still be in a blacklisted range.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#283

Earlier quoted context omitted.

No…? Someone can clearly implement a whitelist system that applies only to ipv6… but that makes no judgement on ipv4.

Let's back up a step. You said by definition a whitelist system would consider every IPv6 suspicious (until it's put on the list, presumably). What is that definition? If "applies only to IPv6" is an optional decision someone could make, then it's not part of the definition of a whitelist system for IPs, right?

What are you talking about?

The prior comment was responding directly to your comment, not any comment preceding that.

Of course it’s no longer by definition if you expand the scope beyond an ipv6 whitelist as there are an infinite number of possible whitelists.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#284
post #276
post #254

Earlier quoted context omitted.

There's a non proof of work challenge: https://anubis.techaro.lol/docs/admin/configuration/challeng... Also: Anubis does not mine cryptocurrency. Proof of work is easy to validate on the server and economically scales poorly in the wild for abusive scrapers.

Thanks for the link. I’ll have a look. I’m glad there’s no cryptocurrency involved (was never a concern) but I worry about the optics of something so closely associated. (I appreciate your commenting on this. I know the project recently blew up in popularity. Keep up the great work)

If you have suggestions for JS based challenges that don't become a case of "read the source code to figure out how to make playwright lie", I'm all ears for the ideas :)

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#285

Earlier quoted context omitted.

I use Firefox with adblocking and some fingerprinting anti-measurements and I rarely hit their challenges. Your IP reputation must be bad. They have an addon [1] that helps you bypass Cloudflare challenges anonymously somehow, but it feels wrong to install a plugin to your browser from the ones who make your web experience worse 1: https://developers.cloudflare.com/waf/tools/privacy-pass/

> Your IP reputation must be bad. And for an extremely large number of honest users, they cannot realistically avoid this. I live in India. Mobile data and fibre are all through tainted CGNAT, and I encounter Cloudflare challenges all the time. The two fibre providers I know about use CGNAT, and I expect others do too. I did (with difficulty!) ask my ISP about getting a static IP address (having in mind maybe ditchin…

Why don't your ISPs just use IPv6?

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#286
post #120

Earlier quoted context omitted.

How is Cloudflare a parasite? I can use Cloudflare, and get their AI protection, for free. I have dozens of domains I have used with Cloudflare at one point and I haven't paid them a dime.

Did you read his comment? He explained the issue he has with Cloudflare...

Yeah but they are a dictator, OpenAI et al are the parasites.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#287

Earlier quoted context omitted.

Including your comment, including this comment. HN itself is routinely scraped. What makes me most uncomfortable is deanonymization via speech analysis. It's something we can already do but is hard to do at scale. This is the ultimate tool for authoritarians. There's no hidden identities because your speech is your identifier. It is without borders. It doesn't matter if your government is good, a bad acting governmen…

The degree to which people say “self-delete” and “unalive” is absurd these days and I now hear it in real life. It’s Orwellian in the truest sense of the word.

Orwell was the optimist. It’s Huxley’s vision we should be really worried about. Brave new world indeed.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#288
post #13

> If A.I. companies freely use data from various websites without permission or payment, people will be discouraged from creating new digital content I don't see a way out of this happening. AI fundamentally discourages other forms of digital interaction as it grows. Its mechanism of growing is killing other kinds of digital content. It will eventually kill the web, which is, ironically, its main source of food.

[flagged]

Now? Always has been. Compared to alternatives it’s still the best economy-scale resource usage optimization framework we’ve got.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#289
post #232
post #161

Earlier quoted context omitted.

Open source software typically has a license. People not following the license isn’t tolerated. This is what AI scrapers are doing. They’re taking your code, your artwork and your writing without any consideration for the license.

Weather training on code is fair use is still an open legal question, and it may well be fair use. The way a license works is by saying "you have my permission to use this code as long as you follow these conditions", but if no license is required than the conditions are irrelevant. There is an active case on this, where Microsoft has been sued over GitHub copilot, and it has been slowly moving through the court syst…

> The way a license works is

Let's actually look at the MIT license, a very permissive license

  > Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to ***use***, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

  > The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
So, you can use it but need to cite the usage. It's not that hard. Fair use if you just acknowledge usage.

Is it really that difficult to acknowledge that you didn't do everything on your own? People aren't asking for money. It's just basic acknowledgement.

Forget the courts for a second, just ask yourself what is the right thing to do. Ethically.

Re: Cloudflare Introduces Default Blocking of A.I. Data Scrapers

#290
post #112

Earlier quoted context omitted.

It's cloudflare and parasites like them that will make the internet un-free. It's already happening, I'm either blocked or back to 1998 load times be cause of "checking your browser". They are destroying the internet and will make it so only people who do approved things on approved browsers (meaning let advertising companies monetize their online activity) will get real access. Cloudflare isn't solving a problem, th…

I use Firefox with adblocking and some fingerprinting anti-measurements and I rarely hit their challenges. Your IP reputation must be bad. They have an addon [1] that helps you bypass Cloudflare challenges anonymously somehow, but it feels wrong to install a plugin to your browser from the ones who make your web experience worse 1: https://developers.cloudflare.com/waf/tools/privacy-pass/

I'm having lots of problems with fingerprinting protection on Librewolf and ungoogled-chromium. I use uBlock Origin and JShelter extensions on both. I'm always getting "your browser is out of date" despite always having the most newest versions.

Some sites like Stackexchange will work after just reloading the page. And rest of the sites usually work when I remove Javascript protection and Fingerprint detection from JShelter. Sill not all of them. So, they maybe/probably want to reliably fingerprint my browser to let me continue.

If I use crappy fingerprint protection, I'm not having problems but if I actually randomize some values then sites wont work. JShelter deterministicly randomizes some values using session identifier and eTLD+1 domain as a key to avoid breaking site functionality but apparently Cloudflare is beeing really picky. Tor browser is not having these problems but it uses different strategy to protect itself from fingerprinting and doesn't randomize values but tries to have unified values across different users making identification impossible.

Post reply on HN