Live data from Hacker News

Cloudflare's new marketplace lets websites charge AI bots for scraping

techcrunch.com

251–260 of 280 posts

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#251
post #166
post #130

Earlier quoted context omitted.

> Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot Those ones are the fucking worst. I've noticed that if I try to succeed in these captchas too quickly, it'll just say "Sorry, try again" even when every click was correct, so instead, I've started going in slow motion and faking "misclicking" which makes it much more likely to accept me as human. I canno…

I always spoil as many of these as possible. Sometimes it takes me a while to prove that I'm human, but I'm dead-set on convincing it that I'm a stupid human. Of course, I fantasize that some day a robo-car will crash because I taught it that there's really no difference between a motorcycle and a flight of stairs.

> but I'm dead-set on convincing it that I'm a stupid human

this guy is really dumb BUT he has a very high quality computer THUS he is in the managerial class Final -> Ramp up the Ads!

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#252
post #166

Earlier quoted context omitted.

I always spoil as many of these as possible. Sometimes it takes me a while to prove that I'm human, but I'm dead-set on convincing it that I'm a stupid human. Of course, I fantasize that some day a robo-car will crash because I taught it that there's really no difference between a motorcycle and a flight of stairs.

https://qntm.org/frame Excellent short story that’s, somewhat related at least.

It seems sort of like over-engineering here - pretty sure this kind of thing would never happen with the Illuminati Ganga Automated Drive-By Solution https://medium.com/luminasticity/the-illuminati-ganga-automa...

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#253

Common Crawl is shown in their screen shot of "Providers" along side OpenAI and Antropic. The challenge is that Common Crawl is used for a lot of things that are not AI training. For example, it's a major source of content for the Wayback machine. In fact, that's the entire point of the Common Crawl project. Instead of dozens of companies writing and running their (poorly) designed crawlers and hitting everyone's sit…

already sites like perplexity have been completed blocked by cloudflare due to some meta signal and can't even load it. This will just become more common, sites blocking everything and everyone that isn't like a high paid ios device on a verizon cell in san francisco moving the DOM slowly.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#254

Earlier quoted context omitted.

Either AI training is fair use or it isn't. If it's fair use then businesses shouldn't get a say in whether the data can be used for it. If it isn't, then the answer to your question is copyright law. Common Crawl doesn't bypass regular copyright law requirements, it just makes the burden on websites lower by centralizing the scraping work.

Its not a legal question but a behavior and sustainability question. If it is fair use, but is undesirable for content makers, then they’re still not under any obligation to allow scraping. So they’ll try stuff like this, and other more restrictive bot blockers. Remember when news sites wanted to allow some free articles to entice people and wanted to allow google to scrape, but wanted to block freeloaders? They deci…

In the end the websites always lose that battle if humans are willing to put effort into sharing it. You see people just pasting full article text or summaries into reddit comments. Those people are probably subscribers.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#255

Is anybody else seeing an absolutely massive amount of Amazonbot crawls on their site? What are they up to? And why so aggressively?

They have documentation on verifying if it is indeed their bot: https://developer.amazon.com/amazonbot

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#256
post #141

Earlier quoted context omitted.

FWIW, it can't be cookies alone that gives you an inordinate number of bot challenges. I use private tabs on Firefox (for Linux and Android) for most of my browsing, and I rarely get any challenges regardless of what I do. The only issues tend to be when I make repeated searches for things with "quotes" and whatnot on Google or on Stack Exchange sites. But for the most part, those challenges aren't particularly drawn…

It varies a lot based on what I'm doing. Sites that rely on ads like english-language¹ recipes or health information have a lot of "you're European so you're blocked altogether" or "let me check that the connection is secure , ah wait, here is a captcha for you to solve" pages. Anything that needs to do fraud detection usually hates me as well, perhaps because I have a phone number and bank account from another count…

    > That German ISPs have daily-rotating IP addresses
This is interesting. What is the purpose? Security? Privacy?

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#257
post #102

Earlier quoted context omitted.

Those "keep clicking until we stop fading in more results" challenges mean they're fairly confident you're a bot and this is the highest difficulty level to prove your lack of guilt. I get these only when using a browser that isn't already full of advertising cookies (edit: which, to be clear, I hope is still considered an acceptable state to have your browser in)

It's acceptable, but suspicious. Two standard deviations away from the median browser (and a lot more like the configuration of a scraper, which would get reloaded in some Docker instance frequently with a fresh empty cookie jar because storing data costs infrastructure).

right, so using the heuristics libraries to determine if you were a bot you are probably already 65% bot, then if the threshold is 70% bot maybe you just need to tab really quick to an input and control-c your password and there you are.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#258
post #129
post #105

Earlier quoted context omitted.

I wonder how many of those captchas are controlled by competitors of Firefox?

ReCAPTCHA absolutely hammers Firefox compared to Chrome for me. On sites that use it for login I rarely just get the "check the box" challenge anymore, and am instead being asked to train their CV algorithms by picking 5+ images of stoplights or motorcycles. Punishment for avoiding the Chrome universe I guess.

part of Google's control of captcha also has to do with knowing who you are, so if you come to a site but google knows who you are and have a 99% surety you are not a bot even if you act very botlike on that site you probably aren't going to get any problems.

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#259

Earlier quoted context omitted.

Oh you will not notice. The pages can easily be spread out between residential IPs using headless browsers (masked as real ones), unless you really pay attention you won't see the ones that want to hide.

How many scrapers are sophisticated enough to go this far though? Most of them are probably of bad quality and can be detected.

Why would those sophisticated enough to go that far, be of low quality

Re: Cloudflare's new marketplace lets websites charge AI bots for scraping

#260
post #141

Earlier quoted context omitted.

It varies a lot based on what I'm doing. Sites that rely on ads like english-language¹ recipes or health information have a lot of "you're European so you're blocked altogether" or "let me check that the connection is secure , ah wait, here is a captcha for you to solve" pages. Anything that needs to do fraud detection usually hates me as well, perhaps because I have a phone number and bank account from another count…

> That German ISPs have daily-rotating IP addresses This is interesting. What is the purpose? Security? Privacy?

Preventing hosting from a home server without paying for a static IP.
Post reply on HN