Live data from Hacker News

The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

positiveblue.substack.com

411–420 of 520 posts

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#411
post #20

Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…

We have thousands of engineers of these companies right here on hackernews and they cry and scream about privacy and data governance on every topic but their own work. If you guys need a mirror to do some self reflection I am offering to buy.

I'll contribute for the mirror. The hypocrisy is so loud, aliens in outer space can hear it (and sound doesn't even travel in vacuum).

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#412
I do like Cloudflare in general, but the whole anti-AI push is just another form of the Luddism surrounding AI since 2022. Cloudflare perhaps wisely picked up on this trend and decided to capitalize on it, but I think it would be a mistake to allow it to become their brand.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#413

Earlier quoted context omitted.

The thing is that rhinoceroses aren't well-behaved. Even if some small fraction of them in theory might be well-behaved, the effort of trying to account for that is too small to bother. If 99% of rhinoceroses aren't well-behaved, the simple and correct response is to ban them all, and then maybe the nice ones can ask for a special permit. You switch from allow-by-default to block-by-default. Similarly it doesn't make…

ISPs are supposed to disconnect abusive customers. The correct thing to do is probably contact the ISP. Don't complain about scraping, complain about the DDOS (which is the actual problem and I'm increasingly beginning to believe the intent.)

Great! How do I get, say, Google's ISP to disconnect them?

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#414

Earlier quoted context omitted.

The biggest issue right now seems to be people renting their residential IP addresses to scraper companies, who then distribute large scrapes across these mostly distinct IPs. These addresses are from all over the world, not just your own country, so we'll either need a World Government, or at least massive intergovernmental cooperation, for regulation to help.

I don't think we need a world government to make progress on that point. The companies buying these services, are buying them from other companies. Countries or larger blocks like the EU can exert significant pressure on such companies by declaring the use of such services as illegal when interacting with websites hosted in the country or block or by companies in them.

It just seems too easy to skirt around via middlemen. The EU (say) could prosecute an EU company directly doing this residential scraping, and it could probably keep tabs on a handful of bank accounts of known bad actors in other countries, and then investigate and prosecute EU companies transferring money to them. But how do you stop an EU company paying a Moldovan company (that has existed for 10 days) for "internet services", that pays a Brazilian company, that pays a Russian company to do the actual residential scraping? And then there's all the crypto channels and other quid pro quo payment possibilities.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#415

Earlier quoted context omitted.

No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…

> "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" What this scenario actually reveals is that the words "open to the public" are not intended to mean "access is completely unrestricted". It's fine to not want to give completely unrestricted access to something. What's not fine, or at least what complicates things unnecessarily, is using words like "open and free" to des…

Nobody has ever meant "access is completely unrestricted".

As a trivial example: what website is going to welcome DDoS attacks or hacking attempts with open arms? Is a website no longer "open to the public" if it has DDoS protection or a WAF? What if the DDoS makes the website unavailable to the vast majority of users: does blocking the DDoS make it more or less open?

Similarly, if a concert is "open to the public", does that mean they'll be totally fine with you bringing a megaphone and yelling through the performance? Will they be okay with you setting the stage on fire? Will they just stand there and say "aw shucks" if you start blocking other people from entering?

You can try to rules-lawyer your way around commonly-understood definitions, but deliberately and obtusely misinterpreting such phrasing isn't going to lead to any kind of productive discussion.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#416
post #20

Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…

We have thousands of engineers of these companies right here on hackernews and they cry and scream about privacy and data governance on every topic but their own work. If you guys need a mirror to do some self reflection I am offering to buy.

In the recent days, the biggest delu-lulz was delivered by that guy who'd bravely decided to boycott Grok out of... environmental concerns, apparently. It's curious how everybody is so anxious these days, about AI among other things in our little corner of the web. I swear, every other day it's some new big fight against something... bad. Surely it couldn't ALL be attributed to policy in the US!

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#417

Earlier quoted context omitted.

You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.

> You have a problem with badly behaved scrapers, not AI. And you have a problem understanding that "freedom and openness" extend only to where the rights (e. g. the freedom) of another legal entity begins. When I don't want "AI" (not just the badly-behaved subset) rifling my website then I should be well within my rights to disallow just that, in the same way as it's your right to allow them access to your playgroun…

This is not what the parent means. What they mean is such behavior is a hypocrisy. Because you are getting access to truly free websites whose owners are interested in having smart chatbots trained on the free web, but you are blocking said chatbots while touting "free Internet" message.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#418
post #415

Earlier quoted context omitted.

> "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" What this scenario actually reveals is that the words "open to the public" are not intended to mean "access is completely unrestricted". It's fine to not want to give completely unrestricted access to something. What's not fine, or at least what complicates things unnecessarily, is using words like "open and free" to des…

Nobody has ever meant "access is completely unrestricted". As a trivial example: what website is going to welcome DDoS attacks or hacking attempts with open arms? Is a website no longer "open to the public" if it has DDoS protection or a WAF? What if the DDoS makes the website unavailable to the vast majority of users: does blocking the DDoS make it more or less open? Similarly, if a concert is "open to the public",…

>You can try to rules-lawyer your way around commonly-understood definitions

Despite your assertions to the contrary, "actually free to use for any purpose" is a commonly understood interpretation of "free to use for any purpose" -- see permissive software licenses, where licensors famously don't get to say "But I didn't mean big companies get to use it for free too!"

The onus is on the person using a term like "free" or "open" to clarify the restrictions they actually intend, if any. Putting the onus anywhere else immediately opens the way for misunderstandings, accidental or otherwise.

To make your concert analogy actually fit: A scraper is like a company that sends 1000 robots with tape recorders to your "open to the public" concert. They do only the things an ordinary member of the public do; they can't do anything else. The most "damage" they can do is to keep humans who would enjoy the concert from being able to attend if there aren't enough seats; whatever additional costs they cause (air conditioning, let's say) are the same as the costs that would have been incurred by that many humans.

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#419

Earlier quoted context omitted.

There's actually not much evidence of this, since the attack traffic is anonymous.

HN people working in these AI companies have commented to say they do this, and the timing correlates with the rise of AI companies/funding. I haven't tried to find it in my own logs, but others have said blocking an identifiable AI bot soon led to the same pattern of requests continuing through a botnet.

Did HN people present evidence?

Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch

#420
post #5

I have zero issue with Ai Agents, if there's a real user behind there somewhere. I DO have a major issue with my sites being crawled extremely aggressively by offenders including Meta, Perplexity and OpenAI - it's really annoying realising that we're tying up several cpu cores on AI crawling. Less than on real users and google et al.

> I DO have a major issue with my sites being crawled extremely aggressively by offenders including Meta, Perplexity and OpenAI Gee, if only we had, like, one central archive of the internet. We could even call it the internet archive. Then, all these AI companies could interface directly with that single entity on terms that are agreeable.

Internet Archive is missing enormous chunks of the internet though. And I don't mean weird parts of the internet, just regional stuff.

Not even news articles from top 10 news websites from my country are usually indexed there.

Post reply on HN