Earlier quoted context omitted.
You don't think that the AI companies will take efforts to detect and filter bad data for training? Do you suppose they are already doing this, knowing that data quality has an impact on model capabilities?
The current state of the art in AI poisoning is Nightshade from the University of Chicago. It's meant to eventually be an addon to their WebGlaze[1] which is an invite-only tool meant for artists to protect their art from AI mimicry If these companies are adding extra code to bypass artists trying to protect their intellectual property from mimicry then that is an obvious and egregious copyright violation More likely…
The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
481–490 of 520 posts
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#482Earlier quoted context omitted.
Perplexity has been one of the AI companies that created the problem that gave rise to this CF proposal. Why doesn't Perplexity invest more into being a responsible scraper? https://blog.cloudflare.com/perplexity-is-using-stealth-unde...
Re-read what I wrote.
What did you say that relates to Perplexity being one of the reasons that Cloudflare and their customers have decided they need better protection from abusive scrapers?
Websites choose their own gatekeepers, Cloudflare is just one provider
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#483Everyone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed d…
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#484Earlier quoted context omitted.
No, that is not true. It is only true if you just equate "AI training bots" with "people" on some kind of nominal basis without considering how they operate in practice. It is like saying "If your grocery store is open to the public, why is it not open to this herd of rhinoceroses?" Well, the reason is because rhinoceroses are simply not going to stroll up and down the aisles and head to the checkout line quietly wit…
You have a problem with badly behaved scrapers, not AI. I can't disagree with being against badly behaved scrapers. But this is neither a new problem or an interesting one from the idea of making information freely available to everyone, even rhinoceroses, assuming they are well behaved. Blocking bad actors is not the same thing as blocking AI.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#485Earlier quoted context omitted.
"If you discuss this thread with your friends are you going to give me credit?" Yes. How else would I enable my friends to look it up for themselves?
6 months from now when you've internalized this entire thread are you even going to remember where you got it from?
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#486Earlier quoted context omitted.
> Asking people to just unilaterally disarm by imposing no restrictions I'm not asking for this. I'm asking for people who want such restrictions (most of which I consider entirely reasonable) to say so explicitly. It would be enough to replace words like "free" or "open" with "fair use", which immediately signals that some restrictions are intended, without getting bogged down in details.
Why? It seems you already know what people mean by "open and free", and it does have a connection to the ideals of openness and freedom, namely in the systemic context that I described above. So why bother about the terminology?
The only sensible way forward is to be explicit.
Why fight this obvious truth? Why does it hurt so much to say what you mean?
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#487Earlier quoted context omitted.
This has been my experience more recently as well, I've finally migrated from google to Brave Search since google was just slow for me. I also appreciate the AI search results a bit when im looking for something very specific (like what the yaml definition for a docker swarm deployment constraint looks like) because the AI just gives me the snippet while the search results are 300 medium blog posts about how to use d…
Not to mention how much worse it is on mobile. Every web site asks me to accept their cookies, close layers of ads with tiny buttons, and loads slowly with ads spread throughout the content. And that’s just to figure out if I’m even on the right page.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#488> An allowlist run by ONE company? An allowlist run by one company that site owners chose to engage with. But the irony of taking an ideological stance about fairness while using AI generated comics for blog posts…
The bots/crawlers/browsers are pre-categorized by CloudFlare.
Defaults matter and how CloudFlare categorizes your privacy-focuses or agentic browser would impact your experience on a good chunk of the web.
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#489Earlier quoted context omitted.
They're getting to the point of 200-300RPS for some of my smaller marketing sites, hallucinating URLs like crazy. It's fucking insane.
You'd think they would have an interest in developing reasonable crawling infrastructure, like Google, Bing or Yandex. Instead they go all in on hosts with no metering. All of the search majors reduce their crawl rate as request times increase. On one hand these companies announce themselves as sophisticated, futuristic and highly-valued, on the other hand we see rampant incompetence, to the point that webmasters eve…
Re: The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
#490Earlier quoted context omitted.
The website owner has rights too. Are you arguing they cannot choose to implement such gatekeeping to keep their site operating in a financially viable manner?
The first article of our constitution says people shall be treated equally in equal situations. I presume that most countries have similar clauses but, beyond legalese, it's also simply in line with my ethics to treat everyone equally There are people behind those connection requests. I don't try to guess on my server who is a bot and who is not; I'll make mistakes and probably bias against people who use uncommon se…
Do not apply laws where they do not apply.