Live data from Hacker News

Sites scramble to block ChatGPT web crawler after instructions emerge

arstechnica.com

11–20 of 38 posts

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#11
I have two sites that provide documentation for open source libraries I've created, and I definitely won't be blocking ChatGPT. It has already read my documentation and can correctly answer most StackOverflow-level questions about my libraries' use. This is seriously impressive and very helpful, as far as I'm concerned.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#12
post #8
post #5

blocked it on every single site I manage there is zero benefit to me in allowing OpenAI to absorb my content it is a parasite, plain and simple (as is GitHub Copilot) and I'll be hooking in the procedurally generated garbage pages for it soon!

What if it becomes the next Google? Are you really sure you want to be removed from their index?

You are promoting FOMO.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#13
post #8
post #5

blocked it on every single site I manage there is zero benefit to me in allowing OpenAI to absorb my content it is a parasite, plain and simple (as is GitHub Copilot) and I'll be hooking in the procedurally generated garbage pages for it soon!

What if it becomes the next Google? Are you really sure you want to be removed from their index?

Not the OP, but I already blocked it with my robots.txt. I am 100% sure I want to be removed from their index, even if they become "the next Google". I would rather have my website fade away into obscurity than increase the usefulness of their AI or any other AI model.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#14
From the article:

> For example, blocking content from future AI models could decrease a site's or a brand's cultural footprint if AI chatbots become a primary user interface in the future.

I would rather leave the internet entirely if AI chatbots become a primary user interface.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#15
post #5

blocked it on every single site I manage there is zero benefit to me in allowing OpenAI to absorb my content it is a parasite, plain and simple (as is GitHub Copilot) and I'll be hooking in the procedurally generated garbage pages for it soon!

How do you assess which robots benefit you?

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#16

I wonder if its worth poisoning the replies for scrapers that don't obey robots.txt. Send back nonsense, lies, and noise. This would be an adversarial approach like https://adnauseam.io/ uses for ad tracking.

If you (or others) come up with a way to build a system to poison AI/LLM/other models to make them useless, count me in to help.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#17

From the article: > For example, blocking content from future AI models could decrease a site's or a brand's cultural footprint if AI chatbots become a primary user interface in the future. I would rather leave the internet entirely if AI chatbots become a primary user interface.

What's interesting about your statement is that if you knew they were chatbots, then by definition, they wouldn't be AI. e.g. they wouldn't have passed the Turing test.

I know what you're saying, and totally agree. Unfortunately the term "AI" is now meaningless.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#19
post #8
post #5

blocked it on every single site I manage there is zero benefit to me in allowing OpenAI to absorb my content it is a parasite, plain and simple (as is GitHub Copilot) and I'll be hooking in the procedurally generated garbage pages for it soon!

What if it becomes the next Google? Are you really sure you want to be removed from their index?

If something happens, then you can do something about it.

In this particular case, if enough people block ChatGPT scraping then it cannot become the next google. Most notably, I imagine all commercial news organizations will block it because they need people to visit their actual website to pay for putting news up on their website. And it will remain that way until it can be demonstrated that ChatGPT drives more traffic to a website than it redirects traffic away from a website. The Microsoft chat in Edge is much closer to that in the way its summaries include clickable quotes from articles.

Re: Sites scramble to block ChatGPT web crawler after instructions emerge

#20

I wonder if its worth poisoning the replies for scrapers that don't obey robots.txt. Send back nonsense, lies, and noise. This would be an adversarial approach like https://adnauseam.io/ uses for ad tracking.

If you (or others) come up with a way to build a system to poison AI/LLM/other models to make them useless, count me in to help.

I’d imagine this is best possible via illegal methods such as mass hacking websites and inserting the appropriate poison
Post reply on HN