Live data from Hacker News

Blocking LLMs from your website cuts you off from next-generation search

johnjianwang.medium.com

71–80 of 82 posts

Re: Blocking LLMs from your website cuts you off from next-generation search

#71

> how many of you wouldn’t hook up your website to Google? If there was a paid-only search engine with dubious ethics practices that was overwhelming my site with traffic in order resell search trained off of (among other things) my personally generated content, I would absolute block it. LLMs are not search engines, and I'm not gaining any followers or customers in any meaningful way because an LLM indexes my site.…

> If there was a paid-only search engine with dubious ethics practices that was overwhelming my site with traffic in order resell search trained off of (among other things) my personally generated content, I would absolute block it.

Let’s compare Google with OpenAI:

Paid-only: neither check; both have free tiers, eventually supported by ads (Google took 10+ years before it got littered with ads, I promise OpenAI will make the ad experience even stinkier because they keep you on the site as opposed to Google who only have you for a few seconds. The ads will be blinky, and they will be nested into the content.

Dubious ethics: both check.

Overwhelming bot traffic: both check.

Make money on your content: both check.

> LLMs are not search engines, and I'm not gaining any followers or customers in any meaningful way because an LLM indexes my site.

So paywall it?

The Anubis PoW captcha is an option, too. Then you will block trainers and allow agents.

Re: Blocking LLMs from your website cuts you off from next-generation search

#72
post #39

> how many of you wouldn’t hook up your website to Google? If there was a paid-only search engine with dubious ethics practices that was overwhelming my site with traffic in order resell search trained off of (among other things) my personally generated content, I would absolute block it. LLMs are not search engines, and I'm not gaining any followers or customers in any meaningful way because an LLM indexes my site.…

> LLMs are not search engines, and I'm not gaining any followers or customers in any meaningful way because an LLM indexes my site. Friends of mine run a service company, and they already see a significant number of customers reach out because they found them using ChatGPT (et al), not Google. By significant I mean ~20% or so. Also, for e-commerce, Deep Research from OpenAI, is way better in doing product recommendat…

> Deep Research from OpenAI, is way better in doing product recommendations than Google

Interesting, that’s not my experience and I’d be the first to replace Google if I could. I’ll have to try again.

Re: Blocking LLMs from your website cuts you off from next-generation search

#73

Good, because this “next generation search” doesn’t cite sources, invents falsehoods, steals content, and doesn’t direct traffic to the site in question, which was the whole point of search engines in the first place . The fact LLM companies constantly keep getting dinged for ignoring every barrier we throw up to stop their scraping short of something like Anubis shows what their real goal is: theft, monopolization,…

> this “next generation search” doesn’t cite sources

The training done on the content does not provide citable references with current models. The agentic search and summary done post-training does.

A lot of the heavy traffic is for training, though, because AI companies are in competition for large amounts of training data.

Re: Blocking LLMs from your website cuts you off from next-generation search

#74
post #37

I agree with the sentiment. I remember Gwern in an interview remarking something to the effect that if you make your writing and thoughts invisible to LLMs, then your thoughts are going to be invisible to the future, as LLMs are here to stay.

Might be thinking of https://www.dwarkesh.com/p/gwern-branwen#%C2%A7influencing-t...

Re: Blocking LLMs from your website cuts you off from next-generation search

#75

The first sentence of the article is literally wrong as it conflates LLM and the search part of a RAG (retrieval augmented generation, when you mix a web search and an LLM). Blocking bots cuts you off from the next-generation search, because it cuts you off from search at all. So far, blocking LLM simply prevents you from being part of the training dataset, which is not the same thing. Please stop upvoting such bad c…

> blocking LLM simply prevents you from being part of the training dataset That's narrow. Perplexity, and other LLM agent services, do perform a regular web search to gain context, before generation their output. How else would they have access to recent data when the underlying LLM's knowledge cutoff is usual at least a few weeks?

They are normal scrapers nothing specific to LLM as they are not yet used for training an LLM, unless I miss something from their architecture. So I don't get why they would be called LLM crawlers, when they are search engine crawlers. At least they could be called RAG crawlers for better nuance. The article linked in the post first sentence is more precise as it deals with scrapers: https://techcrunch.com/2025/08/04/perplexity-accused-of-scra... Some people may be ok with search engines but not LLM training so it's not the same deal.

Re: Blocking LLMs from your website cuts you off from next-generation search

#76
post #69

Earlier quoted context omitted.

Users don't seek content for the attribution; that's extra noise, unless there's reason to contact the attributed. And given that many websites offer an inefficient flow to content, made of ads and/or unnecessarily animated things for example, the LLM is merely improving the experience for the user .

The whole conversation here is about incentives for the content creators, not the user. Yes, as a user I'd like everything served to me on a silver platter, for free, on demand, and completely and 100% aligned with my interests exclusively with no thought given to anybody else... but that's not a realistic world. In the real world, if the content providers have no reason to provide content, they won't. I kind of hate…

There are many content creators out there who create and share with the only incentive being for fun and/or interest. However their content is in most cases down-ranked due to a lack of SEO, because no time to waste on that silliness when there are better things to do. The internet is merely reverting to its previous form where such strings-free content was the order of the day, instead of the highly SEO'd spam that's primarily aimed at getting ad impressions or whatever nonsense that degrades user experience.

Re: Blocking LLMs from your website cuts you off from next-generation search

#77

Earlier quoted context omitted.

> blocking LLM simply prevents you from being part of the training dataset That's narrow. Perplexity, and other LLM agent services, do perform a regular web search to gain context, before generation their output. How else would they have access to recent data when the underlying LLM's knowledge cutoff is usual at least a few weeks?

They are normal scrapers nothing specific to LLM as they are not yet used for training an LLM, unless I miss something from their architecture. So I don't get why they would be called LLM crawlers, when they are search engine crawlers. At least they could be called RAG crawlers for better nuance. The article linked in the post first sentence is more precise as it deals with scrapers: https://techcrunch.com/2025/08/04…

Crawling is done to discover and index content for search results (to relieve dependence on Google, etc). Scraping is done to get relevant content into the LLM's context window. And then the LLM generates the output. All the functions are there, so someone may emphasize just a subset to try making their point (which can cause issues if relevant context is left out, whether accidentally, ignorantly or maliciously).

> RAG crawlers

Very few people know what "RAG" is, so it makes little sense to mention it to any other than a technical audience.

> not LLM training

There's an issue of trust, because once content is scraped, it can also be used to train future models. That's really what ought to be emphasized IMO.

Re: Blocking LLMs from your website cuts you off from next-generation search

#78

And here's me, wondering how anyone in their right mind would ever consider putting anything original or worth reading on the internet ever again. If you do you're just feeding the AI monster for free.

The internet is dead. I might need to stop writing here. I have no clue how many of the users here are bots by now and that dissolves the fun of it. I'll start a book circle or something ...

Re: Blocking LLMs from your website cuts you off from next-generation search

#79
post #70
post #42

Earlier quoted context omitted.

Fair, if your content is your product, but I’m more than happy for every LLM on the planet to summarize my page and hype the virtues of my product to its user.

Enjoy the brief window of LLMs "hyping the virtues of your product" to its users for free. In 2030 that's not going to sound realistic at all. And I feel I'm being generous pushing it up to 2030, the first "sponsored training data" either already exists or will probably be out this year, the only question being whether it will be publicly admitted to or not.

No doubt, but that doesn't change the position.

Re: Blocking LLMs from your website cuts you off from next-generation search

#80

And here's me, wondering how anyone in their right mind would ever consider putting anything original or worth reading on the internet ever again. If you do you're just feeding the AI monster for free.

The internet is dead. I might need to stop writing here. I have no clue how many of the users here are bots by now and that dissolves the fun of it. I'll start a book circle or something ...

I was thinking this last night. Let's start one together!
Post reply on HN