Live data from Hacker News

Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

website-auditor.io

31–40 of 66 posts

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#31
post #20

This is a weird complaint. Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia. There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site? Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most si…

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

Makes sense! Good point.

You're effectively reminding us that brands don't get understand well how to influence ai seo

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#32

It is architecturally impossible for an LLM to associate a link or citation that it crawled with a response that comes out the other end. Every link they're giving you to support their statements is tacked on because it may vaguely match the tokens it just generated. It is perfectly common to find that a "source" does not contain the statements. You cannot expect an LLM to write you a Wikipedia article, much less a l…

This is true, but good LLMs (even local, consumer-sized ones these days) can do this much better than … whatever it is that Google’s AI overviews use. Think about your coding harness: the model reads a bunch of files into the context, and it generally doesn’t forget/hallucinate which lines came from which file.

I'd be surprised if it were anything other than just an LLM only equivalent of Gemma some 4B param model, they can't be spending much money on it because each query essentially needs to be cheaper than the ad revenue per search.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#34
post #20

This is a weird complaint. Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia. There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site? Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most si…

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

That’s why I, from now on, only allow citable information to be scraped. The return of cloaking, lol.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#35
This makes sense. A handful of websites hold most of the “trusted” info because they’re massive. I don’t expect you to quote my blog with only three entries. The real trick would be getting AI companies to stop hammering sites that don’t show up in answers, but even if they don’t use a source for an answer, crawling still provides value.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#36
post #20

This is a weird complaint. Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia. There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site? Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most si…

The question is very important. What if you ask/tell an AI "I need a 2 bedroom vacation rental in Park City for a trip this fall". If you only get Airbnb results you are missing a lot of the true answer.

That's a "problem" of the harness that you use those, not a question of the model. The model can "know" things like "Linus Torvalds created Linux" but things like "Is X available in Y right now?" it obviously cannot, leading to this being a completely different thing compared to "just knowing" something and providing references to "how it knows".

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#37
post #13

It is architecturally impossible for an LLM to associate a link or citation that it crawled with a response that comes out the other end. Every link they're giving you to support their statements is tacked on because it may vaguely match the tokens it just generated. It is perfectly common to find that a "source" does not contain the statements. You cannot expect an LLM to write you a Wikipedia article, much less a l…

Then the least they can do is cite ALL their sources. At least somewhere on their website.

So an HTML page with a list of essentially every page on every site on the internet? What do you think, maybe a trillion URLs?

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#38
post #20

This is a weird complaint. Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia. There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site? Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most si…

Its a marketing post for a tool that offers visibility to site owners. Honestly these type of "neutral data reports" masquerades should be flagged or adequately disclosed.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#39

Earlier quoted context omitted.

This is true, but good LLMs (even local, consumer-sized ones these days) can do this much better than … whatever it is that Google’s AI overviews use. Think about your coding harness: the model reads a bunch of files into the context, and it generally doesn’t forget/hallucinate which lines came from which file.

I'd be surprised if it were anything other than just an LLM only equivalent of Gemma some 4B param model, they can't be spending much money on it because each query essentially needs to be cheaper than the ad revenue per search.

Could they even afford 4B dense parameters touching every token? That seems like a LOT of compute compared to generating the classic Google SERP.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#40
post #31

Earlier quoted context omitted.

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

Makes sense! Good point. You're effectively reminding us that brands don't get understand well how to influence ai seo

Small businesses don't. You're right. But somehow those larger companies seem to get cited. Wouldn't it be nice if we all knew how we were being influenced, both potential consumer and small business owner? Let's have the lid off!
Post reply on HN