Live data from Hacker News

Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

website-auditor.io

51–60 of 66 posts

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#51
post #5

Are we saying that it's now a problem that we're not getting scraped? This appears to be a new generation of "SEO", marking itself as a service for getting into AI results? This is not the future I want to be a part of. Perhaps it's inevitable that after a break from everything being driven by money that LLMs will now be ruined by people spending $X to get into AI to make back $X+1, leading to an arms race of ever in…

> Are we saying that it's now a problem that we're not getting scraped?

No, it is saying that even the ones that are being scraped are not being included in results

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#52
post #20

This is a weird complaint. Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia. There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site? Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most si…

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

>But if I ask an LLM 'who are the best commercial roofers in Houston?', there isn't one right answer. When an LLM answers that roofing question, it typically cites 3 to 5 businesses. The other 40 legitimate roofing companies in the area are left out.

I don't see how mentioning the other 40 should follow from that fact. You get exactly what you're asking for - the best, according to agent's judgement. Judgement and processing of simple search results is exactly what you're using the intelligent agent for. Ask for a list of all commercial roofers in Houston and then you can expect a list.

It can be argued that models are clustering their replies around a few options when asked to choose from a list of equally valid ones (due to mode collapse, distribution biases, primacy/recency biases etc), but these options are not equally valid for the agent, it looks at their sites or possibly in some other places like review sites to rank them. Its ranking criteria might be not good at all, but that's another question.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#53

This is why we need the HTTP 402 standard to become common. If websites charge pennies per AI crawl, they will make more money than ever being reference in that 6.2% of websites that get cited (of which even another small percent get any follow through that leads to a sale or ad click) HTTP 402 also basically extends the pay per token model people have gotten used to with AI model providers, except applied to the who…

Seems like the opposite if the problem you’re trying to solve is “my site isn’t cited over someone else’s”. HTTP 467 - pls cite me, I’ll pay you.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#54

Earlier quoted context omitted.

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

Not really, ahrefs has already started working on this and there are several tools out there that track GEO. Also what is the incentive for me to rank on LLMs? People do not use it to click sources or even take action. Do you have any proof that being ranked as the best roofer in houston drives sales more than not being ranked as it? If not, why should i care as the roofer?

It depends. If you are doing blog content with the idea of upselling users to your product, probably not as much (because the LLM can just give the user the answer).

However, if you are looking for the best product/service/whatever, then yes it really does matter. I've bought _so_ many products because of LLM recommendations. For example, I wanted a new webcam, I asked the LLM to find me the best ones with a large sensor and Linux compatibility. It gave me a shortlist then I chose one and then I bought it.

This experience is far better than trailing through dozens of pages of (even pre LLM) SEO slop.I just tried the same on Google search and all the links recommended a camera with ~10% the sensor size that I bought.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#55
post #49

Earlier quoted context omitted.

If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece. The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself. If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofe…

Good, let’s hope there’s no SEO arms race for this too. Maybe the best roofers have a chance of getting cited over the ones with the best SEO agency. Seems unlikely though.

How are we defining the best roofers?

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#57

This is why we need the HTTP 402 standard to become common. If websites charge pennies per AI crawl, they will make more money than ever being reference in that 6.2% of websites that get cited (of which even another small percent get any follow through that leads to a sale or ad click) HTTP 402 also basically extends the pay per token model people have gotten used to with AI model providers, except applied to the who…

The problems with micropayments are

1. The market for lemons. In fact sites trying to "monetize their content" are the most likely to be lemons, so just asking is a signal that your "content" is not worth anything.

2. If your information does have value, it competes in a market with other sites full of high quality information that aren't trying to monetize it, meaning your specific information has to be specifically very valuable and not available yet for free elsewhere. If it is valuable as information (i.e. not something like creative writing), it will quickly spread and become freely available.

Lots of "content" just isn't worth anything (or has negative worth: it wastes your time). e.g. consider youtubers begging to get viewers to like/subscribe to increase their reach, and people generally don't despite it costing them nothing. Because it's not even worth a click to them.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#58
post #49

Earlier quoted context omitted.

Good, let’s hope there’s no SEO arms race for this too. Maybe the best roofers have a chance of getting cited over the ones with the best SEO agency. Seems unlikely though.

How are we defining the best roofers?

The model does it in this case, not "we". Classic SEO does influence it because it relies on traditional search, but then it estimates the best according to whatever cognitive capabilities it has, to its understanding of user's request, to what its harness tells it to do, and to its character engineered during post-training (and models' biases are very carefully engineered).

I would expect a judgement of a modern model to be superficial for this request, but still better than nothing. A deep research agent might be a lot better and much less susceptible to "optimization".

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#59

It is architecturally impossible for an LLM to associate a link or citation that it crawled with a response that comes out the other end. Every link they're giving you to support their statements is tacked on because it may vaguely match the tokens it just generated. It is perfectly common to find that a "source" does not contain the statements. You cannot expect an LLM to write you a Wikipedia article, much less a l…

Lol, this same behavior comes up more often than you'd expect in academia too. Many many citations that nobody cross checks that also don't support the originally cited piece of information.

with the advantage that you can discover this by reading citations.

Re: Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers

#60

Earlier quoted context omitted.

Lol, this same behavior comes up more often than you'd expect in academia too. Many many citations that nobody cross checks that also don't support the originally cited piece of information.

with the advantage that you can discover this by reading citations.

Yes, but it takes a lot of time to find a lie and very little time to conceive a lie. It's a losing battle as the percentage of honest citations drops (and it doesn't need to be a big drop to have a big impact.. think if the reliability of planes dropped 1%).
Post reply on HN