Earlier quoted context omitted.
Or detect the LLM and serve up an LLM rewritten version of the page. That way you feed it poisonous garbage.
The issue is detecting them when they use random user agents and ip ranges.
From what I've seen, most AI scrapers operate on known cloud IP ranges, usually amazon (Perplexity included), so just check for those.