Earlier quoted context omitted.
Static, Near Static (not generated on demand at least; generated only on real content update), and Login seems likely. AI not caching things is a real issue. Sites being difficult TO cache / failing the 'wget mirror test' is the other side of the issue.
What about AI not respecting robots.txt? I myself have never ran into this, but I've seen complaints of many people who did.
since when actor that want gather your entire data respect things like this??? how can you enforce such things with just "please don't crawl this directory thanks"