Earlier quoted context omitted.
I am not sure. How would crawlers filter this?
Check if the response time, the length of the "main text", or other indicators are in the lowest few percentile -> send to the heap for manual review. Does the inferred "topic" of the domain match the topic of the individual pages? If not -> manual review. And there are many more indicators. Hire a bunch of student jobbers, have them search github for tarpits, and let them write middleware to detect those. If you are…
Do people still do this, or do they just off shore the task?