Ask HN: The amount of AI bot traffic is out control?
1–9 of 9 posts
Re: Ask HN: The amount of AI bot traffic is out control?
#2Re: Ask HN: The amount of AI bot traffic is out control?
#3How do you know it's facebook's crawler if it going through residential proxies?
Re: Ask HN: The amount of AI bot traffic is out control?
#4Re: Ask HN: The amount of AI bot traffic is out control?
#5Anyhow, I occasionally suffer this same problem when happy crawlers find my site and use truesign.ai to block them, so far successfully. It detects residential proxies, fake emails and overall suspicious activity, and - something I really wanted to avoid - there are no captchas.
Re: Ask HN: The amount of AI bot traffic is out control?
#6How do you know it's facebook's crawler if it going through residential proxies?
It had meta in the user-agent headers, it’s trivial to spoof but I’ve read of many others who are seeing the same thing. It’s not just them though, it’s also parallel web systems, and then a bunch of others who aren’t identifying themselves.
Re: Ask HN: The amount of AI bot traffic is out control?
#7How do you differentiate between an Ai scraper and an end user having an agent perform a search and then fetch all the results to analyze them for relevance? (because search has become so bad I need an agent to deal with it before I look at things)
Re: Ask HN: The amount of AI bot traffic is out control?
#8> The AI scrapers are routing requests through residential proxies How do you differentiate between an Ai scraper and an end user having an agent perform a search and then fetch all the results to analyze them for relevance? (because search has become so bad I need an agent to deal with it before I look at things)
Re: Ask HN: The amount of AI bot traffic is out control?
#9> The AI scrapers are routing requests through residential proxies How do you differentiate between an Ai scraper and an end user having an agent perform a search and then fetch all the results to analyze them for relevance? (because search has become so bad I need an agent to deal with it before I look at things)
Because the sheer volume is insane, and when looking through the logs the requests for a short term IP do not make sense / are not sequential, hope that makes sense
It sounds like you are reading tea leaves (based on your other comments) and volume alone is not an indicator to the source. My personal web page access has gone up 10x because I have agents doing things, and they are sloppy af, fetching way more than they need to. Insane volume from individuals, by way of their personal agents, is not surprising to me, given what I have at my hands and what I know others have and do in other harnesses.
Them not being sequential, or typical scraping patterns, would to me lend credence towards people using deep research agents because google search has become so bad. In other words, the SEO industry went too hard, helped ruin search result quality, and now we the people are using Ai to sift through the noise for what's actually useful, work around all that money that manipulates what we see. For me, what used to be a single google search with useful page snippets is now an agent running multiple queries across multiple SERPs and fetching dozens or >100 pages.