Does this author have a big pre-established audience or something? Struggling to understand why this is front-page worthy.
End of an era for me: no more self-hosted git
11–20 of 228 posts
Re: End of an era for me: no more self-hosted git
#12The author of this post could solve their problem with Cloudflare or any of its numerous competitors. Cloudflare will even do it for free.
Re: End of an era for me: no more self-hosted git
#13Does this author have a big pre-established audience or something? Struggling to understand why this is front-page worthy.
Re: End of an era for me: no more self-hosted git
#14I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors.
Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated flows? I have not worked on bot detection the last few years, but it was very common for residential proxy based scrapers to hammer sites for years, so I'm wondering what's different.
Re: End of an era for me: no more self-hosted git
#15Re: End of an era for me: no more self-hosted git
#16The author of this post could solve their problem with Cloudflare or any of its numerous competitors. Cloudflare will even do it for free.
Re: End of an era for me: no more self-hosted git
#17Re: End of an era for me: no more self-hosted git
#18Does this author have a big pre-established audience or something? Struggling to understand why this is front-page worthy.
because he's unable to self-host git anymore because AI bots are hammering it to submit PRs. self-hosting was originally a "right" we had upon gaining access to the internet in the 90s, it was the main point of the hyper text transfer protocol.
It's painful to have your site offline because a scraper has channeled itself 17,000 layers deep through tag links (which are set to nofollow, and ignored in robots.txt, but the scraper doesn't care). And it's especially annoying when that happens on a daily basis.
Not everyone wants to put their site behind Cloudflare.
Re: End of an era for me: no more self-hosted git
#19Does anyone know what's the deal with these scrapers, or why they're attributed to AI? I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors. Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated…
It's a race to the bottom. What's different is we're much closer to the bottom now.
Re: End of an era for me: no more self-hosted git
#20Does anyone know what's the deal with these scrapers, or why they're attributed to AI? I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors. Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated…