Live data from Hacker News

End of an era for me: no more self-hosted git

kraxel.org

191–200 of 228 posts

Re: End of an era for me: no more self-hosted git

#191

I cut traffic to my Forgejo server from about 600K request per day to about 1000: https://honeypot.net/2025/12/22/i-read-yann-espositos-blog.h... 1. Anubis is a miracle. 2. Because most scrapers suck, I require all requests to include a shibboleth cookie, and if they don’t, I set it and use JavaScript to tell them to reload the page. Real browsers don’t bat an eye at this. Most scrapers can’t manage it. (This wasn’t…

IIRC 2 is included in the "go-away" program which is similar to Anubis but without the PoW

Re: End of an era for me: no more self-hosted git

#192

Does anyone know what's the deal with these scrapers, or why they're attributed to AI? I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors. Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated…

I just threw up a public Forjego instance for some lightweight collaboration. About 2 minutes after the certificate was created, I'm guessing they picked up the instance from the transparency logs for certificates, and started going through every commit and so on from the two repositories I had added. Watched it for a while, thinking eventually it'd end. It didn't, seemed like Claudebot and GPTBot (which was the only…

One idea is to not block them, but return a plausible looking page that's not the normal page, then it won't be detected as a block.

Re: End of an era for me: no more self-hosted git

#193
post #56

Does anyone know what's the deal with these scrapers, or why they're attributed to AI? I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors. Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated…

I stopped trying to understand. Encountering a 404 on my site leads directly to a 1 year ban.

[dead]

Re: End of an era for me: no more self-hosted git

#194
post #50

Earlier quoted context omitted.

The dirty secret is a lot of them come through "residential proxies", aka backdoored home routers, iot devices with shitty security, etc. Basically the scrapers who are often also third party, go to these "companies" and buy access to these "residential proxies". Some are more... considerate than others. Why? Data. Every bit of it is it might be valuable. And not to sound tin foil hatty, but we are getting closer to…

Has this actually been investigated and proven to be true? I see allegations, but no facts really. It seems to me to be just as likely that people are installing LLM chatbot apps that do the occasional bit of scraping work on the sly, covered by some agreed EULA.

[dead]

Re: End of an era for me: no more self-hosted git

#195
post #50

Earlier quoted context omitted.

The dirty secret is a lot of them come through "residential proxies", aka backdoored home routers, iot devices with shitty security, etc. Basically the scrapers who are often also third party, go to these "companies" and buy access to these "residential proxies". Some are more... considerate than others. Why? Data. Every bit of it is it might be valuable. And not to sound tin foil hatty, but we are getting closer to…

it isn't that hard to just buy a bunch of sim cards and put them in a modem and use that. it's good enough as a residential proxy. source: I did that before, when I worked on plaid-like thing.

[dead]

Re: End of an era for me: no more self-hosted git

#196

Earlier quoted context omitted.

[flagged]

This is really not the type of legacy that I want to leave behind on hackernews but then again, I have been vocal that I just write what I think. Literally. It has its flaws but I am not sugar coating it. Sometimes I am unable to explain myself but the thing is that I write on HN to point out of some idea, some discussion. It's better written here than lost and yes most of my ideas might be incoherent but they make p…

Well that's the best reply I've seen to someone being cranky on the Internet. I'm not even mad.

Take care.

Re: End of an era for me: no more self-hosted git

#197

Earlier quoted context omitted.

How can I detect if my router is backdoored, or being used as a residential proxy?

> How can I detect if my router is backdoored, or being used as a residential proxy? Aside from the obvious smoke tests (are settings changing without your knowledge? Does your router expose access logs you can check?), I'm not sure there's any general purpose way to check, but 2 things you can do are: 1. search for your router's model number to see if it's known to be vulnerable, and replace it with a brand-new repu…

I work for IPinfo. We track close to a hundred resproxy providers. So, if OP's router is compromised, the device IPs will likely be flagged.

From what I know, whenever a router is backdoored or a resproxy SDK gains access to a device to use their bandwidth, the access to that pool of devices is often shared among multiple resproxy vendors. Many resproxy vendors do not have their own SDKs for their services.

Also, as far as I know, not many resproxy operators manage their sim farms or hardware pools. It is mostly based on compromised devices or SDK access.

Re: End of an era for me: no more self-hosted git

#198

Earlier quoted context omitted.

[flagged]

This is really not the type of legacy that I want to leave behind on hackernews but then again, I have been vocal that I just write what I think. Literally. It has its flaws but I am not sugar coating it. Sometimes I am unable to explain myself but the thing is that I write on HN to point out of some idea, some discussion. It's better written here than lost and yes most of my ideas might be incoherent but they make p…

Great response.

> Teach me instead of such tone for I am interested in learning

Here are some notes:

Run on sentences and lack of punctuation make your writing hard to follow; brevity can be effective.

For each sentence, choose a subject, verb, predicate, proposition, etc. to form a single clause, but don't compound multiple such clauses into a single sentence. Break sentences up with punctuation so that the eye rests more easily when scanning. Eye fatigue is a real thing that good writers know how to manage. Contractions can also help clean up the noise.

It's okay to occasionally have compound sentences, such as this one, but too many of those leave your reader's head spinning.

It's fine and encouraged to write your initial draft in stream-of-consciousness form as you have, but an editing pass would make a worthwhile difference for slightly more effort. You do well at breaking up ideas and sentences into new paragraphs, but within those paragraphs it can be hard to keep up.

As an example, your first sentence could be rewritten from

  This is really not the type of legacy that I want to leave behind on hackernews but then again, I have been vocal that I just write what I think. Literally. It has its flaws but I am not sugar coating it.
to

  This isn't the type of legacy I want to leave behind on Hacker News. I prefer to write in a stream-of-consciousness style. This approach has its flaws, but it feels more natural to me.
Notice I trimmed some unnecessary words such as "really", split up a sentence, removed an unnecessary conjunction, added a comma before the "but" since the sentence contains two independent clauses.

I replaced "I am not sugar coating it" with what I feel is closer to your intended communication. "I'm not sugar coating it" is directed towards the reader and might be interpreted as antagonistic, whereas "it feels more natural to me" is directed towards yourself and can't be misconstrued.

I also compacted the phrase, "but then again, I have been vocal that I just write what I think" to "I prefer to write in a stream-of-consciousness style". The original phrase turns the reader around a bit, it takes a moment to derive intent.

The second phrase reads in a balanced way, `subject -> predicate -> verb -> preposition -> adjective -> noun`. One main clause and a complement, compared to three entire separate clauses in the original phrase. The second phrase flows down well hierarchically, and is easy to follow, while the original phrase turns the reader around and causes real, measurable fatigue when interpreting your communication.

Does this help?

Re: End of an era for me: no more self-hosted git

#199
post #25

Does anyone know what's the deal with these scrapers, or why they're attributed to AI? I would assume any halfway competent LLM driven scraper would see a mass of 404s and stop. If they're just collecting data to train LLMs, these seem like exceptionally poorly written and abusive scrapers written the normal way, but by more bad actors. Are we seeing these scrapers using LLMs to bypass auth or run more sophisticated…

I would love to understand this. Just a few years ago badly behaved scrapers were rare enough not to be worth worrying about. Today they are such a menace that hooking any dynamic site up to a pay-to-scale hosting platform like Vercel or Cloud Run can trigger terrifying bills on very short notice. "It's for AI" feels like lazy reasoning for me... but what IS it for? One guess: maybe there's enough of a market now for…

I'm currently on a free trial of ChatGPT, and one new thing I didn't regularly see before* is that longer tasks perform a lot of web searches when generating non-trivial results.

I wonder if this is part of it? It's not (just) DDOS by crawlers, it's DDOS by the users themselves triggering (albeit indirectly) far more requests than a human normally would? I've seen that happen in a different context, over a decade ago now.

* old models would do this sometimes when you ask for whatever the "deep research" mode was called, but this now seems to happen a lot more and involve a lot more fetches

Re: End of an era for me: no more self-hosted git

#200
post #50

Earlier quoted context omitted.

The dirty secret is a lot of them come through "residential proxies", aka backdoored home routers, iot devices with shitty security, etc. Basically the scrapers who are often also third party, go to these "companies" and buy access to these "residential proxies". Some are more... considerate than others. Why? Data. Every bit of it is it might be valuable. And not to sound tin foil hatty, but we are getting closer to…

How can I detect if my router is backdoored, or being used as a residential proxy?

[dead]
Post reply on HN