Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

241–250 of 450 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#241

Earlier quoted context omitted.

I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.

you really revved the weebs with this one

Nah, I don't especially like the logo myself when I come across those sites, but I do like that "brand" people who want to make money off it and not pay dislike it even more.

Re: A year of fighting scrapers on my 1.5 million-page website

#242

Earlier quoted context omitted.

you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…

Can you recommend a specific home assistant community to check out?

r/homeassistant

Re: A year of fighting scrapers on my 1.5 million-page website

#243
post #9

Is there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.

I have a very low traffic blog and I just use the nearlyfreespeech to host it. So far it’s been less than $1

Re: A year of fighting scrapers on my 1.5 million-page website

#244

Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/

Because of this wasteful crap the internet is so slow nowadays...

Try opening gcc bug tracker on your phone: https://gcc.gnu.org/bugzilla/

Re: A year of fighting scrapers on my 1.5 million-page website

#246
post #211

Earlier quoted context omitted.

you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…

[flagged]

So instead of adding of addressing his point about User Agents, you decide its better to make a low level bullshit comment about the very last sentence to stand up for... checks notes... Corporations and dismiss everything else. Top notch quality content that for sure added to the conversation.

Re: A year of fighting scrapers on my 1.5 million-page website

#247
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

Substitute the word “website” for book, or training course, or documentary or published paper, or patent or one of hundreds of examples.

Now you see the problem.

Re: A year of fighting scrapers on my 1.5 million-page website

#248
post #206

Earlier quoted context omitted.

I don't disagree, but there is a sliding scale here. For instance, I wanted to buy a piece of equipment the other day from a local company for a specific usecase. I wanted to find a specific price/weight/specs ratio, and asked an llm to loop through the 20 or so items, fetch their page and calculate and present some values for each. This then led to me going and buying the one I found. So the llm was mainly just an e…

I don’t see the sliding scale - if the website didn’t want a bot they could have blocked you. They clearly don’t mind so what’s the issue? You still used a bot.

[deleted]

Re: A year of fighting scrapers on my 1.5 million-page website

#249

Earlier quoted context omitted.

How do you monetise bot traffic?

You take your site off the public internet and paywall it off. Websites on the open internet should be accessible and ideally the goal should be to share something cool with the world, not just to make yourself rich.

> You take your site off the public internet and paywall it off

I did that. I even printed and bound it in a bunch of dead trees and put it up for sale. The LLMs still stole it.

Re: A year of fighting scrapers on my 1.5 million-page website

#250

Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/

Because of this wasteful crap the internet is so slow nowadays... Try opening gcc bug tracker on your phone: https://gcc.gnu.org/bugzilla/

This is like getting angry at cookie banners instead of all the companies tracking and selling your data.

You're complaining about the symptom (needing to have these checks) not the cause (if they don't, 99% of their traffic will be bots, the site will slow to a crawl and be unusable anyway).

In any case, I saw the dumb anime girl for about 2s then the site loaded. Not a big deal.

Post reply on HN