Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

441–450 of 453 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#441
post #107

Earlier quoted context omitted.

Yeah this is one of those things that hurts some real users and AI scrapers have figured out how to bypass. Not worth it IMO.

How do you figure it hurts real users? The amount of compute/energy used on the proof of work is pretty minimal. You're using more when you watch a YouTube video or browse a JS-heavy web app. Of course a sophisticated scraper can "figure out" how to bypass. It isn't trying to be foolproof, it's adding an extra cost to deter massive amounts of bot traffic. I put it in front of my hobby project because I can't afford t…

Over a slow mobile connection Anubis doesn't load at all, even if the actual website would load just fine in seconds. Hence the user is locked out from the website.

Re: A year of fighting scrapers on my 1.5 million-page website

#442

Earlier quoted context omitted.

I disagree. I personally consider these my biggest problems with the Web: - Bias, specifically commercial bias - Webpage formatting: every website looks different, hides the information I want in different places. Sure, it "looks pretty", but I don't want pretty; I want info - Scams/SEO/etc... The LLMs are very good at reading from multiple sources, parsing, and presenting only the data in a consistent format. There…

Why even use a LLM, just buy a cookbook and you can even waste less time and know the information is valid. And again if your logic is consistent, google routes your search to an LLM automatically and renders the reply - that will save as much time and do the same thing since LLMs are so good at parsing and presenting information. So why should I pick an LLM over google's inbuilt LLM or an actual cookbook on the desk…

From Google's perspective you might as well be a bot because when you search there you don't see ads. If you don't have adblock installed, your parent comment applies because you get worse experience as LLMs don't serve ads.

Re: A year of fighting scrapers on my 1.5 million-page website

#443

Earlier quoted context omitted.

How is that grim? It's the dream of the Semantic Web coming true, just by different means than planed.

Because the economic, social, and climate impacts are at best uncertain and at worst devastating to the majority of human beings on the planet. It's always surprising to me that people don't intuitively separate the usefulness of AI from its risks. It feels like everyone has to be a doomer or a booster.

> at worst devastating

This is such a low bar literally anything would clear it. Even tomato farming.

Re: A year of fighting scrapers on my 1.5 million-page website

#444
post #343

Earlier quoted context omitted.

Pagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.

Do you think the LLM reads every page on the internet before generating your answer? Of course not. What happens is that it use some sort of ranking algorithm to pick the pages that are most likely to answer your query and reads *them* (at best. At worst it just makes something up). You aren't avoiding the problems with ranking algorithms by asking an LLM, you're taking all of those problems, adding more problems on…

> Do you think the LLM reads every page on the internet before generating your answer? Of course not.

The way LLMs are trained the answer is of course yes, but they don't remember them all exactly, of course.

Re: A year of fighting scrapers on my 1.5 million-page website

#445
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and move on with my life.

The situation is getting worse in some OSS communities too. I go to a bugtracker just to read the discussion on the issue I am facing, often to understand what's the roadmap to fix it if any, and some of them immediately demand me to enable JS and solve a captcha. Even GitHub doesn't do it!

Re: A year of fighting scrapers on my 1.5 million-page website

#446
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. And this is a complaint? I find nothing in this statement to sympathize with whatsoever. Isn't the contract of, essentially everything, that effort is required to obtain something worthwhile? Please let me know where I'm getting this long-standing, fairly fundamental understanding of the wo…

Because the effort you are talking about is completely unnecessary. Imagine every grocery store would require you do a little dance when you buy a carton of milk. Your vision needs some narrowing (which I believe will justify the parent point), because as-is it has obvious counterexamples.

I believe your problem is that "effort" is unspecified. "some effort" would make the statement correct, but some effort does not justify arbitrary effort, therefore you have no point here.

Re: A year of fighting scrapers on my 1.5 million-page website

#447
post #55
post #9

Is there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.

"A $5/month VPS won't cut it anymore." Are you speaking from experience, or inferring from articles like this? I serve a static site on the lowest Linode $5/month VPS and it is grotesquely overprovisioned for that use case. It is not the case that every site is getting slammed every second by hundreds of requests per second. Now, if you have some sort of dynamically-computed website that is generated by a slow script…

Yes! There are linux boxes that stip and resume like a lambda, but really linux. And you pay only for what you use. Look for e.g. at https://shellbox.dev, it starts at $0.02/hr

Re: A year of fighting scrapers on my 1.5 million-page website

#448
post #321

Earlier quoted context omitted.

Scenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense

You’re missing the part about the human who interacts with the webpage.

Doesn't matter.

Even by the typical robots.txt definition, a bot has to operate more or less fully autonomously. If the user initiated the request, it doesn't count, and it shouldn't follow robots.txt.

Re: A year of fighting scrapers on my 1.5 million-page website

#450

Earlier quoted context omitted.

You are assuming that because you are giving the operator money that you are entitled to use their website however you please -- that's not how it works. If you enter a brick & mortar establishment and break their rules, even as a paying customer, you might risk e.g. being kicked out and banned, depending on the behavior. This is not unique to e-commerce To be fair, I do not generally support wholesale banning of scr…

That's an incredibly dishonest characterization of what I wrote. Please re-read the guidelines of the site, they say exactly not to do that. > saying that you should be entitled to behave as you please because you are a paying customer No, I'm saying my usage of the site was exactly as if I used my browser agent instead of the llm agent. It visited the category for the product I was interested it, then it visited the…

The site guidelines also say to assume good faith, to converse curiously, edit out swipes, not go fulminate, and to avoid grandstanding.

Before calling me dishonest, did you stop to consider that maybe you didn't understand my comment, or that you weren't clear with your meaning?

I genuinely don't think you understood what I wrote.

If I operate a business, you are not necessarily entitled to transact with me how you please. I can implement whatever rules I want so long as I am not outright committing discrimination against a legally protected category and so long as I am upholding other fiduciary obligations.

For all you know, the site operator is opposed to AI agents on moral grounds, or has beef with the developer of the specific LLM interface you are using.

More practically, an e-commerce website often won't make a profit if a purchase is too small. If they find that LLM users are easily bypassing their upselling tactics that usually result in higher basket sizes, then ideally they would adapt their approach to LLMs or implement a minimum order size, but they are also perfectly within their rights to simply ban you.

I am not saying that this is the likely reason -- just that there are reasons, many, and often they aren't very obvious unless you are in their shoes.

Every kind of business does this sort of thing! A software consultant or designer might only take money from healthcare companies, and not talk to anyone offering (apparently, to a naive outsider) equivalent work in a different industry. A restaurant might implement a dress code. A cab driver might insist you not wear strong perfume.

Some rules that externally seem flippant and unreasonable exist for reasons that are entirely sensible, but also entirely nonobvious externally. In any event, I don't agree that your usage is equivalent to a manual session, if it were, the operator wouldn't be able to detect it! Therefore they are within their rights to detect this and disallow it. There are plenty of reasons that they might do this, and unfortunately you probably don't have the information needed to interrogate those in a meaningful way

Post reply on HN