Earlier quoted context omitted.
Yeah this is one of those things that hurts some real users and AI scrapers have figured out how to bypass. Not worth it IMO.
How do you figure it hurts real users? The amount of compute/energy used on the proof of work is pretty minimal. You're using more when you watch a YouTube video or browse a JS-heavy web app. Of course a sophisticated scraper can "figure out" how to bypass. It isn't trying to be foolproof, it's adding an extra cost to deter massive amounts of bot traffic. I put it in front of my hobby project because I can't afford t…
A year of fighting scrapers on my 1.5 million-page website
441–450 of 453 posts
Re: A year of fighting scrapers on my 1.5 million-page website
#442Earlier quoted context omitted.
I disagree. I personally consider these my biggest problems with the Web: - Bias, specifically commercial bias - Webpage formatting: every website looks different, hides the information I want in different places. Sure, it "looks pretty", but I don't want pretty; I want info - Scams/SEO/etc... The LLMs are very good at reading from multiple sources, parsing, and presenting only the data in a consistent format. There…
Why even use a LLM, just buy a cookbook and you can even waste less time and know the information is valid. And again if your logic is consistent, google routes your search to an LLM automatically and renders the reply - that will save as much time and do the same thing since LLMs are so good at parsing and presenting information. So why should I pick an LLM over google's inbuilt LLM or an actual cookbook on the desk…
Re: A year of fighting scrapers on my 1.5 million-page website
#443Earlier quoted context omitted.
How is that grim? It's the dream of the Semantic Web coming true, just by different means than planed.
Because the economic, social, and climate impacts are at best uncertain and at worst devastating to the majority of human beings on the planet. It's always surprising to me that people don't intuitively separate the usefulness of AI from its risks. It feels like everyone has to be a doomer or a booster.
This is such a low bar literally anything would clear it. Even tomato farming.
Re: A year of fighting scrapers on my 1.5 million-page website
#444Earlier quoted context omitted.
Pagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.
Do you think the LLM reads every page on the internet before generating your answer? Of course not. What happens is that it use some sort of ranking algorithm to pick the pages that are most likely to answer your query and reads *them* (at best. At worst it just makes something up). You aren't avoiding the problems with ranking algorithms by asking an LLM, you're taking all of those problems, adding more problems on…
The way LLMs are trained the answer is of course yes, but they don't remember them all exactly, of course.
Re: A year of fighting scrapers on my 1.5 million-page website
#445The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and move on with my life.
Re: A year of fighting scrapers on my 1.5 million-page website
#446The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. And this is a complaint? I find nothing in this statement to sympathize with whatsoever. Isn't the contract of, essentially everything, that effort is required to obtain something worthwhile? Please let me know where I'm getting this long-standing, fairly fundamental understanding of the wo…
I believe your problem is that "effort" is unspecified. "some effort" would make the statement correct, but some effort does not justify arbitrary effort, therefore you have no point here.
Re: A year of fighting scrapers on my 1.5 million-page website
#447Is there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.
"A $5/month VPS won't cut it anymore." Are you speaking from experience, or inferring from articles like this? I serve a static site on the lowest Linode $5/month VPS and it is grotesquely overprovisioned for that use case. It is not the case that every site is getting slammed every second by hundreds of requests per second. Now, if you have some sort of dynamically-computed website that is generated by a slow script…
Re: A year of fighting scrapers on my 1.5 million-page website
#448Earlier quoted context omitted.
Scenario one: you use software to connect to their server and download a webpage. You are a user. Scenario two: you use software to connect to their server and download a webpage. You are a "bot". Make it make sense
You’re missing the part about the human who interacts with the webpage.
Even by the typical robots.txt definition, a bot has to operate more or less fully autonomously. If the user initiated the request, it doesn't count, and it shouldn't follow robots.txt.
Re: A year of fighting scrapers on my 1.5 million-page website
#449Is there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.
Re: A year of fighting scrapers on my 1.5 million-page website
#450Earlier quoted context omitted.
You are assuming that because you are giving the operator money that you are entitled to use their website however you please -- that's not how it works. If you enter a brick & mortar establishment and break their rules, even as a paying customer, you might risk e.g. being kicked out and banned, depending on the behavior. This is not unique to e-commerce To be fair, I do not generally support wholesale banning of scr…
That's an incredibly dishonest characterization of what I wrote. Please re-read the guidelines of the site, they say exactly not to do that. > saying that you should be entitled to behave as you please because you are a paying customer No, I'm saying my usage of the site was exactly as if I used my browser agent instead of the llm agent. It visited the category for the product I was interested it, then it visited the…
Before calling me dishonest, did you stop to consider that maybe you didn't understand my comment, or that you weren't clear with your meaning?
I genuinely don't think you understood what I wrote.
If I operate a business, you are not necessarily entitled to transact with me how you please. I can implement whatever rules I want so long as I am not outright committing discrimination against a legally protected category and so long as I am upholding other fiduciary obligations.
For all you know, the site operator is opposed to AI agents on moral grounds, or has beef with the developer of the specific LLM interface you are using.
More practically, an e-commerce website often won't make a profit if a purchase is too small. If they find that LLM users are easily bypassing their upselling tactics that usually result in higher basket sizes, then ideally they would adapt their approach to LLMs or implement a minimum order size, but they are also perfectly within their rights to simply ban you.
I am not saying that this is the likely reason -- just that there are reasons, many, and often they aren't very obvious unless you are in their shoes.
Every kind of business does this sort of thing! A software consultant or designer might only take money from healthcare companies, and not talk to anyone offering (apparently, to a naive outsider) equivalent work in a different industry. A restaurant might implement a dress code. A cab driver might insist you not wear strong perfume.
Some rules that externally seem flippant and unreasonable exist for reasons that are entirely sensible, but also entirely nonobvious externally. In any event, I don't agree that your usage is equivalent to a manual session, if it were, the operator wouldn't be able to detect it! Therefore they are within their rights to detect this and disallow it. There are plenty of reasons that they might do this, and unfortunately you probably don't have the information needed to interrogate those in a meaningful way