pretty crazy that this article is seemingly written by an AI used a detector and it is 95% confident its generated
A year of fighting scrapers on my 1.5 million-page website
191–200 of 461 posts
Re: A year of fighting scrapers on my 1.5 million-page website
#192Earlier quoted context omitted.
The best way to read the information on the internet today is via a LLM. Just like you don't process raw data in a data warehouse or datalake by looking at tables, you use a SQL or BI tool, to process all the information on the internet you need to tool to digest it for you. For many, today that tool is a LLM chat interface or agent.
This is such a grim thing to read.
Re: A year of fighting scrapers on my 1.5 million-page website
#193Earlier quoted context omitted.
The best way to read the information on the internet today is via a LLM. Just like you don't process raw data in a data warehouse or datalake by looking at tables, you use a SQL or BI tool, to process all the information on the internet you need to tool to digest it for you. For many, today that tool is a LLM chat interface or agent.
This is such a grim thing to read.
Re: A year of fighting scrapers on my 1.5 million-page website
#194Earlier quoted context omitted.
> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.
The social contract of accepting scraping for visibility was always tenuous and is now fully dead. But it’s shocking the number of people who are essentially victim blaming here. You should spend YOUR time to optimize your free website so MY use case is not impacted. How is that not incredibly selfish on its face?
Re: A year of fighting scrapers on my 1.5 million-page website
#195The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and move on with my life.
Re: A year of fighting scrapers on my 1.5 million-page website
#196Earlier quoted context omitted.
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; in this case you are not a user, you are a bot user.
you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…
Re: A year of fighting scrapers on my 1.5 million-page website
#197Earlier quoted context omitted.
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; in this case you are not a user, you are a bot user.
you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…
Identified by the user agent header. Which most bots fake or leave out to increase the chance of getting where they are not wanted.
Your bot is a good bot? Great, let us know when you've dealt with all the bad bots and we'll open the doors to the remaining (good) bots again.
Re: A year of fighting scrapers on my 1.5 million-page website
#198The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
Re: A year of fighting scrapers on my 1.5 million-page website
#199> There's a real conflict here. I want Google and Bing and DuckDuckGo to crawl my site and send me new readers. But I don't want everyone else strip-mining it. Kinda sounds like we're missing a peer to peer network here. Instead of downloading the same data over and over again we can just download it once and then share it. Wouldn't that be better for everyone involved? It would also function as a distributed WayBack…
The problem there is trust
The domain name system is now just a rent seeking scheme. It was great as a temp solution but over time it has deleted more content than preserved. I might in theory be billed for having a country name but I don't pay for a city, street name, house number or postal code. Online you should be able to move your widget shop to widget street. Can bill people who want to live on real estate street or used car street and/or set some requirements.
Re: A year of fighting scrapers on my 1.5 million-page website
#200I stopped posting to my website. Why should it be so much work to stop this theft?
Would it be helpful to have geofencing and regulation?