Live data from Hacker News

A year of fighting scrapers on my 1.5 million-page website

patronview.com

341–350 of 451 posts

Re: A year of fighting scrapers on my 1.5 million-page website

#341

Earlier quoted context omitted.

I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.

you really revved the weebs with this one

It seems to come up every time this complaint arises. "I want to use an open source project without being associated with furries" isn't that unreasonable of a request.

Re: A year of fighting scrapers on my 1.5 million-page website

#343
post #237

Earlier quoted context omitted.

Why not ? Can google be trusted ? Or facebook ? How do people have been accessing the internet for the last decade you reckon and how's this any different ?

Google (as it originally) existed wasn't a source of information, it was an *index* of it. You were trusting google as a source for {search_query}, you were trusting the links it gave as a source (based on your own evaluation). LLMs are fundamentally different, because you are trusting the software to actually generate the information in a truthful way. If you use google (sans AI), you're putting some trust in their…

Pagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.

Re: A year of fighting scrapers on my 1.5 million-page website

#344

Earlier quoted context omitted.

Don't tell this person about translation plugins. They change everything! The end user never reads what the website actually served!

It's probably a pretty common misconception that websites should have total control over how their content is viewed, but it's strange to see that on a website like this. Part of the reason apps are so popular with companies is that it gives them control over how content is presented which is something webpages were never intended to give them.

The entire internet was built on advertising money. That's why any of this even exists.

Re: A year of fighting scrapers on my 1.5 million-page website

#345
post #101

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…

How do you monetise bot traffic?

Ibwas looking for something specific late last year, and I made three bookings (ca €500 each) for places found on the $site with the $bot.

I didn't ask the not to use this site, my query was broad and complex and would take me days to do it myself. I wouldn't

I'm quite certain they earn a hearty commission off it, and I think it was mostly possible because the site was "friendly" to bots. Otherwise I probably wouldn't choose the site because it's never any of my top choices when I look for this myself.

So, maybe you monetise it like this? You asked, I answered. Doesn't fit every site or business profile.

Re: A year of fighting scrapers on my 1.5 million-page website

#346

Earlier quoted context omitted.

It's probably a pretty common misconception that websites should have total control over how their content is viewed, but it's strange to see that on a website like this. Part of the reason apps are so popular with companies is that it gives them control over how content is presented which is something webpages were never intended to give them.

The entire internet was built on advertising money. That's why any of this even exists.

The internet existed and thrived long before the advertisers infested it. The internet was different, but in many ways better. It was still useful and amazing. That was why they came. It would be still be useful and amazing if every advertiser on Earth disappeared tomorrow and took their ads with them.

Re: A year of fighting scrapers on my 1.5 million-page website

#347

Earlier quoted context omitted.

I guess the disconnect here is a bunch of HN'ers believing professional companies and websites want to attach their branding to a sexualized anime character and that they are willing to pay to remove it. Which one then wonders why they would install it in the first place.

Sexualized? It’s just a cartoon/anime character holding a magnifying glass

In any case, if you run a professional website, this immediately comes off as juvenile and/or amateurish. And am y people just assume it's part of your website.

Re: A year of fighting scrapers on my 1.5 million-page website

#348
post #117

Earlier quoted context omitted.

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; in this case you are not a user, you are a bot user.

you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…

It's not a simple semantic distinction. The mechanism matters because the monetization is built around it, and the monetization (as we've built it) only works because there's a human on the other end.

Can we build other monetization? Maybe. There's certainly proposals. It's real hard when there's too many layers between the user and the output though. I suspect the solutions will be worse than what we have now. For now the answer is "just paywall", but given your invocation of "corpo" here I suspect that's not an outcome you'd be too keen on ;)

Re: A year of fighting scrapers on my 1.5 million-page website

#349
post #343

Earlier quoted context omitted.

Google (as it originally) existed wasn't a source of information, it was an *index* of it. You were trusting google as a source for {search_query}, you were trusting the links it gave as a source (based on your own evaluation). LLMs are fundamentally different, because you are trusting the software to actually generate the information in a truthful way. If you use google (sans AI), you're putting some trust in their…

Pagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.

Do you think the LLM reads every page on the internet before generating your answer? Of course not. What happens is that it use some sort of ranking algorithm to pick the pages that are most likely to answer your query and reads *them* (at best. At worst it just makes something up). You aren't avoiding the problems with ranking algorithms by asking an LLM, you're taking all of those problems, adding more problems on top, and pretending that this is somehow better.

Re: A year of fighting scrapers on my 1.5 million-page website

#350
post #347

Earlier quoted context omitted.

Sexualized? It’s just a cartoon/anime character holding a magnifying glass

In any case, if you run a professional website, this immediately comes off as juvenile and/or amateurish. And am y people just assume it's part of your website.

Money can be used to purchase goods or services, such as an unbranded version.
Post reply on HN