Earlier quoted context omitted.
I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures.
you really revved the weebs with this one
A year of fighting scrapers on my 1.5 million-page website
341–350 of 453 posts
Re: A year of fighting scrapers on my 1.5 million-page website
#342Re: A year of fighting scrapers on my 1.5 million-page website
#343Earlier quoted context omitted.
Why not ? Can google be trusted ? Or facebook ? How do people have been accessing the internet for the last decade you reckon and how's this any different ?
Google (as it originally) existed wasn't a source of information, it was an *index* of it. You were trusting google as a source for {search_query}, you were trusting the links it gave as a source (based on your own evaluation). LLMs are fundamentally different, because you are trusting the software to actually generate the information in a truthful way. If you use google (sans AI), you're putting some trust in their…
Re: A year of fighting scrapers on my 1.5 million-page website
#344Earlier quoted context omitted.
Don't tell this person about translation plugins. They change everything! The end user never reads what the website actually served!
It's probably a pretty common misconception that websites should have total control over how their content is viewed, but it's strange to see that on a website like this. Part of the reason apps are so popular with companies is that it gives them control over how content is presented which is something webpages were never intended to give them.
Re: A year of fighting scrapers on my 1.5 million-page website
#345The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots…
How do you monetise bot traffic?
I didn't ask the not to use this site, my query was broad and complex and would take me days to do it myself. I wouldn't
I'm quite certain they earn a hearty commission off it, and I think it was mostly possible because the site was "friendly" to bots. Otherwise I probably wouldn't choose the site because it's never any of my top choices when I look for this myself.
So, maybe you monetise it like this? You asked, I answered. Doesn't fit every site or business profile.
Re: A year of fighting scrapers on my 1.5 million-page website
#346Earlier quoted context omitted.
It's probably a pretty common misconception that websites should have total control over how their content is viewed, but it's strange to see that on a website like this. Part of the reason apps are so popular with companies is that it gives them control over how content is presented which is something webpages were never intended to give them.
The entire internet was built on advertising money. That's why any of this even exists.
Re: A year of fighting scrapers on my 1.5 million-page website
#347Earlier quoted context omitted.
I guess the disconnect here is a bunch of HN'ers believing professional companies and websites want to attach their branding to a sexualized anime character and that they are willing to pay to remove it. Which one then wonders why they would install it in the first place.
Sexualized? It’s just a cartoon/anime character holding a magnifying glass
Re: A year of fighting scrapers on my 1.5 million-page website
#348Earlier quoted context omitted.
> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. No; in this case you are not a user, you are a bot user.
you know what a browser is called by the web site? check the header that tells the version. USER AGENT not user, an agent on behalf of the user. the web was intended to be usable by agents belonging to users. in our path of the multiverse, the main agent we think of happens to be live interactive browser. but that need not have been, nor will it be in the future, the dominant USER AGENT. for a glimpse at one possible…
Can we build other monetization? Maybe. There's certainly proposals. It's real hard when there's too many layers between the user and the output though. I suspect the solutions will be worse than what we have now. For now the answer is "just paywall", but given your invocation of "corpo" here I suspect that's not an outcome you'd be too keen on ;)
Re: A year of fighting scrapers on my 1.5 million-page website
#349Earlier quoted context omitted.
Google (as it originally) existed wasn't a source of information, it was an *index* of it. You were trusting google as a source for {search_query}, you were trusting the links it gave as a source (based on your own evaluation). LLMs are fundamentally different, because you are trusting the software to actually generate the information in a truthful way. If you use google (sans AI), you're putting some trust in their…
Pagerank is just a popularity contest, doesn't say anything about trust. AI probably does a better job determining trustworthyness than pagerank.
Re: A year of fighting scrapers on my 1.5 million-page website
#350Earlier quoted context omitted.
Sexualized? It’s just a cartoon/anime character holding a magnifying glass
In any case, if you run a professional website, this immediately comes off as juvenile and/or amateurish. And am y people just assume it's part of your website.