Live data from Hacker News

Creepy Crawlies

people.kernel.org

631–640 of 705 posts

Re: Creepy Crawlies

#631

Earlier quoted context omitted.

You don’t need to spend $5000 to obtain the hash rate of a $5000 device on a rental basis. You may have heard of this thing called “the cloud”. Obtaining very high hash rates is effectively free, largely as a side effect of the crypto bust. Not sure why anyone would characterize these scrapers im general as all being fly-by-night operations that don’t have two cents to scrape together.

And yet you have not addressed the point that many report that tricks like Anubis work. If they are so stupid an idea that they could never work, why do they seem to having the desired effect?

I never said the approach was so stupid it could never work. I just shared how it’s annoying and how I worked around my annoyance.

An even stupider approach would work exactly as well or better, like a form saying “type the letter y in this box to continue”. The only benefit is as a road bump that makes the site in any way custom. The moment anyone you are defending against so much as looks at the mechanics of solving the challenge it completely falls apart.

Re: Creepy Crawlies

#632
post #524

> because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge. This statement holds the core misapprehension behind Anubis. It’s not a ton of cycles. There is no difficulty setting that would be inconvenient for bots but usable for humans on mobile devices. I noticed the other day that lists.ffmpeg.org had moved to Anubis difficulty level 6, which takes ~180sec for my…

Does it really take 3 minutes on your iPhone? My pixel 8 does it in slightly less than a minute in Firefox.

iphone 16 and i gave up after about 30 seconds, it was maybe a quarter done

Re: Creepy Crawlies

#633

Time to F*$k the internet. The whole concept of anonymous IP addresses was broken but worked for a long time. Now it is just stupid. Just like domain names. (Are more names used by squatters than real?). And email as identity? Time to engineer solutions and create a new protocol layer. This is not a hard problem. It just requires that someone build a certification wall. The IETF should have done this long ago, right?

It's not a hard problem, the hard part is doing it while keeping a similar level of privacy/anonimity

Re: Creepy Crawlies

#634
post #141

> [...] when a source is guaranteed to be LLM-free, like the entire history of kernel commits [...] Is that really the case? It was my understanding that LLM-based agents were explicitly allowed as long as their users follow certain guidelines [1]? And more generally: Somehow the theory of "essentially all bot traffic is AI labs crawling the Internet for LLM training data" doesn't make sense to me at all. There are a…

The use of LLM agents was allowed relatively recently, you still have 20 years of guaranteed no llm commits in the history. As for personal LLMs being the cause of the traffic, I think you are overestimating the number of users of LLMs generally, and even then most of those users are not actually developers. And there are even fewer cases still where the llm being used actually needed something from the linux kernel, personally after quite heavy use of LLMs on linux, I have never seen it try to fetch the linux kernel directly once.

Re: Creepy Crawlies

#635

Earlier quoted context omitted.

> rather than moving to a challenge system that actually impacts scrapers At this point isn't it basically auth-only? Rant: (genuinely wondering too, and RFC, request for conversation) at this point don't we have Google, etc. basically doing Real World ID Verification, but without an open protocol backing, using it to corral users into their ecosystem and gather data, and leaving us without some open and distributed…

Botnets will just borrow your id.

Maybe, but at least you can more easily assign responsibility in this case. If you soft-block or hard-block someone because they're in a botnet, you can get them to change (expire the old ones) their credentials and start with clean systems. Maybe they'll realize their TV is part of a botnet if they have to keep doing it, or their PC has malware.

Re: Creepy Crawlies

#636
post #627

Earlier quoted context omitted.

I’m not a participant in this race. Are the AI companies worried about anything but their valuations? I don’t care about their valuations, but I do care about the risks that they are creating for the economy, society, and the technological advancement, at large. Micro transactions [in this case] are a great idea, these crawlers need to be taxed and made to pay for the unaccounted external costs. Furthermore, we need…

> Are the AI companies worried about anything but their valuations? >I don’t care about their valuations, but I do care about the risks that they are creating for the economy, society, and the technological advancement, at large. Doesn't the same thing apply to every company? And also every government? And also individual? We're all doing the "locally rational" thing, and globally doing... well... I think the main fa…

Yes, this applies to every company. The institutions have evolved a lot, and how we look at social safety and fairness have both evolved and devolved at the same time.

Companies, individuals, and governments have different roles. Rationality from their perspective is specific to their perspective, and I don’t believe just because evolution or physics are rational we should die off from a new virus or asteroid impact. That would be perfectly rational, but I don’t care about the global rationality of the universe.

What is best for my family and other families is not what is best for Sam Altman or Sataya, or Elon Mask.

We’re just talking about market players. What I’m saying is that these market forces need an outlet via some new infrastructure and regulation. Maybe it’s not micro-payments, maybe it’s some other new ideas.

I just don’t want the next century to be the century of East India AI sticking me in a Matrix pod, rationally.

Re: Creepy Crawlies

#637

Earlier quoted context omitted.

The crawlers are not AI. The crawlers are deterministic. They are collecting data to train AIs.

But gitweb is probably the second most used method of hosting a git repo and easily recognizable through heuristics. If it’s gitweb, fallback to git access and save everyone, including the crawler, time and resources.

Understood. The comment I replied to suggests the crawlers should figure this out on their own.

Re: Creepy Crawlies

#638

Instead of trying to block why not monetize? So the proof of work can be directed at something you can be paid for (bitcoin mining)?

If everyone had devices that could do some compute worth paying for, people would be doing the work all of the time. The problem is actually useful POW is a lot more expensive compute-wise, making it non viable for normal users

Re: Creepy Crawlies

#639
post #594

I've spent the last few days adding traps to one of my websites, ironically using LLMs of course, and I've been having quite a lot of fun doing it. Instead of the proof-of-work system of Anubis, I've gone down the iocaine route but implemented it in my application itself, as it's built in Elixir and causing problems for scrapers is really fun when it takes almost no server resources. Currently I trick bad scrapers in…

My team runs a quite popular website, #1 or #2 in the market depending on the region. Several million visits per day. Around June/July we got a 10x boost out of nowhere, and it started affecting performance for users, increased hosting costs, and random bursts would bring the website down. We spent some time trying out solutions, from Cloudflare and Anubis to AWS, but it ended up affecting real users, and we got comp…

In an odd sequence of events, I've sorta got a weird tiny following in China due to having a vintage camera stall in an antiques market here in Wales that gets shared on Xiaohongshu (https://en.wikipedia.org/wiki/Xiaohongshu) sometimes when I have Chinese customers. So I'd feel bad blocking the entire country myself, though thankfully I'm not being inundated with traffic from there yet - the most egregious bots I've caught so far are actually ones from Meta which ignore robots.txt.

Re: Creepy Crawlies

#640
This greatly implies online advertising is dead, that the internet as something useful to humans is soon becoming questionable.
Post reply on HN