Live data from Hacker News

Wikipedia is struggling with voracious AI bot crawlers

engadget.com

71–80 of 105 posts

Re: Wikipedia is struggling with voracious AI bot crawlers

#71
post #47

Earlier quoted context omitted.

This is an interesting model in general: free for humans, pay for automation. How do you enforce that though? Captchas sounds like a waste.

Any plan that starts with "Step one: Apply the tool that almost perfectly distinguishes human traffic from non-human traffic" is doomed to failure. That's whatever the engineering equivalent of "begging the question" is, where the solution to the problem is that we assume that we have the solution to the problem.

Identity verification is not that far fetched these days. For europeans you got eIDAS and related tech, some other places have similar stuff, for rest of world you can do video based id checks. There are plenty of providers that handle this, it's pretty commonplace stuff.

Re: Wikipedia is struggling with voracious AI bot crawlers

#72

The weird thing is their own data does not reflect this at all. The number of articles accessed by users, spiders and bots alike has not moved significantly over the last few years. Why these strange wordings like "65 percent of the resource-consuming traffic"? Is there non-resource consuming traffic? Is this just another fundraising marketing drive? Wikimedia has been know to be less than truthful wrt their funding…

The graph you linked seems to be about article viewing ("page views", like a GET request to https://en.wikipedia.org/wiki/Democracy for example), while the article mentions multimedia content, so fetching the actual bytes of https://en.wikipedia.org/wiki/Democracy#/media/File:Economis... for example, which would consume more content than just loading article pages, as far as I understand.

Re: Wikipedia is struggling with voracious AI bot crawlers

#73
post #71
post #47

Earlier quoted context omitted.

Any plan that starts with "Step one: Apply the tool that almost perfectly distinguishes human traffic from non-human traffic" is doomed to failure. That's whatever the engineering equivalent of "begging the question" is, where the solution to the problem is that we assume that we have the solution to the problem.

Identity verification is not that far fetched these days. For europeans you got eIDAS and related tech, some other places have similar stuff, for rest of world you can do video based id checks. There are plenty of providers that handle this, it's pretty commonplace stuff.

This is a terrible idea. Consider how that will be abused as soon as a government who hates you comes in to power.

Re: Wikipedia is struggling with voracious AI bot crawlers

#74
> "expansion happened largely without sufficient attribution, which is key to drive new users to participate in the movement."

these multi$B corps continue to leech off of everyone's labors, and no one seems able to stop them; at what level can entities take action? the courts? legislation?

we've basically handed over the Internet to a cabal of Big Tech

Re: Wikipedia is struggling with voracious AI bot crawlers

#77
post #31

Wikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.

BTW, I do not know why you are getting downvoted, this is a real concern that someone needs to tackle one day.

Real concern or not, this is not related to the discussion at hand, which is AI crawlers hammering Wikipedia, which is related to AI crawlers hammering everything these days. Here's the concern at hand.

I would like to read on Wikipedia corruption with quality sources (in a separate HN post, which would probably be successful), but that's not quite on-topic here. Not only it's off-topic and borderline whataboutism, it's also not sourced, so the comment doesn't actually help someone who isn't in the knows. Thus, as is, it's not much interesting and kinda useless.

These reasons are probably why it has been downvoted: off topic, not helping, not well researched.

Re: Wikipedia is struggling with voracious AI bot crawlers

#78
post #31

Wikipedia spends 1% of its budget on hosting fees. It can spend a bit more given the rest of their corruptions.

> the rest of their corruptions

the "corruption" accusations are mostly BS and the usual ideological differences

I'll take Wikipedia, with all its warts, over $BigTech and $VC-driven (==Ad-driven) companies/orgs any day, and it's not even close.

Re: Wikipedia is struggling with voracious AI bot crawlers

#79

Earlier quoted context omitted.

BTW, I do not know why you are getting downvoted, this is a real concern that someone needs to tackle one day.

It’s getting downvoted because the parent comment aligns with what Elon said about Wikipedia; so it’s a knee jerk reaction. Though the sentiment is factual. Previous discussion: (2022) https://news.ycombinator.com/item?id=32840097

> sentiment is factual

the sentiment might exist, but that doesn't mean it's based on facts

that discussion you linked to can be broken down into:

- people upset because they thought Wikipedia was almost bankrupt and it turns out its not (though Wikipedia never claimed to be in its fund-raising)

- people upset because they see too many requests for donations

- people upset because Wikimedia execs are getting "high" salaries (though they are much _much_ lower than at private co's)

- people upset because they think Wikipedia spreads "left-wing ideologies"

None of this has anything to do with "corruption".

Re: Wikipedia is struggling with voracious AI bot crawlers

#80

what's the best way to stop the bots? cloudflare?

Why should we stop the bots? Wikipedia supposedly wants the world to have this free information, a bit is just another way of supporting that goal.

Not at all costs though, including disrupting access to said free information.

And this free information is not free from rights to respect neither, it's under CC-BY-SA, which requires attribution and sharing under the same conditions, the kind of "subtleties" and "details" with which AI companies have been wiping their big arses.

Post reply on HN