Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

281–290 of 374 posts

Re: An update on Wayback Machine access

#281

Earlier quoted context omitted.

Then make agent friendly content. Take the text and make a markdown version.

People doing this say it makes things worse because then the bots download both.

Because not enough people do this earnestly, and many more do it maliciously (bot endpoints that lie, or provide significantly less information than people endpoints) or put it behind a business contract (yes, APIs), so the bots or agents can't trust it in general.

Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason for the tension is simply that most websites and businesses on-line rely on users wasting their time, and cannot abide any form of end-user automation. Their business plans hinge on their ability to force themselves on you.

Re: An update on Wayback Machine access

#282
post #115

Earlier quoted context omitted.

AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.

If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS`

The "Bros"? Maybe.

I wouldn't want them to. The whole point of using agents to do stuff on the web for me, is for them to do the stuff on the web for me.

This is the reverse of "do not track" case. It'll not be effective because every service will set it to DISALLOW by default anyway, because it costs them nothing, and for most services, it actually is what they want anyway - most of businesses on the web are making money on wasting people's time, and for that, they need to force themselves on people; end-user automation defeats that, so they actively fight it (and complain a lot).

Re: An update on Wayback Machine access

#283

Earlier quoted context omitted.

I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.

I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.

It wouldn't be behind their back, like I said, "Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it."

Re: An update on Wayback Machine access

#284

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.

This! But only wayback machine, other services I'm not sure.

But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.

Re: An update on Wayback Machine access

#285

Earlier quoted context omitted.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them

Because there's no working alternative.

Re: An update on Wayback Machine access

#286
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out. Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you…

[flagged]

Re: An update on Wayback Machine access

#287

Earlier quoted context omitted.

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…

Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.

Re: An update on Wayback Machine access

#288

Earlier quoted context omitted.

One site was doing that which archive in their name but wasn't apart of archive.org

archive.is or archive.today? Why being shy in the era of stealing AI?

In some parts of the internet you can't mention a pirate site (or left wing stuff, anything sexual, or Palestine) without being banned. HN isn't one of them, but people have learned to be overly cautious.

Re: An update on Wayback Machine access

#289
post #92

Earlier quoted context omitted.

They're different sites, with different goals, run by different people.

That provide the same functional service.... Hence, distinction without a difference.

No they don't. Archive.org is co-operative, it respects robots.txt and allows deletion. It's also very slow. Archive.* is adversarial and archives sites that don't like it. That's why the FBI is trying to take it down.

Re: An update on Wayback Machine access

#290

Earlier quoted context omitted.

I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.

It wouldn't be behind their back, like I said, "Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it."

You wouldn't need archive.org for that - you could negotiate with the actual website. I think they'd demand quite a lot of money.
Post reply on HN