Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

151–160 of 374 posts

Re: An update on Wayback Machine access

#151
post #97

Earlier quoted context omitted.

> Asking users to email them with details of their OS, browser, IP address is just crazy. It's surely to serve as data to help tell humans apart from bots. > Changes made by IA shouldn't become my responsibility. They're a free service. It's ultimately not their responsibility to service you either.

Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request. IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

No, it's more like you're requesting something from them and they're telling you they may need some technical, non-personally-identifiable info from you to fulfill your request.

Re: An update on Wayback Machine access

#153
post #115
post #77

Earlier quoted context omitted.

Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.

AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.

If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS`

Re: An update on Wayback Machine access

#154
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)

Re: An update on Wayback Machine access

#155

Earlier quoted context omitted.

No matter what they did, you'd have a new excuse for why you won't pay.

Ah, the old ad hominem attack. How refreshing. But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but". It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

Is it really an ad hom if he doesn't know the hom?

Their reply is 100% based on the content of your post.

Re: An update on Wayback Machine access

#156
post #45
post #35

AI companies should pay billions to wayback machine for access

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.

Re: An update on Wayback Machine access

#157
post #55
post #45

Earlier quoted context omitted.

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.

I'd have a lot less of a problem with AI if everything that went into their training was public domain and made easily available to anyone for any use. It'd feel less like AI companies were just stealing the work of others and charging for it.

Re: An update on Wayback Machine access

#158

I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.

Plenty of threads on HN about this, Anubis does not work.

It always seems to keep me, a normal human, locked out of any site that uses it.

Re: An update on Wayback Machine access

#159

Can't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?

Not while the new dukes of the internet wielding massive armies of hijacked smart TVs have a better time browsing the web than I have; as a mere peasant with just a few IP addresses. There would be no reason to sign up and pay up for bulk access, unless open access is shut down.
Post reply on HN