Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

171–180 of 374 posts

Re: An update on Wayback Machine access

#171
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point.

Open access doesn't seem sustainable.

But I might just grumpy about spending another hour this week adjusting rules to prevent bots.

Re: An update on Wayback Machine access

#172

Earlier quoted context omitted.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

Ironically, your framing of the situation is infinitely more disingenuous to push a personal agenda versus anything the archive.today guy did.

Re: An update on Wayback Machine access

#173
post #111

Earlier quoted context omitted.

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.

The minuscule traffic generated by the wayback machine, which serves to preserve the content for years to come, is completely incomparable to the scrapers that hammer every single href linked on a website.

Re: An update on Wayback Machine access

#174
post #92

Earlier quoted context omitted.

That provide the same functional service.... Hence, distinction without a difference.

Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.

Exactly this

archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.

Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.

It's nice to have both options. When I archive a site, I usually use both for added resiliency.

Re: An update on Wayback Machine access

#175

Earlier quoted context omitted.

What if they charged money? Is it something you'd pay for?

I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.

I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.

Re: An update on Wayback Machine access

#176

Earlier quoted context omitted.

Do they offer bulk torrent downloads as an alternative?

Once upon a time some people explored backing up the Internet Archive. However, that experiment ended. They mention there were some learnings and they then say: > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to…

The internet archive's decentralization project is paused as far as I can tell. They have too many things to do and too little funding to do it all. Their current strategy seems to be establishing new legal entities outside the us like in Canada and Switzerland, but they don't accept web traffic even though they hold full copies of the internet archive. There used to be a full copy in Egypt at the library of Alexandria and another in the Netherlands. Not sure if they're still in use, but they did accept web traffic. They hold a decentralized web camp every year in the middle of a forest

Re: An update on Wayback Machine access

#177
post #109

Earlier quoted context omitted.

They get marked as inaccessible, but still exist in IA data

is it possible to access somehow? it seems the site got excluded because of robots.txt set by some domain squatter, not manually.

They have all the WARC file cataloged on their site outside the way back machine

Re: An update on Wayback Machine access

#179

Earlier quoted context omitted.

Ah, the old ad hominem attack. How refreshing. But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but". It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

Is it really an ad hom if he doesn't know the hom? Their reply is 100% based on the content of your post.

If you make up a person to insult then yeah it's still ad hominem.

Re: An update on Wayback Machine access

#180

Earlier quoted context omitted.

Gatekeeping information is not the solution

Why not? If its the difference between the information being available at all, then I choose login any day of the week.

> If its the difference between the information being available at all

That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.

Post reply on HN