Live data from Hacker News

The Internet Archive is back online

arstechnica.com

61–70 of 107 posts

Re: The Internet Archive is back online

#61
post #2

It would be better if the Internet Archive was decentralized without a central point of failure, maybe run on something like bittorent.

I'm working on this, ArchiveBox v0.8 adds the beginnings of a content addressable store, with plans for bittorrent-backed instance-to-instance sharing in a later version.

I think Archive.org should still exist too (and ArchiveBox donates + submits URLs to Archive.org too), but having a self-hosted option where you can archive personal stuff that requires a login, and do P2P sharing with with fine grained permissions is a gap that should be filled.

Aiming to archive the entire internet is Archive.org's goal, aiming to archive the part of the internet YOU care about is our goal.

Re: The Internet Archive is back online

#62
post #3
post #2

It would be better if the Internet Archive was decentralized without a central point of failure, maybe run on something like bittorent.

I kind of agree, but the way the internet is going, with everyone being behind carrier-grade nat, it's not much of a decentralized network of computers anymore, not to mention all the kids with their laptops and tablets not even hosting anything :(

There are ways around this, I've experimented with setting up a cluster of ArchiveBox instances that share snapshots over Tailscale. Tailscale lets users sign up for free accounts, and you can share machines between separate accounts. A (CGNAT-compatible) decentralized invite-only network could concievably spread that way.

Re: The Internet Archive is back online

#63
post #13

Earlier quoted context omitted.

I don't think they censor anything, strictly archiving. Do you know of any instance in which they censored a site?

Kiwi farms

There are people that maintain "non-public archives" of stuff like that for litigation, long-term archival storage (think sealed boxes intended for future generations of historians. (Libraries, laywers, journalists can run their own WebRecorder, Perma.cc, ArchiveBox, etc. instances)

I think that's a reasonable middle ground, we don't necessarily need every single piece of heinous content mirrored for free access 24/7 the moment it appears anywhere on the internet, as long as there is some historic record somewhere that's probably ok.

Re: The Internet Archive is back online

#65
post #58

Earlier quoted context omitted.

If I help seed this DWeb and it turns out it has some copyrighted materials in it, will I be potentially held liable?

You're always responsible for what you, yourself and your computer does. There is a chance EFF/some other organization could help you out in case you end up in court, but that's a maybe, not a guarantee.

Harder to make this argument with encrypted distributed filesystems. If I'm storing a single chunk of an encrypted blob on Filecoin, am I responsible for the entire file even if I don't know what's in it, and I'm only storing a single fragment?

Re: The Internet Archive is back online

#66
post #51
post #25

Earlier quoted context omitted.

Incentivising seeding is hard. Maybe cryptocurrencies can be useful here, but I understand not everyone likes them especially here on HN. In retrospect the ideal setup would have been if archiving was included into the core HTTP protocol.

Maybe someone can invent a proof of seeding protocol? So that would bring some good to the public instead of just burning energy. Don't ask me how it would work...

Storj, Filecoin, etc. fill this gap but it's still really hard to earn enough to justify the effort at small scales.

Re: The Internet Archive is back online

#67
post #28

Earlier quoted context omitted.

I would assume they delete illegal stuff as they are compelled to. What I'd like to know is their policy for legal stuff that they exclude that is not as a result of DMCA.

Can you give an example? EDIT: Just seen your other reply. Perhaps it was excluded due to right to forget laws?

ploetzblog was available and is now completly gone :( "lost" some recipes that he didn't migrate that I used to bake all the time. Used to look it up on the IA and was pissed when it was deleted

Re: The Internet Archive is back online

#68
Does someone know where i can download MIT OCW videos ?

As the videos are present on archive.org but it is down and i was unable to find them anywhere else online ?

Also, yt-dlp is also not working: https://github.com/yt-dlp/yt-dlp/issues/10128

Example: https://ocw.mit.edu/courses/7-016-introductory-biology-fall-...

Re: The Internet Archive is back online

#69

Earlier quoted context omitted.

Kiwi farms

There are people that maintain "non-public archives" of stuff like that for litigation, long-term archival storage (think sealed boxes intended for future generations of historians. (Libraries, laywers, journalists can run their own WebRecorder, Perma.cc, ArchiveBox, etc. instances) I think that's a reasonable middle ground, we don't necessarily need every single piece of heinous content mirrored for free access 24/7…

[deleted]

Re: The Internet Archive is back online

#70
post #13

When the internet archive censors a website is it deleted permanently or just not publicly available?

I don't think they censor anything, strictly archiving. Do you know of any instance in which they censored a site?

I take it you’ve never encountered the dreaded message, “The item is not available due to issues with the item's content”?

There was a news item here on HN about something available on the Internet Archive: https://news.ycombinator.com/item?id=16725526> This is now gone from IA. Old page with links to IA which are no longer working: https://web.archive.org/web/20180331224513/http://profileeng...>

Post reply on HN