Live data from Hacker News

Internet Archive breached again through stolen access tokens

bleepingcomputer.com

271–280 of 376 posts

Re: Internet Archive breached again through stolen access tokens

#271

Earlier quoted context omitted.

Their torrents suck and IME don’t update to changes in the archive.

Torrents are immutable in principle, which is good for preserving things. A new version of a set of files should be a new torrent.

How would preservationists go about automatically updating the torrent and data they seed? Or would they need to manually regularly check, if they are still seeding the up-to-date content?

Re: Internet Archive breached again through stolen access tokens

#272
post #158

Earlier quoted context omitted.

The group that claimed to be responsible for the first hack was said to be Russian-based, anti-U.S., pro-Palestine, and their reasoning for the attack was because of IA's violation of copyright.... I think you should draw your own more informed conclusions, but it smells a lot like feds to me.

What do Palestine, Russia, and the U.S. have to do with the Internet Archive? The Internet Archive is a supremely boring target politically.

> The Internet Archive is a supremely boring target politically.

I mean it's where we go to prove a politician said something after they deleted it or where a government changes the wording of something...

I'd argue it's one of the juicer political targets if you're actually wanting to do something.

Re: Internet Archive breached again through stolen access tokens

#273

We need archives built on decentralized storage. Don't get me wrong, I really like and support the work Internet Archive is doing, but preserving history is too important to entrust it solely to singular entities, which means singular points of failure.

I designed a system where you could say "donate this spare 2 TB of my disk space to the Internet Archive" and the IA would push 2 TB of data to you. This system also has the property that it can be reconstructed if the IA (or whatever provider) goes away. Unfortunately, when I talked to a few archival teams (including the IA) about whether they'd be interested in using it, I either got no response or a negative one.

Because the incentive will be archiving things they believe should be archived, you need the process to begin with what urls do you want to be archiving, then people will be incentivized for archiving the juicy stuff IA is used for and you just throw some stuff they didn't ask to archive in the remit of them storing what they want.

Re: Internet Archive breached again through stolen access tokens

#275

We need archives built on decentralized storage. Don't get me wrong, I really like and support the work Internet Archive is doing, but preserving history is too important to entrust it solely to singular entities, which means singular points of failure.

LigGen / SciHub are model citizens in this regard. Use the best tech for the job . Torrents + IPFS + simple http mirrors all over . While their data is big its not as big as Archive.org https://libgen.is/repository_torrent/ so I guess one still needs the funding to go to these decentralized nodes and they are there mostly for resiliancy

Re: Internet Archive breached again through stolen access tokens

#276

Earlier quoted context omitted.

I designed a system where you could say "donate this spare 2 TB of my disk space to the Internet Archive" and the IA would push 2 TB of data to you. This system also has the property that it can be reconstructed if the IA (or whatever provider) goes away. Unfortunately, when I talked to a few archival teams (including the IA) about whether they'd be interested in using it, I either got no response or a negative one.

Why reinvent the wheel ? There are so many proven distributed archiving systems, a lot of which are mentioned in these comments.

What system of these allows me to donate a bunch of disk space to a provider of my choosing, without thinking about it afterwards?

Re: Internet Archive breached again through stolen access tokens

#277
post #158

Earlier quoted context omitted.

The group that claimed to be responsible for the first hack was said to be Russian-based, anti-U.S., pro-Palestine, and their reasoning for the attack was because of IA's violation of copyright.... I think you should draw your own more informed conclusions, but it smells a lot like feds to me.

What do Palestine, Russia, and the U.S. have to do with the Internet Archive? The Internet Archive is a supremely boring target politically.

>supremely boring target politically.

Oh how wrong you are.

There is nothing boring in a target that can be used to validate others' lies and potential hypocrisy, changes in their policies etc.

The wayback machine itself serves as a truly priceless way to go back through someone's public life.

Re: Internet Archive breached again through stolen access tokens

#278

Earlier quoted context omitted.

Nearly every entry in the library has a torrent file (which is a distributed storage system), but with the index pages down, they're not accessible.

They're not using DHT?

They're not talking about peer discovery, they're talking about .torrent file discovery.

Re: Internet Archive breached again through stolen access tokens

#279

We need archives built on decentralized storage. Don't get me wrong, I really like and support the work Internet Archive is doing, but preserving history is too important to entrust it solely to singular entities, which means singular points of failure.

I designed a system where you could say "donate this spare 2 TB of my disk space to the Internet Archive" and the IA would push 2 TB of data to you. This system also has the property that it can be reconstructed if the IA (or whatever provider) goes away. Unfortunately, when I talked to a few archival teams (including the IA) about whether they'd be interested in using it, I either got no response or a negative one.

Is this open source or do you have any design docs? I love the idea and would love to learn more about it.

Re: Internet Archive breached again through stolen access tokens

#280

Earlier quoted context omitted.

Our digital memory shouldn't be in the hands of a small number of organizations in my view. You're right about cost effectiveness. There are pros and cons to both but it's not just external threats that have to be considered. History has always gotten rewritten throughout time. If you have a giant library it's easier for bad actors to gain influence and alter certain books, or remove them. This isn't just theoretical…

>Our digital memory shouldn't be in the hands of a small number of organizations in my view. I would wager at least 95% of "digital memory" archived is just absolute garbage from SEO spam to just some small websites holding no actual value. The true digital memory of the world is almost entirely behind the walls of reddit, twitter, facebook, and very few other sites. The internet landscape has changed massively from…

We are currently in the middle of an information dark age and not many people have realised this yet.
Post reply on HN