Live data from Hacker News

Internet Archive breached again through stolen access tokens

bleepingcomputer.com

351–360 of 376 posts

Re: Internet Archive breached again through stolen access tokens

#351
post #275

Earlier quoted context omitted.

LigGen / SciHub are model citizens in this regard. Use the best tech for the job . Torrents + IPFS + simple http mirrors all over . While their data is big its not as big as Archive.org https://libgen.is/repository_torrent/ so I guess one still needs the funding to go to these decentralized nodes and they are there mostly for resiliancy

If so, why do they keep jumping from one dubious domain and TLD from another?

That's mentioned in the comment you're replying to as "simple HTTP mirrors." Those are vulnerable to takedowns but are kept up for the sake of easy accessibility to the public that doesn't know how to access any of the other options. Their availability has no real bearing on the availability of the other options beyond the tech savviness of the general public.

Re: Internet Archive breached again through stolen access tokens

#352

Earlier quoted context omitted.

Internet archive allows full text search of books, newspapers, etc.. Or anyway it did, before being breached.

It does transcribe books (through imperfect OCR) so I guess that's possible. Never relied on it as I search by title and author. But anyways not the case for the wayback product which is the unique core to IA.

That's not unique, not a product, and not the part I use most.

Well, OK, maybe other webpage archives don't work as well, I haven't tried them, but there are others. And they're newer, so don't have such extensive historical pages.

Large numbers of Wikipedia references (which relied on IA to prevent link rot) must be completely broken now.

Re: Internet Archive breached again through stolen access tokens

#353

Earlier quoted context omitted.

A million is out of the parameters of the case. Realistically you won't get enough volunteer-storage to cover one IA. And even if you did, it wouldn't satisfy the mission requirements, which is to store reliably for decades all of the data.

This isn't meant to be storage for IA, it's meant to be a distributed backup.

Ah my bad, so it's not a replacement of IA. In that case it makes sense

Re: Internet Archive breached again through stolen access tokens

#354

Earlier quoted context omitted.

Ooo, excellent. Yes, hiding items is imperfect, but I understood that it was legally required or something. (IANAL and IDFK, TBH) I wonder how perma.cc gets around that.

I'm afraid that it just hasn't been tested in court yet. I haven't read this paper yet, but... https://www.tesble.com/10.1080/0270319x.2021.1886785 from the abstract: > The article concludes that Perma.cc's archival use is neither firmly grounded in existing fair use nor library exemptions; that Perma.cc, its "registrar" library, institutional affiliates, and its contributors have some (at least theoretical) exposure…

So, precisely the same constraints that IA operates under, just perma.cc isn't big enough yet to have been forced to comply with them?

I'll hold my breath.

Re: Internet Archive breached again through stolen access tokens

#355
post #21

Earlier quoted context omitted.

This seems to get brought at least once in the comments for every one of these articles that pops up. The IA has tried distributing their stores, but nowhere near enough people actually put their storage where their mouths are.

Perhaps a naïve question, but hasn't this problem been solved by the FreeNet Project (now HyphaNet) [0]? (the re-write — current FreeNet — was previously called Locutus, IIRC [1]). Side note: As an outsider, and someone who hasn't tried either version of FreeNet in more than almost 2 decades, was this kind of a schism like the Python 2 vs. Python 3 kerfuffle? Is there more to it? [0]: https://www.hyphanet.org/ [1]: h…

Hi, Freenet's FAQ explains the renaming/rebranding here: [1]

Neither version of Freenet is designed for long-term archiving of large amounts of data so it probably isn't ideally suited to replacing archive.org, but we are planning to build decentralized alternatives to services like wikipedia on top of Freenet.

[1] https://freenet.org/faq/#why-was-freenet-rearchitected-and-r...

Re: Internet Archive breached again through stolen access tokens

#356

We need archives built on decentralized storage. Don't get me wrong, I really like and support the work Internet Archive is doing, but preserving history is too important to entrust it solely to singular entities, which means singular points of failure.

This sounds like a task that could be taken care of by public libraries. That, and supporting Tor. I just don't see it happening anytime soon, at least in the US.

Re: Internet Archive breached again through stolen access tokens

#357
To everyone who wants a better alternative to IA, who thinks they have a different solution, who thinks it should be run by a different organization, etc.

Nobody has ever stopped a competitive alternative from existing. Feel free to give it a shot. You have a head start with all the work that they've done and shared.

Re: Internet Archive breached again through stolen access tokens

#358

Earlier quoted context omitted.

This isn't meant to be storage for IA, it's meant to be a distributed backup.

Ah my bad, so it's not a replacement of IA. In that case it makes sense

Yes, the idea is that this is a replacement for the torrents they make public. In case the IA goes away, we'll have this distributed dataset to fall back on.

Re: Internet Archive breached again through stolen access tokens

#359
post #305
post #296

Earlier quoted context omitted.

I'm not going to look up legal precedent, hire a lawyer if you want that. You are wrong, copyright specifically prohibits copying, not distribution. They can get a cease and desist that requests you destroy property and they ca get a court order backing that which will put you into contempt of court if you fail to do so. Proving damages is easier with distribution, but that is a civil matter not a criminal matter.

They’re not making copies either.

So they aren't making copies? How then do they have an archive of internet resources if not by copying said resources?

You do realise the "downloading" is implicitly a copy.

If you want to actually have a civil discussion then you need to make some reasonable argument than "They're not making copies either."

Sounds like whatever role you played at IA when you were there didn't give you any actual insight into what happens in operation and you simply tried to prove your point with an appeal to authority instead of backing it with facts and reason.

Re: Internet Archive breached again through stolen access tokens

#360

Earlier quoted context omitted.

Ah my bad, so it's not a replacement of IA. In that case it makes sense

Yes, the idea is that this is a replacement for the torrents they make public. In case the IA goes away, we'll have this distributed dataset to fall back on.

An archive of an archive
Post reply on HN