Live data from Hacker News

The Internet Archive is under a DDoS attack

mastodon.archive.org

191–200 of 227 posts

Re: The Internet Archive is under a DDoS attack

#191
post #137

Earlier quoted context omitted.

would a decentralized internet archive make sense? impossible to ddos.

That exists and it's called Arweave. You even collect imaginary brownie points for maintaining the archive.

> Arweave

looks a bit more broad than i wanted. as is its like a thin wrapper for IPFS.

Re: The Internet Archive is under a DDoS attack

#192
post #191

Earlier quoted context omitted.

That exists and it's called Arweave. You even collect imaginary brownie points for maintaining the archive.

> Arweave looks a bit more broad than i wanted. as is its like a thin wrapper for IPFS.

It's a fake brownie point scheme (i.e. cryptocurrency) to let the people who archive the most stuff decide what goes into the archive. There's also IPFS but that has absolutely no way to decide what gets archived.

Re: The Internet Archive is under a DDoS attack

#194
post #51

Earlier quoted context omitted.

Botnets usually, sometimes amplification attacks against NTP or DNS, although the Chinese government’s Great Firewall also has offensive capabilities known as the Great Cannon, although they are generally used against GitHub because it hosts censorship-circumvention software like VPNs.

Are botnets usually hosted on personal computers or servers or IoT? I'm thinking maybe archive.org can block a whole range of IPs if needed.

All of the above, and also increasingly IoT devices, specially printers and routers. There is no simple way to block them.

Re: The Internet Archive is under a DDoS attack

#196
post #86

Earlier quoted context omitted.

Doubt cloudflare has anything to do with it. The operators most likely don't want to openly expose their website's ip addresses.

That is exactly the problem. These services are constantly at war with each other and are attacked by competitors. Cloudflare provides DDoS protection to the DDoS providers so they can keep their services online, which directly benefits Cloudflare by DDoS being a bigger problem than if they were all busy attacking each other. This is a sampling of currently available services and who they use for DDoS protection: str…

[dead]

Re: The Internet Archive is under a DDoS attack

#197

Earlier quoted context omitted.

Assume everyone is familiar with this project, dating back to 1996: https://en.wikipedia.org/wiki/WWWOFFLE https://ftp.netbsd.org/pub/pkgsrc/distfiles/wwwoffle-2.9j.tg... The way the www is going, it seems like downloading a copy of libgen, i.e., nonfiction books, and scimag, i.e., academic journals, via torrent, would be more valuable than archiving websites, in general. These primary sources are part of the materia…

Do we know for sure that they trained on data from libgen etc? It's such a powerful source of information you'd assume they must have, although they would never admit it. There must be a way to test if they have, via enquiring about some niche information only found in certain books.

It is apparently widely suspected that a certain "Books2" dataset mentioned by OpenAI is basically just LibGen:

https://blusharkmedia.medium.com/the-ongoing-battle-against-...

https://techhq.com/2023/09/can-libgen-shadow-library-survive...

https://www.twitter.com/theshawwn/status/1320282152689336320

https://qz.com/openai-books-piracy-microsoft-meta-google-cha...

https://qz.com/shadow-libraries-are-at-the-heart-of-the-moun...

https://goodereader.com/blog/e-book-news/authors-file-lawsui...

When asked about whether this was true, they refused to answer based on confidentiality concerns, then said they had deleted all copies of the dataset, stopped using it, and no longer employed the individuals that compiled it:

https://www.businessinsider.com/openai-destroyed-ai-training...

We do know for a fact that the (non-OpenAI-controlled) "Books3" dataset is just "all of bibliotik":

https://www.twitter.com/theshawwn/status/1320282149329784833

https://github.com/soskek/bookcorpus/issues/27

And we also apparently know for a fact that this was included in the datasets used to train LLAMA:

https://en.wikipedia.org/wiki/The_Pile_(dataset)

https://aicopyright.substack.com/p/the-books-used-to-train-l...

https://aicopyright.substack.com/p/has-your-book-been-used-t...

Re: The Internet Archive is under a DDoS attack

#198

Earlier quoted context omitted.

Same. I've been scraping PDF'ed magazines, etc. and keeping them locally. In addition to feeding my byte-hoarding tendencies, I like the idea I could be off-grid in my van/RV somewhere and reading a "Popular Electronics" magazine from 1972 on my laptop. (Oh, never mind YouTube videos that I once added to playlists ... that later disappear leaving only holes in my playlists.)

My problem with this approach is that the stuff I want to look at in 10 yrs time is never the stuff I think of saving right now. In the 2000s there were browser extensions I've forgotten the names of (shelf? slogger?) that would automatically save local copies of every webpage on page load. But I don't think they're around anymore and have no idea how you could achieve similar functionality with dynamic pages anyway.

> But I don't think they're around anymore and have no idea how you could achieve similar functionality with dynamic pages anyway.

Chromium's MHTML "Save as…" and the SingleFile WebExtension should both save copies of the rendered DOM.

Apparently Safari has WebArchive and Mozilla had MAFF for similar use cases.

I think WARC is supposed to save enough data about network streams for dynamic pages to work. At least on the Wayback Machine, infinite scrolling and "Load More" buttons do kinda work sometimes. You may have to load the archived pages in a browser and try to use each dynamic feature at least once, to trigger requests for needed resources.

SingleFile: https://github.com/gildas-lormeau/SingleFile

LWN on WARC, tools: https://anarc.at/blog/2018-10-04-archiving-web-sites/

Self-hostable web archives: https://awesome-selfhosted.net/tags/archiving-and-digital-pr...

Wayback Machine addons, bookmarklets: https://help.archive.org/help/save-pages-in-the-wayback-mach...

Re: The Internet Archive is under a DDoS attack

#199
post #173

Earlier quoted context omitted.

Did you read the blog post? It doesn't include the entire correspondence so it's not clear how explicit Cloudflare was about this but the Enterprise plan they were trying to upsell them includes BYOIP. It's clear to me that Cloudflare insisted they buy the enterprise plan because it includes BYOIP. So in other words, Cloudflare noticed the author was running a gambling site, they decided that this was negatively impa…

To me it’s similar to the whole “SSO wall of shame” thing, where a vital feature is locked behind more expensive pricing. As said in the article: “We tried saying that we don't need any number of the 14 features that are included” Which, to me, is the crux of the issue. Is it fair for Cloudflare to say “You are breaking the terms of service if you do not change your set up in this specific way, and also the way you n…

I'd say at that point it's essentially compensation for personal suffering.

Yes, BYOIP as a feature does not seem complex enough to warrant paying for an Entperise license. But the kind of customers who need BYOIP (especially if they need it to avoid harming your IP reputation) are likely to be at a higher risk of being flaky or otherwise painful so this is very much a tax on running that kind of business (just as porn sites often find it hard to find payment processors because of the high risk of credit card fraud).

As a freelancer I have absolutely made offers at 10x my going rate for client I did not want. The idea is that if they really want me to work for them, at least I get reimbursed for the suffering that will entail. This kind of pricing structure is no different.

Re: The Internet Archive is under a DDoS attack

#200
post #184

Earlier quoted context omitted.

Shallow dismissal anyway, even if he was the Supreme Majestic King of New Americania. He might further explain his answer. And I'm truly sorry for the DDos happening to this guy's organisation!

What is there to explain further?

Some evidence or reasoning that there's nothing damning anywhere on IA that anyone could possibly need kept quiet to the point of having ordered this particular attack. Just being the target of an attack doesn't mean you have perfect information about the perpetrator or their motives.

This could even be as simple as "Some aspect of the attack pattern is inconsistent with such a motive", or "We spotted the perpetrator credibly gloating about it". But just from IA's public statements, the pattern ("launching tens of thousands of fake information requests per second") is quite consistent with simple denial.

Post reply on HN