Live data from Hacker News

The Internet Archive is under a DDoS attack

mastodon.archive.org

201–210 of 227 posts

Re: The Internet Archive is under a DDoS attack

#201
post #105

This is why I’ve gotten into the habit of maintaining my own WWW archive of sites I find interesting. Probably have around 1 TiB now, and One Of These Days I’d like to set my network up so it can serve arbitrary sites directly from local archive to revive any site I want. I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this co…

Assume everyone is familiar with this project, dating back to 1996: https://en.wikipedia.org/wiki/WWWOFFLE https://ftp.netbsd.org/pub/pkgsrc/distfiles/wwwoffle-2.9j.tg... The way the www is going, it seems like downloading a copy of libgen, i.e., nonfiction books, and scimag, i.e., academic journals, via torrent, would be more valuable than archiving websites, in general. These primary sources are part of the materia…

https://news.ycombinator.com/item?id=34616256

https://news.ycombinator.com/item?id=40258584

https://arxiv.org/pdf/2005.14165.pdf

https://www.wired.com/story/battle-over-books3

https://www.washingtonpost.com/technology/interactive/2023/a...

https://www.theguardian.com/technology/2023/apr/20/fresh-con...

https://storage.courtlistener.com/recap/gov.uscourts.cand.41...

See 40-45.

https://storage.courtlistener.com/recap/gov.uscourts.nysd.60...

See 87-116.

Re: The Internet Archive is under a DDoS attack

#202
post #135
post #105

This is why I’ve gotten into the habit of maintaining my own WWW archive of sites I find interesting. Probably have around 1 TiB now, and One Of These Days I’d like to set my network up so it can serve arbitrary sites directly from local archive to revive any site I want. I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this co…

yes please. extra credit for anyone who shares instructions on how to inject this into every website i browse sans blocklist

https://archivebox.io/ perhaps?

Re: The Internet Archive is under a DDoS attack

#205
post #105

This is why I’ve gotten into the habit of maintaining my own WWW archive of sites I find interesting. Probably have around 1 TiB now, and One Of These Days I’d like to set my network up so it can serve arbitrary sites directly from local archive to revive any site I want. I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this co…

You can also use archivebox very user friendly

Re: The Internet Archive is under a DDoS attack

#206
post #105

This is why I’ve gotten into the habit of maintaining my own WWW archive of sites I find interesting. Probably have around 1 TiB now, and One Of These Days I’d like to set my network up so it can serve arbitrary sites directly from local archive to revive any site I want. I have a `wget-mirror` shell function invoking wget with all the trimmings that takes care of 99% of sites. I’ll edit the full command into this co…

Assume everyone is familiar with this project, dating back to 1996: https://en.wikipedia.org/wiki/WWWOFFLE https://ftp.netbsd.org/pub/pkgsrc/distfiles/wwwoffle-2.9j.tg... The way the www is going, it seems like downloading a copy of libgen, i.e., nonfiction books, and scimag, i.e., academic journals, via torrent, would be more valuable than archiving websites, in general. These primary sources are part of the materia…

This "AI" nonsense seems like a legitimate threat to literacy. Why would young people read a nonfiction book when they can just send questions to a so-called "tech" company that has used the book in training a LLM. These companies, needless intermediaries with zero experise on the subject matter of the book, exist only to collect data and use it for commercial purposes. Unlike the books' authors and publishers they have no responsibility for publishing information that is adequately researched and factually correct.

Re: The Internet Archive is under a DDoS attack

#207
post #153
post #149

Earlier quoted context omitted.

how is this a solution? The Archive performs a valuable service. They're collecting wahy more of the internet than you are (I assume) so when that thing you didn't back up today is not available in 10yrs it's more likely to be on the archive. I donate to The Archive. More people should too.

I don't know why you're treating them as mutually exclusive. Single points of failure are as bad when it comes to organizations as they are with anything else. Internet Archive (the org) could stop existing with the flick of a pen. I don't think “Let Somebody Else Do It” is a healthy attitude to take, and I'm going to keep doing what I'm doing. Plus for as great of a service as Wayback Machine is, it can be very unpl…

Heck IA will even temporarily IP-block you just for loading an archived page with too many images. It's a very useful resource but often also very painful to use.

Re: The Internet Archive is under a DDoS attack

#208
post #153

Earlier quoted context omitted.

I don't know why you're treating them as mutually exclusive. Single points of failure are as bad when it comes to organizations as they are with anything else. Internet Archive (the org) could stop existing with the flick of a pen. I don't think “Let Somebody Else Do It” is a healthy attitude to take, and I'm going to keep doing what I'm doing. Plus for as great of a service as Wayback Machine is, it can be very unpl…

Heck IA will even temporarily IP-block you just for loading an archived page with too many images. It's a very useful resource but often also very painful to use.

Even if you're logged in?

Re: The Internet Archive is under a DDoS attack

#209

Earlier quoted context omitted.

Cloudflare is not "an extortionist firm". It is a large tech company, where occasionally teams employ shitty sales tactics to meet their numbers, but generally provides a valuable service and acts reasonably ethically. There are open source tools to mitigate DDoS, but all of them will have some marginal cost to run, and they will all be significantly worse than Cloudflare as they benefit from neither Cloudflare's dat…

No, thanks. Cloudflare acts ethically only until it suits them. It is the pre-exploitation phase to lure a customer. We are not fools here. The report at https://news.ycombinator.com/item?id=40481808 says it all. Secondly, considering Cloudflare would MITM all traffic, it would make a data good source for the NSA, thereby violating all user privacy.

That argument doesn't really fly. The "poor little customer" was an online casino who was using Cloudflare to avoid getting taken down in countries where online gambling isn't allowed.

This had a high risk of getting Cloudflare's limited ipv4 addresses to a blacklist - affecting ALL of their customers.

All CF did was ask them to switch to an Enterprise plan and bring their own IP-addresses. They refused to do either and rather cried on the internet claiming CF to be bullies. It's not like the price they asked was even a fraction of the profits an online casino brings in every DAY.

Re: The Internet Archive is under a DDoS attack

#210

Earlier quoted context omitted.

What's the significance of that? (Googling "Jason Scott TIA" gives me "Dr Jason Scott is a Senior Research Fellow in the Tasmanian Institute of Agriculture" which doesn't explain much to me)

Jason Scott works at the Internet Archive[1]. [1]: https://en.wikipedia.org/wiki/Jason_Scott

And he knows about every single file in it?
Post reply on HN