Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

181–190 of 207 posts

Re: Thank you for helping us increase our bandwidth

#181
post #62

Earlier quoted context omitted.

> . I think that a crisis like this is exactly the time to try new and experimental ways of making things available. Yeah a crisis when the stress level of anyone is already much higher is the best time to crank the stress of authors even higher by experimenting with their livelihood without consulting them!

There is a huge difference between a book and code that I write. A book can be read once and the reader benefits. Useful software is usually useful for far more than one use. Free trials and demos are how we sell more software. A book? It's hard to have a free trial. That said, I really have a hard time assailing libraries. Ebooks are sold to libraries at 2-5x the price I can buy the same ebook for as a consumer. Pri…

> A book? It's hard to have a free trial.

Publish one chapter of the book in an anthology, on the web, etc. I've purchased several books after reading one chapter in this manner.

Re: Thank you for helping us increase our bandwidth

#183
post #100
post #88

Earlier quoted context omitted.

You can get an 8TB HDD for $150 right now. That's 6250 drives. That's about $1MM in drives, which doesn't sound that cost-prohibitive. Obviously that's not the whole cost since you need to pay for bandwidth, replication and other infrastructure like the host node, but it sounds like something that could be even hosted by a number of volunteers on r/datahoarder or r/homelab. I also remember reading about Sia on HN, wh…

BackBlaze B2 is $5/TB/month Azure Archive is $2/TB/month ($1.68 if reserved) AWS Glacier Deep Archive is $1/TB/month GCP Cloud Storage Archive is $1.20/TB/month Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.

The majority of cloud costs of storage are bandwidth, so ignoring this makes the analysis meaningless.

Re: Thank you for helping us increase our bandwidth

#184

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

There are only two sites I donate to: wikipedia and archive.org. I honestly don't know what I'd do without them. I have been able to find web sites from 20 years ago on archive.org, it's an absolute treasure.

Re: Thank you for helping us increase our bandwidth

#185

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

They could set up a torrent that people could download parts of, I think it's an ideal system for a distributed backup. But it probably needs tweaking for that purpose; for one you need to ensure that the data is evenly distributed, and second you're dealing with data that is appended to regularly. But I think it can be done.

It exists! The ArchiveTeam had a project, you can see its status here: http://iabak.archiveteam.org/

Here's an overview: https://www.archiveteam.org/index.php?title=INTERNETARCHIVE....

Sadly, it seems the project has fallen out of repair.

Re: Thank you for helping us increase our bandwidth

#186

Earlier quoted context omitted.

> Also, I'd trust these guys not to mishandle .org As much as I like the idea of running a TLD, if anyone gives me a TLD, I'm gonna put an MX record on it and be jonah@org. That's a promise. EDIT: To be fair, I would probably be overwhelmed with the usual sense of responsibility and not do this. But, the temptation

Would be a nightmare with all the crappy email validators out there..

I recall a similar story involving amateur radio. By ITU convention, amateur callsigns are [prefix][digit][suffix], like W1AW.

So one night, a new ham is operating his friend's fancier station, which has the HF equipment to talk clear around the world. And they make contact with another ham identifying himself as JY1. No suffix. Well that's weird, so they look it up, the JY prefix is Jordan, okay. So the guy on the mic just asks, "why doesn't your callsign have any letters after the digit?" and the reply comes back with a chuckle, "Oh, because I am the King."

That was the late King Hussein. about whom much has been written. I don't know if his call ever caused trouble with software validators, but I'd certainly believe it.

Re: Thank you for helping us increase our bandwidth

#187
post #24
post #9

The wayback machine has been essential through COVID as many government just publish the numbers and data "of the day" and the only way to compare to the day before is to look at IA.

Do you have any examples of government sites that are doing this? I have a side-hobby of setting up scrapers which pull scraped data into a git repository, precisely for this kind of thing. I'd be happy to set a few up. Some of my posts about this technique (which I call "git scraping"): https://simonwillison.net/tags/gitscraping/

Government of Ontario has been doing it. They now have some APIs that link to the number of cases per day, but for the first month of the pandemic, they were literally overwriting the page every day and when you look at github there was tons of researchers that came up with their own parser in order to save this data before it is deleted.

I've also been running bash scripts every day to save an archive of some of those pages https://github.com/jeromegv/covid_data

Re: Thank you for helping us increase our bandwidth

#188
post #169

I wish that Archive.org would start collecting funds to finance an external backup site. I strongly feel that archive.org is an irreplaceable treasure and that making sure that it would resist natural disasters should be a priority...

That doesn't require a special fund, donate to them. They'd make backups if they had enough donations to cover it.

Re: Thank you for helping us increase our bandwidth

#189
post #101

Earlier quoted context omitted.

They have lots of copyrighted content. Such as virtually every commercial retro video game ever made it seems. I personally think it’s great but surely companies aren’t too pleased? How does archive.org avoid being sued into oblivion?

There are exemptions to DMCA for certain content. > Computer programs and video games distributed in formats that have become obsolete and which require the original media or hardware as a condition of access. A format shall be considered obsolete if the machine or system necessary to render perceptible a work stored in that format is no longer manufactured or is no longer reasonably available in the commercial marke…

That doesn't seem to stop Nintendo from shutting down ROM sites left and right. For example, emuparadise has complied and taken down all ROMs: https://www.emuparadise.me/emuparadise-changing.php

But here is a complete NES ROM set, just sitting there, on archive.org: https://archive.org/details/NESrompack

I just find it strange. Nintendo is vehement about their IP. For archive.org to get an exception doesn't make sense to me.

Re: Thank you for helping us increase our bandwidth

#190
post #101

Earlier quoted context omitted.

They have lots of copyrighted content. Such as virtually every commercial retro video game ever made it seems. I personally think it’s great but surely companies aren’t too pleased? How does archive.org avoid being sued into oblivion?

It's a non-profit library and they don't publicly provide a way to download most of the copyrighted content, so rights holders aren't too concerned. I imagine most people, even copyright lawyers, personally support the free archiving of everything made in the current age as long as it doesn't detract from current business operations and the copyrights are respected for their duration.

It's very easy to download the copyrighted content. For example, here is a full NES ROM set: https://archive.org/details/NESrompack
Post reply on HN