Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

91–100 of 207 posts

Re: Thank you for helping us increase our bandwidth

#91
post #21

Earlier quoted context omitted.

Here's another one: "some of our old tweets were wrong so we just quietly deleted them lol": https://twitter.com/voxdotcom/status/1242537366620966912

This is necessary on Twitter. It's easy for fake or incorrect news to spread, you can't edit it and updates/self-replies are often hidden.

Sure, but they should also own up to their mistake. They flat out said "It won't be a deadly epidemic".

Re: Thank you for helping us increase our bandwidth

#92

Earlier quoted context omitted.

This is necessary on Twitter. It's easy for fake or incorrect news to spread, you can't edit it and updates/self-replies are often hidden.

Sure, but they should also own up to their mistake. They flat out said "It won't be a deadly epidemic".

They made a tweet to explicitly say they deleted it because it's wrong. How is that not "owning up"?

Re: Thank you for helping us increase our bandwidth

#93
Kinda off-topic, but I noticed that it seems the owners of the sites can ask Archive.org to remove existing archives. I've lost a few I "created" myself this way.

Not say it's not reasonable, but how do we work around this? Any alternative services that are more "resilient" in this regard?

Re: Thank you for helping us increase our bandwidth

#94
Maybe my sense of scale is warped but I’m amazed at how little BW the archive uses. Since this data appears to be averaged over a 1 minute window, I wonder what their p99 usage is. The article says they bought a 2nd router. Were they only using a single router before (no redundancy or protected BW)? I realize the Archive probably has a small budget but one can buy a 3.2T switch relatively cheaply these days.

Re: Thank you for helping us increase our bandwidth

#95
post #17

60 Gbit/s of continuous traffic is a lot. If I'm reading the graphs right, Wikipedia "only" has 13.4 Gbit/s of outbound traffic [1]. Of course still well below the single-digit Tbit/s traffic of large internet exchanges [2] [3], but still unexpectedly large. [1]: adding up the outbound numbers for each datacenter: 1.888 + 8.003 + 810 + 1.958 + 807: https://grafana.wikimedia.org/d/000000605/datacenter-global-... [2]:…

Wikipedia is mostly text with some images. Internet archive offers every hung from websites and text to movies and games, which obviously are much larger and don't compress as nicely either.

Re: Thank you for helping us increase our bandwidth

#96
post #66
post #59

What's the easiest way to seed the most popular content ala torrent?

ipfs: https://betanews.com/2018/08/09/decentralized-archive-org

I would love to donate bandwidth/storage. I haven't a clue how. I wish there was software I could throw in my server, set how much storage and bandwidth I can donate and run 24/7.

Re: Thank you for helping us increase our bandwidth

#97
post #64
post #61

Earlier quoted context omitted.

I'm very worried about its backups and what happens when the next Big One hits SF. As far as I've ever been able to determine from talking to anyone at IA (e.g., Kahle, Scott), they don't really have any sort of backups that could actually be restored from in a disaster situation.

There is http://iabak.archiveteam.org , but it’s not exactly large.

It definitely stalled out. It needs a Windows version to get real traction IMO.

Re: Thank you for helping us increase our bandwidth

#98
post #64

Earlier quoted context omitted.

There is http://iabak.archiveteam.org , but it’s not exactly large.

If I'm reading that correctly, it would only cost a bit over 500 bucks a month to host that whole archive on BackBlaze B2. Furthermore it would not be so hard to translate Archive.org items to IPFS objects, if there were an effort to pin a significant number of them to storage and network.

The effort is the issue here. There was this comment back when IA.BAK was in design phase https://news.ycombinator.com/item?id=9148576 And then this is all there was to show for it: https://www.archiveteam.org/index.php?title=INTERNETARCHIVE....

Re: Thank you for helping us increase our bandwidth

#99

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

They should team up with Backblaze...

Re: Thank you for helping us increase our bandwidth

#100
post #88
post #71

Earlier quoted context omitted.

The IA is about 50 PB. IABAK stores 100 TB, or 0.2% of it.

You can get an 8TB HDD for $150 right now. That's 6250 drives. That's about $1MM in drives, which doesn't sound that cost-prohibitive. Obviously that's not the whole cost since you need to pay for bandwidth, replication and other infrastructure like the host node, but it sounds like something that could be even hosted by a number of volunteers on r/datahoarder or r/homelab. I also remember reading about Sia on HN, wh…

BackBlaze B2 is $5/TB/month

Azure Archive is $2/TB/month ($1.68 if reserved)

AWS Glacier Deep Archive is $1/TB/month

GCP Cloud Storage Archive is $1.20/TB/month

Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.

Post reply on HN