Live data from Hacker News

Thank you for helping us increase our bandwidth

blog.archive.org

111–120 of 207 posts

Re: Thank you for helping us increase our bandwidth

#111
post #101

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

They have lots of copyrighted content. Such as virtually every commercial retro video game ever made it seems. I personally think it’s great but surely companies aren’t too pleased? How does archive.org avoid being sued into oblivion?

It's a non-profit library and they don't publicly provide a way to download most of the copyrighted content, so rights holders aren't too concerned. I imagine most people, even copyright lawyers, personally support the free archiving of everything made in the current age as long as it doesn't detract from current business operations and the copyrights are respected for their duration.

Re: Thank you for helping us increase our bandwidth

#112
post #101

archive.org feels like an irreplaceable treasure, the Wayback Machine alone is a time capsule of our digital history. I donate to them monthly and know a lot of other people do as well, so I don't worry much about their financial stability. I'm more worried about external pressures taking content down. I hope the data is backed up six ways to sunday, and that somewhere there's a plan to make it all accessible if Inte…

They have lots of copyrighted content. Such as virtually every commercial retro video game ever made it seems. I personally think it’s great but surely companies aren’t too pleased? How does archive.org avoid being sued into oblivion?

As far as I know, they limit public access to basically any material with a complaint or request, but keep copies. They may do more, but that seems to be the default response.

Which seems smart enough: minimize litigation costs while not losing content permanently via court orders or dmcas or similar threats.

But it's frustrating for the wayback machine- where sites like Snopes and some newspapers have opted out of having their history published (after being accused of ghost edits to articles).

I don't know what the right answer is, but even if they didn't display the page, it would be nice to see the diffs (ala wikipedia), or at least show if and when changes to a page were made. Maybe that's beyond their mission scope, though.

Re: Thank you for helping us increase our bandwidth

#113

I'm really surprised they don't use more CDN for this. Anyone know the reason why it isn't served by something like cloudflare?

Sibling comments are exactly correct -- we're a library. We don't track or record the activity of our patrons (see e.g. https://archive.org/services/docs/api/views.html#footnote-wh... ), and since most CDNs cannot offer the same guarantee, we can't use them.

Re: Thank you for helping us increase our bandwidth

#114

For the video and audio content. Could they take advantage of something like Webtorrent.

Every item has a torrent file, and can be retrieved with a BitTorrent client. https://help.archive.org/hc/en-us/articles/360004715251-Arch...

Brilliant!

Re: Thank you for helping us increase our bandwidth

#115
post #17

60 Gbit/s of continuous traffic is a lot. If I'm reading the graphs right, Wikipedia "only" has 13.4 Gbit/s of outbound traffic [1]. Of course still well below the single-digit Tbit/s traffic of large internet exchanges [2] [3], but still unexpectedly large. [1]: adding up the outbound numbers for each datacenter: 1.888 + 8.003 + 810 + 1.958 + 807: https://grafana.wikimedia.org/d/000000605/datacenter-global-... [2]:…

60 Gbps is about 1/3 of the traffic served by a single Netflix CDN node.

It's harder when you don't make money, and everyone's not downloading the same things.

Re: Thank you for helping us increase our bandwidth

#116

Earlier quoted context omitted.

We don't save logs of who hits us, but we watch who hits us. It's not bots. But good thinking.

How do you know? Not doubting you, just curious as to how you figure out if a request is from a bot or not.

Well, obviously, a sneaky bot is a sneaky sneak and acts like a person. But conversely, people act in a way that could be like a bot if anyone casually looked at them - mass downloads, no theme or meaning to content. In general, however, it's all pretty much people. Millions of people a day.

Re: Thank you for helping us increase our bandwidth

#118

Earlier quoted context omitted.

At the bandwidth levels they are using they would need to use Cloudflare Enterprise, and in my experience that is way more expensive than other CDN providers. Also, since archive.org has so much content, the caching ratio is going to be very bad and kill CDN efficiency while still requiring lots of direct bandwidth. Cheap direct bandwidth in their case looks best(which is what they seem to be doing).

Another aspect is they are actually using bandwidth to max and have a cap on usage spending. No CDN I know of allows capping bandwidth usage. The way they use bandwidth means the most efficent GB/dollar cost although at the cost of poor performance.

[deleted]

Re: Thank you for helping us increase our bandwidth

#119

Earlier quoted context omitted.

> single Netflix CDN node. Looks like they max out at 100 Gbps, or as low as 40 Gbps, depending on appliance model and link aggregation configuration. No argument either way, just thought it was cool info. https://openconnect.zendesk.com/hc/en-us/articles/3600345383...

That's out of date. Here's a presentation from last year where they talk about >190 Gbps: https://people.freebsd.org/~gallatin/talks/euro2019.pdf

The video accompanying this slide deck is also a great watch: https://www.youtube.com/watch?v=8NSzkYSX5nY

Re: Thank you for helping us increase our bandwidth

#120
post #100
post #88

Earlier quoted context omitted.

You can get an 8TB HDD for $150 right now. That's 6250 drives. That's about $1MM in drives, which doesn't sound that cost-prohibitive. Obviously that's not the whole cost since you need to pay for bandwidth, replication and other infrastructure like the host node, but it sounds like something that could be even hosted by a number of volunteers on r/datahoarder or r/homelab. I also remember reading about Sia on HN, wh…

BackBlaze B2 is $5/TB/month Azure Archive is $2/TB/month ($1.68 if reserved) AWS Glacier Deep Archive is $1/TB/month GCP Cloud Storage Archive is $1.20/TB/month Of course, there can be i/o and network charges, and different levels of redundancy (but possibly bulk discounts)...but the bare storage costs for for 50 PB per year would be roughly $600k - $3 MM/y.

The cloud business is a small fraction of Amazon's revenue but a large part of their profits. It's extremely profitable for them. That's why there is such a large discrepancy between (non bulk) HDD price and (non bulk) per month cost for archival.
Post reply on HN