Live data from Hacker News

Internet Archive Infrastructure

archive.org

41–50 of 114 posts

Re: Internet Archive Infrastructure

#41
post #5

This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…

I assume the 200PB of storage and 60Gbps egress bandwidth 24/365 they do would be _extremely_ pricy on AWS...

A hybrid solution might be possible. Put your core infra on AWS, but get a very cheap CDN (or custom solution) in front of it to handle the 60Gbps so that only a small fraction will hit your AWS infra. Do the same for storage, e.g. build your own ceph cluster on bare-metal instead of Amazon S3.

Re: Internet Archive Infrastructure

#42
It wasn't clear to me, but I might have just missed it. Is there a backup strategy beyond just duplicate data?

Have they lost data during rebuild?I know he briefly talked about that risk with different HD sizes.

Re: Internet Archive Infrastructure

#43

Earlier quoted context omitted.

On the scale of big hosting operations, 60Gbps outbound is not that much. If you're buying full tables IP transit from major carriers at IX points, I've seen 10GbE for $700-900/mo, and 100GbE circuits for under $7k/month. Of course you wouldn't want to have just one transit, but I'm fairly sure if somebody said to me 'here's $20,000 a month to buy transit', on the west coast, it's within the realm of possible. Ideall…

It seems that companies can be too big for the cloud and too small for the cloud (don't need k8s). I wonder what's the sweet spot.

horizontal and vertical scaling are the latest push, but diagonal would be the sweet spot

Re: Internet Archive Infrastructure

#44

Earlier quoted context omitted.

I assume the 200PB of storage and 60Gbps egress bandwidth 24/365 they do would be _extremely_ pricy on AWS...

I have no clue how people can afford what AWS charges for bandwidth. I did the math once for migrating a project to AWS and the bandwidth alone costed 10x my entire current infrastructure for that project, which is something I run for free.

Because people have nothing to compare their AWS cost to. They don't know how much it would cost them to host their service outside of AWS.

And it is not only a cost comparison. You need different kind of people to manage in-house vs cloud, not that you need less or more, just different skills

Re: Internet Archive Infrastructure

#46
When I was a kid, my dreams was to work at Google, Microsoft, Apple or any of these big companies. Now that I am reaching 30 (and becoming very nostlagic of the old web), I think the company that would make me most happy of login every morning to get some work done would be Internet Archive.

Re: Internet Archive Infrastructure

#47
post #5

This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…

I have ADD and typically eschew watching video if I can get the same, quality content faster in text.

I loved this and watched it to the end.

To any that feel the need to hide or disparage this because it seems to promote doing things on your own vs in the cloud, this isn’t some tech ops Total Money Makeover where you read a book and you’re suddenly in some sort of anti-credit cult. This is hard shit, and it’s the basics of the hard shit that I grew up with as being the only shit.

Yes, you can serve your own data. No one should fault you for doing that if you want. It takes the humble intelligence of the core team and everyone at IA to pull that off at this scale. If you don’t want to do the hard things, you could use the cloud. There are financial reasons also for one or the other, just as there are reasons people live with their family, rent, lease, and buy homes and office space- an imperfect analogy of course.

I hope that some of the others that could go on to work at the big guys or have been working there and want a challenge consider applying to IA when there’s an opening. They’ve done an incredible job, and I look forward to the cool things they accomplish in the future.

Re: Internet Archive Infrastructure

#48
post #28
post #5

This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…

Curious about the cost. Does that already include manpower and various acquisition cost of constructing their internal network (hardware, fiber link between site)? I guess the biggest downside is the speed of scaling that they can do. As it is limited by how fast they can purchase and install new storage device. But with the use case of Internet Archive, that shouldn't matter much.

He said his storage pricing is 2-5x cheaper than google archive line. That is $1.2/3x= $0.4/TB/month. Compare that against $20/month S3. He has 50x less cost. He can afford to overprovision.

Re: Internet Archive Infrastructure

#49
I see incredible value in IA collection of books, videos and software. OTOH I'm puzzled by lack of organization.

Take for example this newer document: https://archive.org/details/manualzilla-id-5695071 The document has horrible name and useless tags, and the content seems to be only section 2 of some SW manual. How would I ever hope to find it if I needed that exact document?

Obviously such huge archive cannot be categorized and annotated by small team, so it would make sense to crowdsource the labeling process. Yet, as registered user I can only flag the item, or write a review. Why doesn't IA let users label content and build their own curated collections of items?

Post reply on HN