This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…
I assume the 200PB of storage and 60Gbps egress bandwidth 24/365 they do would be _extremely_ pricy on AWS...
Internet Archive Infrastructure
41–50 of 114 posts
Re: Internet Archive Infrastructure
#42Have they lost data during rebuild?I know he briefly talked about that risk with different HD sizes.
Re: Internet Archive Infrastructure
#43Earlier quoted context omitted.
On the scale of big hosting operations, 60Gbps outbound is not that much. If you're buying full tables IP transit from major carriers at IX points, I've seen 10GbE for $700-900/mo, and 100GbE circuits for under $7k/month. Of course you wouldn't want to have just one transit, but I'm fairly sure if somebody said to me 'here's $20,000 a month to buy transit', on the west coast, it's within the realm of possible. Ideall…
It seems that companies can be too big for the cloud and too small for the cloud (don't need k8s). I wonder what's the sweet spot.
Re: Internet Archive Infrastructure
#44Earlier quoted context omitted.
I assume the 200PB of storage and 60Gbps egress bandwidth 24/365 they do would be _extremely_ pricy on AWS...
I have no clue how people can afford what AWS charges for bandwidth. I did the math once for migrating a project to AWS and the bandwidth alone costed 10x my entire current infrastructure for that project, which is something I run for free.
And it is not only a cost comparison. You need different kind of people to manage in-house vs cloud, not that you need less or more, just different skills
Re: Internet Archive Infrastructure
#45Want to be part of changing the world?
Re: Internet Archive Infrastructure
#46Re: Internet Archive Infrastructure
#47This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…
I loved this and watched it to the end.
To any that feel the need to hide or disparage this because it seems to promote doing things on your own vs in the cloud, this isn’t some tech ops Total Money Makeover where you read a book and you’re suddenly in some sort of anti-credit cult. This is hard shit, and it’s the basics of the hard shit that I grew up with as being the only shit.
Yes, you can serve your own data. No one should fault you for doing that if you want. It takes the humble intelligence of the core team and everyone at IA to pull that off at this scale. If you don’t want to do the hard things, you could use the cloud. There are financial reasons also for one or the other, just as there are reasons people live with their family, rent, lease, and buy homes and office space- an imperfect analogy of course.
I hope that some of the others that could go on to work at the big guys or have been working there and want a challenge consider applying to IA when there’s an opening. They’ve done an incredible job, and I look forward to the cool things they accomplish in the future.
Re: Internet Archive Infrastructure
#48This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…
Curious about the cost. Does that already include manpower and various acquisition cost of constructing their internal network (hardware, fiber link between site)? I guess the biggest downside is the speed of scaling that they can do. As it is limited by how fast they can purchase and install new storage device. But with the use case of Internet Archive, that shouldn't matter much.
Re: Internet Archive Infrastructure
#49Take for example this newer document: https://archive.org/details/manualzilla-id-5695071 The document has horrible name and useless tags, and the content seems to be only section 2 of some SW manual. How would I ever hope to find it if I needed that exact document?
Obviously such huge archive cannot be categorized and annotated by small team, so it would make sense to crowdsource the labeling process. Yet, as registered user I can only flag the item, or write a review. Why doesn't IA let users label content and build their own curated collections of items?
Re: Internet Archive Infrastructure
#50I'm looking for some badass engineers to join our team to create something that is going to take over the advertising model. YES, its true and it has tobe done.... Want to be part of changing the world?