Live data from Hacker News

Internet Archive Infrastructure

archive.org

91–100 of 114 posts

Re: Internet Archive Infrastructure

#91
post #46

When I was a kid, my dreams was to work at Google, Microsoft, Apple or any of these big companies. Now that I am reaching 30 (and becoming very nostlagic of the old web), I think the company that would make me most happy of login every morning to get some work done would be Internet Archive.

I love archive.org. Hate their "front end". One of the most disorganized user-facing sites. I would love to see an effort to address that. Would it be possible for archive.org to offer an API to allow other sites to present archive.org using their own front-end? We could see lots of specialty sites that focus on the user experience for some slice of archive.org.

There are APIs and it would be great if more people and organizations built on top of them, and specifically build content or collection-specific interfaces.

Here is the entry point for API documentation: https://archive.org/services/docs/api/

Hot linking, CORS, and other things to support third-party integration are usually supported, though there are a lot of special cases for security or to prevent malicious use. If you run in to technical problems we are usually responsive to the main contact routes on the archive.org site.

The system is not designed to allow multi-party curation and editing of metadata, but there is nothing stopping folks from building third-party catalogs on top of the content stored (and served) from IA. That is sort of what openlibrary.org is for books. The same thing could be done for music, video, specific documents, etc.

Note: I work at IA but not on the APIs or archive.org collections

Re: Internet Archive Infrastructure

#92
post #17

Earlier quoted context omitted.

I get the impression Wayback Machine data is stored in powered down drives and they only spin them up when someone accesses the data. That would explain the several second delay and it'd make sense that an archive wouldn't need 95% of its data ready to go at a moments notice since that'd be a terrible waste of power.

The disks are spinning all the time, and most disks are seeing fairly frequent reads to some content or another. A lot of content is very rarely accesses, but almost every disk has some content which gets accessed. If spinning disks had only frequently-accessed content, they would be unable to keep up with the read rate or read throughput, things balance out reasonably on average. Wayback content is on the same disks…

Interesting thanks for the insights! Then would the few second delay be more a matter of time it takes to decompress the contents, or that files are stored on disks which are being accessed a lot, or something else? Always been curious about it.

Re: Internet Archive Infrastructure

#93

Earlier quoted context omitted.

Tracking should be STRICTLY illegal and ONLY acceptable with a verifiable OPT-IN and a transparency on to WHOM the data was sent/read-by/received. With STILL an option to selectively opt-out.

I agree with the sentiment, but there needs to be some degree of leeway. For example are server logs considered tracking? It seems unreasonable to require that logs not be kept. Edit: On further thought, I don't even know that tracking should be banned. Instead, I would argue that advertising in the way enabled by tracking should be banned. That way the incentive is removed with less bureaucracy.

bureaucracy IS the incentive.

We need to kill bureaucratic interests in EVERY THING.

Politicians should be conscripted servants with no method for empowering or financing themselves.

They should run on policy ALONE.

Re: Internet Archive Infrastructure

#94
post #71

Earlier quoted context omitted.

On the scale of big hosting operations, 60Gbps outbound is not that much. If you're buying full tables IP transit from major carriers at IX points, I've seen 10GbE for $700-900/mo, and 100GbE circuits for under $7k/month. Of course you wouldn't want to have just one transit, but I'm fairly sure if somebody said to me 'here's $20,000 a month to buy transit', on the west coast, it's within the realm of possible. Ideall…

It seems they are regularly maxing out their network infrastructure. If it's so cheap, how come they don't just buy more? Is it the cost of the actual hardware? (I know they recently upgraded)

They are maxing out the fiber links between their own datacenters, which is in the process of being addressed. If the bits can't get from the datacenter full of hard drives to the datacenter that connects to the internet, not much point in buying additional transit capacity.

Re: Internet Archive Infrastructure

#95
post #47
post #5

This is a recent video presentation by Jonah Edwards, who runs the Core Infrastructure Team at the Internet Archive. He explains the IA’s server, storage, and networking infrastructure, and then takes questions from other people at the Archive. I found it all interesting. But the main takeaway for me is his response to Brewster Kahle’s question, beginning at 13:06, about why the IA does everything in-house rather tha…

I have ADD and typically eschew watching video if I can get the same, quality content faster in text. I loved this and watched it to the end. To any that feel the need to hide or disparage this because it seems to promote doing things on your own vs in the cloud, this isn’t some tech ops Total Money Makeover where you read a book and you’re suddenly in some sort of anti-credit cult. This is hard shit, and it’s the ba…

Thank you for this! It was definitely geared towards an internal audience but it makes me very happy to know that it was enjoyed and appreciated more broadly.

I am going to get a transcript done and up soon as well -- I just gave the talk on Friday so haven't had time to do so yet.

Re: Internet Archive Infrastructure

#96

Earlier quoted context omitted.

I assume the 200PB of storage and 60Gbps egress bandwidth 24/365 they do would be _extremely_ pricy on AWS...

On the scale of big hosting operations, 60Gbps outbound is not that much. If you're buying full tables IP transit from major carriers at IX points, I've seen 10GbE for $700-900/mo, and 100GbE circuits for under $7k/month. Of course you wouldn't want to have just one transit, but I'm fairly sure if somebody said to me 'here's $20,000 a month to buy transit', on the west coast, it's within the realm of possible. Ideall…

Exactly this. There are some logistical complexities (e.g. some of our bandwidth is funded by the E-Rate Universal Service Program for libraries which runs on a July-June fiscal year and so rapid upgrades on that front aren't possible), but by and large egress bandwidth isn't our primary challenge. Intersite links as I noted in the video are the current big one, and that can and does involve occasional time-consuming construction -- but honestly over the past year, a combination of total blowout of my usual capacity planning (including equipment budgets) plus the logistical complexities of lockdown have resulted in slowness to upgrade as fast as we'd like to.

Re: Internet Archive Infrastructure

#97

Earlier quoted context omitted.

I agree with the sentiment, but there needs to be some degree of leeway. For example are server logs considered tracking? It seems unreasonable to require that logs not be kept. Edit: On further thought, I don't even know that tracking should be banned. Instead, I would argue that advertising in the way enabled by tracking should be banned. That way the incentive is removed with less bureaucracy.

bureaucracy IS the incentive. We need to kill bureaucratic interests in EVERY THING. Politicians should be conscripted servants with no method for empowering or financing themselves. They should run on policy ALONE.

> bureaucracy IS the incentive

I don't agree there. There is no reason that someone has building bureaucracy as they're goal. The goal is to get something out of it, which is accomplished via the bureaucracy.

> We need to kill bureaucratic interests in EVERY THING.

While I don't necessarily disagree, this is orthogonal to the original issue.

> Politicians should be conscripted servants with no method for empowering or financing themselves.

I don't think conscripted means what you think it means...

Also, that would be impossible to implement. Either you don't try to cut off every path, in which case you have the option of limiting bureaucracy, or you try to cut off every path, increasing bureaucracy.

> They should run on policy ALONE.

I agree. Also orthogonal to the topic at hand. Also, please let me know if you find a way of implementing this without requiring bureaucracy as a critical component.

Re: Internet Archive Infrastructure

#98
post #38
post #22

This is awesome! I love seeing companies run their own infrastructure. I wonder if they are using ZFS or just traditional RAID?

This is mentioned in the video, but they don’t use any form of RAID, just paired/mirrored drives in different physical locations. This is preferred for its simplicity and performance, and if I were in their position I would do the same thing.

Exactly correct. I suspect that in the not-too-distant future we will need to move to multi-disk filesystem-level clustering on the storage nodes for some of the reasons laid out in the talk but it's not at all unlikely that we retain the "disconnected mirror" abstraction for redundancy.

Additionally, as I alluded to in the video, we regularly end up going down into the details and when this happens being able to closely examine and understand a disk's contents as written directly at the LBA using fibmaps and similar tooling is invaluable. I have personally been involved in the discovery of multiple possible-data-loss hard drive firmware bugs, and our catalog system is very paranoid about integrity checking. Modern hard drives are computers unto themselves and layering additional complexity on top of that (particularly complexity which might obscure or silently correct errors in the underlying datastream) is something I approach very carefully.

Re: Internet Archive Infrastructure

#99
post #71

Earlier quoted context omitted.

On the scale of big hosting operations, 60Gbps outbound is not that much. If you're buying full tables IP transit from major carriers at IX points, I've seen 10GbE for $700-900/mo, and 100GbE circuits for under $7k/month. Of course you wouldn't want to have just one transit, but I'm fairly sure if somebody said to me 'here's $20,000 a month to buy transit', on the west coast, it's within the realm of possible. Ideall…

It seems they are regularly maxing out their network infrastructure. If it's so cheap, how come they don't just buy more? Is it the cost of the actual hardware? (I know they recently upgraded)

We are upgrading again-- the pandemic has made us kind-of popular.

Because of budget we tend to use things up 100% before putting more in.

Re: Internet Archive Infrastructure

#100

Earlier quoted context omitted.

I love archive.org. Hate their "front end". One of the most disorganized user-facing sites. I would love to see an effort to address that. Would it be possible for archive.org to offer an API to allow other sites to present archive.org using their own front-end? We could see lots of specialty sites that focus on the user experience for some slice of archive.org.

There are APIs and it would be great if more people and organizations built on top of them, and specifically build content or collection-specific interfaces. Here is the entry point for API documentation: https://archive.org/services/docs/api/ Hot linking, CORS, and other things to support third-party integration are usually supported, though there are a lot of special cases for security or to prevent malicious use.…

Besides using the cli, how do you upload HTML URLs to be indexed?
Post reply on HN