Live data from Hacker News

Running ArchiveTeam's Warrior in Kubernetes

gabrielsimmer.com

1–10 of 40 posts

Re: Running ArchiveTeam's Warrior in Kubernetes

#2
Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular.

The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page into the Wayback Machine every week. Or at least trying to.

https://web.archive.org/web/20250122000033/www.google.com

Like so many things about archive.org, when you dig in you start to find wonder and craziness at every turn.

Re: Running ArchiveTeam's Warrior in Kubernetes

#3

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

> by proper entities as required by federal law.

What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?

Re: Running ArchiveTeam's Warrior in Kubernetes

#4

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

How do I as a non-US citizen get access to information from those "proper entities"? Is it even possible for US citizens? This is often a surprise for some visitors of this fine website, but there's a large world outside the US where "federal law" does not apply.

Re: Running ArchiveTeam's Warrior in Kubernetes

#5

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

> by proper entities as required by federal law. What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?

Some of the mass deletions are merely a new administration setting up shop. Policies from the previous administration don't belong on the current whitehouse.gov. They wind up here instead https://bidenwhitehouse.archives.gov/

We pay half a billion in tax dollars for the National Archives, and nearly a billion to the Library of Congress to preserve these records. Others are managed as part of Presidential Libraries.

Thousands of employees, dozens of facilities, billions of dollars.

Meanwhile archive.org doesn't have air conditioning and preserves physical material within the blast radius of an oil refinery. They let vagrants sleep on their steps yet seem surprised when they set the utility pole outsides on fire.

I didn't say it didn't need to be done. I said the whole process needs to be rethought with professional supervision. Setting up more volunteer K8 clusters so that more copies of the Google Home Page can be captured with the wrong user agent isn't going to save democracy.

Re: Running ArchiveTeam's Warrior in Kubernetes

#6

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

How do I as a non-US citizen get access to information from those "proper entities"? Is it even possible for US citizens? This is often a surprise for some visitors of this fine website, but there's a large world outside the US where "federal law" does not apply.

We fund the Library of Congress (largest library in the world) and the National Archives (NARA) who make all of this stuff public. Other goverments do similar things. It's all on the web.

https://www.archives.gov/presidential-records/research/archi...

There are other agencies and data sources to be monitored of course but I'm not seeing a lot of nuance in those efforts yet.

Re: Running ArchiveTeam's Warrior in Kubernetes

#7

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

[flagged]

Re: Running ArchiveTeam's Warrior in Kubernetes

#8
post #7

Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…

[flagged]

Suspicion warranted, but citation needed. For now my money is on archives.gov over archive.org.

And loc.gov over a website that prioritizes making Pac-Man and Donkey Kong playable in the browser yet leaked the drivers licenses and passports of its patrons, and whose public policy on their wonky javascript UX is "don't read books on a phone."

Re: Running ArchiveTeam's Warrior in Kubernetes

#9

Earlier quoted context omitted.

> by proper entities as required by federal law. What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?

Some of the mass deletions are merely a new administration setting up shop. Policies from the previous administration don't belong on the current whitehouse.gov. They wind up here instead https://bidenwhitehouse.archives.gov/ We pay half a billion in tax dollars for the National Archives, and nearly a billion to the Library of Congress to preserve these records. Others are managed as part of Presidential Libraries. T…

Archive.org is outside of the reach of the US government, and is globally distributed. When the US government deletes or darks data (as it has recently done across wide swaths of the federal government website properties), you have no recourse. This means your argument about the resources that go into the US government as a data custodian are meaningless: the outcome is what is material, which is the archival and long term custody & availability of the data sets in scope. Arguably, the Internet Archive has recently proven better at this job than the US government (unsurprising).

You're angry at a high value non profit operating on a limited budget. It's weird. I recommend focusing on more important issues than "it is icky around the richmond facility, the power goes out once in a while, and they use ambient air and convection for system cooling which I don't like."

If you want to save democracy, the Internet Archive doesn't do that itself. It protects the historical record. If you want to save democracy, that's a different conversation.

https://blog.archive.org/2024/05/08/end-of-term-web-archive/

https://web.archive.org/collection-search/EndOfTerm2024PreEl...

(no affiliation)

Post reply on HN