Running ArchiveTeam's Warrior in Kubernetes
gabrielsimmer.com
Running ArchiveTeam's Warrior in Kubernetes
1–10 of 40 posts
Re: Running ArchiveTeam's Warrior in Kubernetes
#2The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page into the Wayback Machine every week. Or at least trying to.
https://web.archive.org/web/20250122000033/www.google.com
Like so many things about archive.org, when you dig in you start to find wonder and craziness at every turn.
Re: Running ArchiveTeam's Warrior in Kubernetes
#3Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?
Re: Running ArchiveTeam's Warrior in Kubernetes
#4Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
Re: Running ArchiveTeam's Warrior in Kubernetes
#5Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
> by proper entities as required by federal law. What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?
We pay half a billion in tax dollars for the National Archives, and nearly a billion to the Library of Congress to preserve these records. Others are managed as part of Presidential Libraries.
Thousands of employees, dozens of facilities, billions of dollars.
Meanwhile archive.org doesn't have air conditioning and preserves physical material within the blast radius of an oil refinery. They let vagrants sleep on their steps yet seem surprised when they set the utility pole outsides on fire.
I didn't say it didn't need to be done. I said the whole process needs to be rethought with professional supervision. Setting up more volunteer K8 clusters so that more copies of the Google Home Page can be captured with the wrong user agent isn't going to save democracy.
Re: Running ArchiveTeam's Warrior in Kubernetes
#6Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
How do I as a non-US citizen get access to information from those "proper entities"? Is it even possible for US citizens? This is often a surprise for some visitors of this fine website, but there's a large world outside the US where "federal law" does not apply.
https://www.archives.gov/presidential-records/research/archi...
There are other agencies and data sources to be monitored of course but I'm not seeing a lot of nuance in those efforts yet.
Re: Running ArchiveTeam's Warrior in Kubernetes
#7Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
Re: Running ArchiveTeam's Warrior in Kubernetes
#8Many of these sites are already captured and archived by proper entities as required by federal law. More is better, I guess, except when it isn't. Duplication of effort is a huge problem in the humanities in general and with archiving in particular. The whole concept needs to be rethought. Captures from these tools show up under "ArchiveTeam" which is currently pumping thousands of copies of the Google Home Page int…
[flagged]
And loc.gov over a website that prioritizes making Pac-Man and Donkey Kong playable in the browser yet leaked the drivers licenses and passports of its patrons, and whose public policy on their wonky javascript UX is "don't read books on a phone."
Re: Running ArchiveTeam's Warrior in Kubernetes
#9Earlier quoted context omitted.
> by proper entities as required by federal law. What federal law do you suppose is guiding the mass deletions? That doesn't look like archiving to me. Now that the foxes are running the henhouse, how reliable do you suppose their own archives are?
Some of the mass deletions are merely a new administration setting up shop. Policies from the previous administration don't belong on the current whitehouse.gov. They wind up here instead https://bidenwhitehouse.archives.gov/ We pay half a billion in tax dollars for the National Archives, and nearly a billion to the Library of Congress to preserve these records. Others are managed as part of Presidential Libraries. T…
You're angry at a high value non profit operating on a limited budget. It's weird. I recommend focusing on more important issues than "it is icky around the richmond facility, the power goes out once in a while, and they use ambient air and convection for system cooling which I don't like."
If you want to save democracy, the Internet Archive doesn't do that itself. It protects the historical record. If you want to save democracy, that's a different conversation.
https://blog.archive.org/2024/05/08/end-of-term-web-archive/
https://web.archive.org/collection-search/EndOfTerm2024PreEl...
(no affiliation)
Re: Running ArchiveTeam's Warrior in Kubernetes
#10https://github.com/ArchiveTeam/warrior-dockerfile/blob/maste...