Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

31–40 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#31
post #15

Title is imprecise, it's Archiveteam.org, not Archive.org. The Internet Archive is providing free hosting, but the archival work was done by Archiveteam members.

What exactly is archiveteam's contribution? I don't fully understand. Edit: Like they kinda seem like an unnecessary middle-man between the archive and archivee, but maybe I'm missing something.

> Like they kinda seem like an unnecessary middle-man between the archive and archivee

They are the middlemen that collects the data to be archived.

In this example the archivee (goo.gl/Alphabet) is simply shutting the service down and has no interest in archiving it. Archive.org is willing to host the data, but only if somebody brings it to them. Archiveteam writes and organises crawlers to collect the data and send it to Archive.org

Re: ArchiveTeam has finished archiving all goo.gl short links

#32
post #6
post #4

Earlier quoted context omitted.

They iterated the entire URL namespace by having volunteers run a client so they didn't get IP banned.

Beautiful. I wish I had seen this and could have helped.

they are still archiving other url shorteners https://tracker.archiveteam.org:1338/ you can participate in that

Re: ArchiveTeam has finished archiving all goo.gl short links

#33
Excellent! ArchiveTeam have always been impressive this way. Some years ago, I was working at a video platform that had just announced it would be shutting down fairly soon. I forget how, but one way or another I got connected with someone at ArchiveTeam who expressed their interest in archiving it all before it was too late. Believing this to be a good idea, I gave them a couple of tips about where some of our device-sniffing server endpoints were likely to give them a little trouble, and temporarily "donated" a couple EC2 instances to them to put towards their archiving tasks.

Since the servers were mine, I could see what was happening, and I was very impressed. Within I want to say two minutes, the instances had been fully provisioned and were actively archiving videos as fast as was possible, fully saturating the connection, with each instance knowing to only grab videos the other instances had not already gotten. Basically they have always struck me as not only having a solid mission, but also being ultra-efficient in how they carry it out.

Re: ArchiveTeam has finished archiving all goo.gl short links

#34
post #15

Earlier quoted context omitted.

What exactly is archiveteam's contribution? I don't fully understand. Edit: Like they kinda seem like an unnecessary middle-man between the archive and archivee, but maybe I'm missing something.

What ArchiveTeam mainly does is provide hand-made scripts to aggressively archive specific websites that are about to die, with a prioritization for things the community deems most endangered and most important. They provide a bot you can run to grab these scripts automatically and run them on your own hardware, to join the volunteer effort. This is in contrast to the Wayback Machine's builtin crawler, which is just…

I just made a root comment with my experience seeing their process at work, but yeah it really cannot be overstated how efficient and effective their archiving process is

Re: ArchiveTeam has finished archiving all goo.gl short links

#35
post #29

I wonder how many of them lead to private YouTube videos, Google documents, etc.

I was going to be cheeky and say "well, now you can download them and search" but it seems it's "Access-restricted-item: true" for some reason, above and beyond being 10G a pop https://archive.org/details/archiveteam_googl_20250228144231...>

Re: ArchiveTeam has finished archiving all goo.gl short links

#38

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

I did some ridiculous napkin math. A random URL I pulled from a Google search was 705 bytes. A googl link is 22 bytes but if you only store the ID, it'd be 6 bytes. Some URLs are going to be shorter, some longer, but just ballparking it all, that lands us in the neighborhood of hundreds of billions of URLs, up to trillions of URLs.

Re: ArchiveTeam has finished archiving all goo.gl short links

#39
post #30
post #23

Earlier quoted context omitted.

I don't understand the data on ArchiveTeam's page but, it seems like they have 35 terabytes of data (286.56TiB)? It's a lot larger than I'd have thought.

FYI, "TiB" means terabytes with a base of 1024, ie. the units you'd typically use for measuring memory rather than the units you'd typically see drive vendors using. The factor of 8 you divided by only applies to units based on bits rather than bytes , and those units use "b" rather than "B", and are only used for capacity measurements when talking about individual memory dies (though they're normal for talking about…

The binary units like GiB, TiB, are technically supposed to be Gibibytes and Tebibytes. Thought it was a bit silly when they first popped up but now I find them adorkably endearing, and a good way to disambiguate something that's often left vague at your expense.

Re: ArchiveTeam has finished archiving all goo.gl short links

#40
post #9
post #7

Recent update from Google: https://blog.google/technology/developers/googl-link-shorten...

This leaves me wondering what the point is? What could it possibly cost to keep redirecting existing shortlinks that they consider unused/low activity already anyway ? (In addition to the higher activity ones parent link says they'll now continue to redirect.)

In another submission someone speculated the reason might be the unending churn of the Google tech stack that just makes low-maintenance stuff impossible.
Post reply on HN