Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

51–60 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#51
post #15

Title is imprecise, it's Archiveteam.org, not Archive.org. The Internet Archive is providing free hosting, but the archival work was done by Archiveteam members.

What exactly is archiveteam's contribution? I don't fully understand. Edit: Like they kinda seem like an unnecessary middle-man between the archive and archivee, but maybe I'm missing something.

Archive Team is carrying books in a bucket brigade out of the burning library. Archive.org is giving them a place to put the books they saved.

Re: ArchiveTeam has finished archiving all goo.gl short links

#52

Why? Did they ask anyone if it was okay? Anything sensitive at those links? Anything at those links people didn't want or need anymore? Maybe people thought those links were dead? Did Google provide a way to cancel those links first? It's like when the GPT links were archived and publicly available that contained sensitive information.

Sometimes to preserve history, you just have to go do what you gotta do.

After all, these are just short links. They link to other things on the Internet. Which is inherently public anyways.

You cannot expect privacy via a simple URL. These short URLs are short, hence programmatically scraping all the URLs.

The GPT Links situation is nothing like this imo. Both however do come down to the stupid human aspect.

Re: ArchiveTeam has finished archiving all goo.gl short links

#54

Why? Did they ask anyone if it was okay? Anything sensitive at those links? Anything at those links people didn't want or need anymore? Maybe people thought those links were dead? Did Google provide a way to cancel those links first? It's like when the GPT links were archived and publicly available that contained sensitive information.

It's a link, what privacy can one expect?

Especially with short links there's always the possibility of entering ~6 characters and getting a hit. So I believe expecting any secrecy from urls is silly.

That's like posting your passwords on Twitter because "Why would anyone find my account"

Re: ArchiveTeam has finished archiving all goo.gl short links

#56
post #35

Earlier quoted context omitted.

I was going to be cheeky and say "well, now you can download them and search" but it seems it's "Access-restricted-item: true" for some reason, above and beyond being 10G a pop https://archive.org/details/archiveteam_googl_20250228144231... >

you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets

No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org

Re: ArchiveTeam has finished archiving all goo.gl short links

#57
post #38

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

I did some ridiculous napkin math. A random URL I pulled from a Google search was 705 bytes. A googl link is 22 bytes but if you only store the ID, it'd be 6 bytes. Some URLs are going to be shorter, some longer, but just ballparking it all, that lands us in the neighborhood of hundreds of billions of URLs, up to trillions of URLs.

> A random URL I pulled from a Google search was 705 bytes.

705 bytes is an extremely long URL. Even if we assume that URLs that get shortened tend to be longer than URLs overall, that’s still an unrealistic average.

Re: ArchiveTeam has finished archiving all goo.gl short links

#58
post #56

Earlier quoted context omitted.

you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets

No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org

I can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.

Re: ArchiveTeam has finished archiving all goo.gl short links

#60

Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.

ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project.

https://github.com/ArthurHeitmann/arctic_shift

Post reply on HN