Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

61–70 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#62
post #18

Earlier quoted context omitted.

ArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)

Sidenote, but you can also run a Warrior in Docker, which is sometimes easier to set up (e.g. if you already have a server with other apps in containers).

Yep, I have my archiveteam warrior running in the built-in Docker GUI on my Synology NAS. Just a few clicks to set up and it just runs there silently in the background, helping out with whatever tasks it needs to.

Re: ArchiveTeam has finished archiving all goo.gl short links

#63

Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.

ArchiveTeam was doing that, but their stuff no longer works due to changes at Reddit. The wiki page about it links to some other groups doing Reddit archiving.

https://wiki.archiveteam.org/index.php/Reddit

Re: ArchiveTeam has finished archiving all goo.gl short links

#64

Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.

ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project. https://github.com/ArthurHeitmann/arctic_shift

Viewer and stats for ArcticShift: https://photon-reddit.com/ https://arctic-shift.photon-reddit.com/

Re: ArchiveTeam has finished archiving all goo.gl short links

#65
post #39
post #30

Earlier quoted context omitted.

FYI, "TiB" means terabytes with a base of 1024, ie. the units you'd typically use for measuring memory rather than the units you'd typically see drive vendors using. The factor of 8 you divided by only applies to units based on bits rather than bytes , and those units use "b" rather than "B", and are only used for capacity measurements when talking about individual memory dies (though they're normal for talking about…

The binary units like GiB, TiB, are technically supposed to be Gibibytes and Tebibytes. Thought it was a bit silly when they first popped up but now I find them adorkably endearing, and a good way to disambiguate something that's often left vague at your expense.

In my experience, nobody actually says "Tebibytes" out loud; it's just that silly. In writing, when the precision is necessary, the abbreviation "TiB" does see some actual use.

Re: ArchiveTeam has finished archiving all goo.gl short links

#66
post #58
post #56

Earlier quoted context omitted.

No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org

I can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.

Then why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret?

I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings

Re: ArchiveTeam has finished archiving all goo.gl short links

#67

Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.

Academictorrents has monthly dumps of all reddit submissions and comments even after the API restrictions.

https://academictorrents.com/browse.php?search=stuck_in_the_...

Re: ArchiveTeam has finished archiving all goo.gl short links

#69
post #66
post #58

Earlier quoted context omitted.

I can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.

Then why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret? I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings

Because then you can access the archived destination if you already know the short URL. You just can't get a full list of potentially sensitive short URL/destination pairs.
Post reply on HN