Ok how do I access them, or is that not the point?
ArchiveTeam has finished archiving all goo.gl short links
61–70 of 112 posts
Re: ArchiveTeam has finished archiving all goo.gl short links
#62Earlier quoted context omitted.
ArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)
Sidenote, but you can also run a Warrior in Docker, which is sometimes easier to set up (e.g. if you already have a server with other apps in containers).
Re: ArchiveTeam has finished archiving all goo.gl short links
#63Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.
Re: ArchiveTeam has finished archiving all goo.gl short links
#64Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.
ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project. https://github.com/ArthurHeitmann/arctic_shift
Re: ArchiveTeam has finished archiving all goo.gl short links
#65Earlier quoted context omitted.
FYI, "TiB" means terabytes with a base of 1024, ie. the units you'd typically use for measuring memory rather than the units you'd typically see drive vendors using. The factor of 8 you divided by only applies to units based on bits rather than bytes , and those units use "b" rather than "B", and are only used for capacity measurements when talking about individual memory dies (though they're normal for talking about…
The binary units like GiB, TiB, are technically supposed to be Gibibytes and Tebibytes. Thought it was a bit silly when they first popped up but now I find them adorkably endearing, and a good way to disambiguate something that's often left vague at your expense.
Re: ArchiveTeam has finished archiving all goo.gl short links
#66Earlier quoted context omitted.
No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org
I can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.
I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings
Re: ArchiveTeam has finished archiving all goo.gl short links
#67Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.
Academictorrents has monthly dumps of all reddit submissions and comments even after the API restrictions.
Re: ArchiveTeam has finished archiving all goo.gl short links
#68Ok how do I access them, or is that not the point?
Re: ArchiveTeam has finished archiving all goo.gl short links
#69Earlier quoted context omitted.
I can see some reasonable arguments for not publishing the full dataset. People undoubtedly shortened lots of links to unlisted videos/documents/pages under the assumption that the short link, like the original link, would be unguessable.
Then why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret? I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings