Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

101–110 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#101
post #95

Gamefaqs remains unarchived.

Be the change you want to see in the world. Contributing .warc files to Archive.org isn't a gated club. My understanding of calling down the Warrior team is when something is time sensitive and needs to pseudo-ddos the site to get the bytes right now. Unless you know something about the demise of Gamefaqs, you have the rest of your life to archive a page at a time

Re: ArchiveTeam has finished archiving all goo.gl short links

#102
post #98

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

This whole thread is starting to read like some kind of misguided practical joke. I also recognize that it may seem like this is directed toward you, but I'm not shooting the messenger I'm just anchoring my reply under this new information. Sorry about that. But, ok, let's continue in good faith scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive…

[dead]

Re: ArchiveTeam has finished archiving all goo.gl short links

#103
post #98

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

This whole thread is starting to read like some kind of misguided practical joke. I also recognize that it may seem like this is directed toward you, but I'm not shooting the messenger I'm just anchoring my reply under this new information. Sorry about that. But, ok, let's continue in good faith scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive…

i think you're missing some key information. the warcs do not just contain the location header information. and their methods are fully public/open source so scenario 1 makes no sense.

sure maybe the warcs will be unlocked at some point in the future. this is a fairly small volunteer effort. i doubt there is some "unlock in 100 years" feature on IA.

Re: ArchiveTeam has finished archiving all goo.gl short links

#104

Why? Did they ask anyone if it was okay? Anything sensitive at those links? Anything at those links people didn't want or need anymore? Maybe people thought those links were dead? Did Google provide a way to cancel those links first? It's like when the GPT links were archived and publicly available that contained sensitive information.

If you want something to remain private, don't post it on the public internet.

Re: ArchiveTeam has finished archiving all goo.gl short links

#105
post #98

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

This whole thread is starting to read like some kind of misguided practical joke. I also recognize that it may seem like this is directed toward you, but I'm not shooting the messenger I'm just anchoring my reply under this new information. Sorry about that. But, ok, let's continue in good faith scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive…

Yes exactly, Wayback Machine can use the warc files despite them being blocked for direct download.

Re: ArchiveTeam has finished archiving all goo.gl short links

#106

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

The 91 TiB includes not just the URL mappings but the actual content of all destination pages, which ArchiveTeam captures to ensure the links remain functional even if original destinations disappear.

Ok but the destination pages are not at risk (or at least not any more than any random page on the web) so why spend any effort crawling them before all shortcuts have been saved?

Re: ArchiveTeam has finished archiving all goo.gl short links

#107
post #69
post #66

Earlier quoted context omitted.

Then why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret? I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings

Because then you can access the archived destination if you already know the short URL. You just can't get a full list of potentially sensitive short URL/destination pairs.

Yeah what they did is probably the best way to handle it.

Re: ArchiveTeam has finished archiving all goo.gl short links

#108
post #9

Earlier quoted context omitted.

This leaves me wondering what the point is? What could it possibly cost to keep redirecting existing shortlinks that they consider unused/low activity already anyway ? (In addition to the higher activity ones parent link says they'll now continue to redirect.)

In another submission someone speculated the reason might be the unending churn of the Google tech stack that just makes low-maintenance stuff impossible.

My guess is that plus not having a single person left to maintain it due to the similarly unending people churn.

Re: ArchiveTeam has finished archiving all goo.gl short links

#109

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

They might be storing in WARC format, which records all the request and response headers and maybe even TLS certificates and things.

Re: ArchiveTeam has finished archiving all goo.gl short links

#110
post #76

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Feels like a bit of a kick in the teeth that I contributed towards archiving something that I don’t even get access to. What happens if they disappear? The dataset is gone forever.

This does seem off. Especially as I can navigate to any of those URLs myself. Hell, if I wanted to spin up 50 virtual servers and go crazy I could probably pay a few thousand bucks to re-scrape the thing myself.
Post reply on HN