I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?
The data is saved as a WARC file, which contains the entire HTTP request and response (compressed, of course). So it's much bigger than just a short -> long URL mapping.
ArchiveTeam has finished archiving all goo.gl short links
91–100 of 112 posts
Re: ArchiveTeam has finished archiving all goo.gl short links
#92Earlier quoted context omitted.
you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets
No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org
Re: ArchiveTeam has finished archiving all goo.gl short links
#93Earlier quoted context omitted.
I did some ridiculous napkin math. A random URL I pulled from a Google search was 705 bytes. A googl link is 22 bytes but if you only store the ID, it'd be 6 bytes. Some URLs are going to be shorter, some longer, but just ballparking it all, that lands us in the neighborhood of hundreds of billions of URLs, up to trillions of URLs.
> A random URL I pulled from a Google search was 705 bytes. 705 bytes is an extremely long URL. Even if we assume that URLs that get shortened tend to be longer than URLs overall, that’s still an unrealistic average.
Re: ArchiveTeam has finished archiving all goo.gl short links
#94Earlier quoted context omitted.
No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org
Tangentially related but I've seen twitter links that used to be on the wayback machine disappear from it at some point, presumably due to personal request from the owner.
Re: ArchiveTeam has finished archiving all goo.gl short links
#95Re: ArchiveTeam has finished archiving all goo.gl short links
#96Earlier quoted context omitted.
i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.
Feels like a bit of a kick in the teeth that I contributed towards archiving something that I don’t even get access to. What happens if they disappear? The dataset is gone forever.
Re: ArchiveTeam has finished archiving all goo.gl short links
#97Earlier quoted context omitted.
i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.
Who fears they will get blocked by whom?
Re: ArchiveTeam has finished archiving all goo.gl short links
#98Earlier quoted context omitted.
You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyruss…
i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.
But, ok, let's continue in good faith
scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive Warrior or its usages
scenario 2: they don't want to expose the target of the redirects because it will feed the boundaries of the ravenous AI slurp machines
If it's scenario 1, then CSV exists and allows mapping from the 00aa11 codes to the "location:" header, no means and methods necessary
If it's scenario 2, then what the hell were they expecting to happen? Embargo the .warc until the AI hype blows over so their great grand children can read about how the Internet was back in the day? I guess the real question is "archive for whom?" because right now unless they have a back-channel way to feed the Wayback Machine's boundary using the .warc files, and thus it secretly populates the Wayback without wholesale feeding the AI boundary, this whole thing is just mysterious
Re: ArchiveTeam has finished archiving all goo.gl short links
#99Earlier quoted context omitted.
ArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)
Ran archive warrior a while back but hadde to shut it down AS i sterted seeing the VM was compromised trying to spam ssh and other login attemps in my local network.
Is that the story, or you are saying that the machine was secured correctly but that running Warrior somehow introduced your network to risk?
Re: ArchiveTeam has finished archiving all goo.gl short links
#100Earlier quoted context omitted.
ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project. https://github.com/ArthurHeitmann/arctic_shift
Thanks. I wonder if anyone does this for hacker news.
Given that Firebase (which powers the API link at the bottom of this page) is a Google property, I cannot possibly imagine why they'd differ