Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

91–100 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#91

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

The data is saved as a WARC file, which contains the entire HTTP request and response (compressed, of course). So it's much bigger than just a short -> long URL mapping.

[deleted]

Re: ArchiveTeam has finished archiving all goo.gl short links

#92
post #56

Earlier quoted context omitted.

you'd have to rescrape them all from https://web.archive.org/cdx/search?url=goo.gl/* - they don't publish the whole datasets

No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org

Tangentially related but I've seen twitter links that used to be on the wayback machine disappear from it at some point, presumably due to personal request from the owner.

Re: ArchiveTeam has finished archiving all goo.gl short links

#93
post #57
post #38

Earlier quoted context omitted.

I did some ridiculous napkin math. A random URL I pulled from a Google search was 705 bytes. A googl link is 22 bytes but if you only store the ID, it'd be 6 bytes. Some URLs are going to be shorter, some longer, but just ballparking it all, that lands us in the neighborhood of hundreds of billions of URLs, up to trillions of URLs.

> A random URL I pulled from a Google search was 705 bytes. 705 bytes is an extremely long URL. Even if we assume that URLs that get shortened tend to be longer than URLs overall, that’s still an unrealistic average.

It is long, it represents the lower hundreds of billions bound in my awful napkin math.

Re: ArchiveTeam has finished archiving all goo.gl short links

#94
post #92
post #56

Earlier quoted context omitted.

No, I meant the .warc.zst files on archive.org that were the result of the ArchiveTeam's work. However, it seems they're under some kind of embargo - which is the first I've ever seen a private link on archive.org

Tangentially related but I've seen twitter links that used to be on the wayback machine disappear from it at some point, presumably due to personal request from the owner.

Pretty sure you can nuke all your domains old content by blocking archive.org in robots.txt

Re: ArchiveTeam has finished archiving all goo.gl short links

#96
post #76

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Feels like a bit of a kick in the teeth that I contributed towards archiving something that I don’t even get access to. What happens if they disappear? The dataset is gone forever.

You get access to it via the wayback machine

Re: ArchiveTeam has finished archiving all goo.gl short links

#97

Earlier quoted context omitted.

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Who fears they will get blocked by whom?

Archive team blocked by hosts wanting to protect their data from AI companies (presumably because they want to extract money from them)

Re: ArchiveTeam has finished archiving all goo.gl short links

#98
post #71

Earlier quoted context omitted.

You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyruss…

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

This whole thread is starting to read like some kind of misguided practical joke. I also recognize that it may seem like this is directed toward you, but I'm not shooting the messenger I'm just anchoring my reply under this new information. Sorry about that.

But, ok, let's continue in good faith

scenario 1: they don't want to uncork the .warc files because it will potentially leak the means and methods of the Archive Warrior or its usages

scenario 2: they don't want to expose the target of the redirects because it will feed the boundaries of the ravenous AI slurp machines

If it's scenario 1, then CSV exists and allows mapping from the 00aa11 codes to the "location:" header, no means and methods necessary

If it's scenario 2, then what the hell were they expecting to happen? Embargo the .warc until the AI hype blows over so their great grand children can read about how the Internet was back in the day? I guess the real question is "archive for whom?" because right now unless they have a back-channel way to feed the Wayback Machine's boundary using the .warc files, and thus it secretly populates the Wayback without wholesale feeding the AI boundary, this whole thing is just mysterious

Re: ArchiveTeam has finished archiving all goo.gl short links

#99
post #85
post #18

Earlier quoted context omitted.

ArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)

Ran archive warrior a while back but hadde to shut it down AS i sterted seeing the VM was compromised trying to spam ssh and other login attemps in my local network.

This smells like a one-click bringup went wrong, and not that the Warrior software was compromised

Is that the story, or you are saying that the machine was secured correctly but that running Warrior somehow introduced your network to risk?

Re: ArchiveTeam has finished archiving all goo.gl short links

#100

Earlier quoted context omitted.

ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project. https://github.com/ArthurHeitmann/arctic_shift

Thanks. I wonder if anyone does this for hacker news.

I believe there is a dataset in BigQuery but I haven't tried looking at it in order to know how uptodate it is https://news.ycombinator.com/item?id=10440502>

Given that Firebase (which powers the API link at the bottom of this page) is a Google property, I cannot possibly imagine why they'd differ

Post reply on HN