Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

81–90 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#81
post #71

Earlier quoted context omitted.

You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyruss…

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Who fears they will get blocked by whom?

Re: ArchiveTeam has finished archiving all goo.gl short links

#82

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

3.75 billion URLs, according to this[1] the average URL is 76.97 characters would be ~268.8 GiB without the goo.gl id/metadata. So I also wonder whats up with that.

https://web.archive.org/web/20250125064617/http://www.superm...

Re: ArchiveTeam has finished archiving all goo.gl short links

#83

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

The data is saved as a WARC file, which contains the entire HTTP request and response (compressed, of course). So it's much bigger than just a short -> long URL mapping.

did they follow the redirect and archive the page content? but why?

Re: ArchiveTeam has finished archiving all goo.gl short links

#84
post #59

Google said they would keep hosting any recently-clicked link; does this mean that all the links are now recently-clicked?

“Recently clicked” wasn’t the criterium, it was “showed activity in late 2024”. So nothing that anybody has done this year – including this archiving – will affect which links Google keep alive.

Re: ArchiveTeam has finished archiving all goo.gl short links

#85
post #18
post #15

Earlier quoted context omitted.

What exactly is archiveteam's contribution? I don't fully understand. Edit: Like they kinda seem like an unnecessary middle-man between the archive and archivee, but maybe I'm missing something.

ArchiveTeam delegates tasks to volunteers and themselves running the Archive Warrior VM, which does the actual archiving. The resultant archives are then centralized by ArchiveTeam and uploaded to the Internet Archive. (Source: ran a Warrior)

Ran archive warrior a while back but hadde to shut it down AS i sterted seeing the VM was compromised trying to spam ssh and other login attemps in my local network.

Re: ArchiveTeam has finished archiving all goo.gl short links

#86

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

The 91 TiB includes not just the URL mappings but the actual content of all destination pages, which ArchiveTeam captures to ensure the links remain functional even if original destinations disappear.

Re: ArchiveTeam has finished archiving all goo.gl short links

#87
post #78

Can we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot m…

> This would mean there is an "official" source of all web data. LLM people can use snapshots of this that already exists, its called CommonCrawl: https://commoncrawl.org/

Common Crawl, while a massive dataset of the web does not represent the entirety of the web.

It’s smaller than Google’s index and Google does not represent the entirety of the web either.

For LLM training purposes this may or may not matter, since it does have a large amount of the web. It’s hard to prove scientifically whether the additional data would train a better model, because no one (afaik) not Google not common crawl not Facebook not Internet Archive have a copy that holds the entirety of the currently accessible web (let alone dead links). I’m often surprised using GoogleFu at how many pages I know exist even with famous authors that just don’t appear in googles index, common crawl or IA.

Re: ArchiveTeam has finished archiving all goo.gl short links

#88

Earlier quoted context omitted.

> This would mean there is an "official" source of all web data. LLM people can use snapshots of this that already exists, its called CommonCrawl: https://commoncrawl.org/

Common Crawl, while a massive dataset of the web does not represent the entirety of the web. It’s smaller than Google’s index and Google does not represent the entirety of the web either. For LLM training purposes this may or may not matter, since it does have a large amount of the web. It’s hard to prove scientifically whether the additional data would train a better model, because no one (afaik) not Google not comm…

Is there any way to find patterns in what doesn't make it into Common Crawl, and perhaps help them become more comprehensive?

Hopefully it's not people intentionally allowing the Google crawler and intentionally excluding Common Crawl with robots.txt?

Re: ArchiveTeam has finished archiving all goo.gl short links

#89
post #78

Can we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot m…

> This would mean there is an "official" source of all web data. LLM people can use snapshots of this that already exists, its called CommonCrawl: https://commoncrawl.org/

Cool! I will check it out

Re: ArchiveTeam has finished archiving all goo.gl short links

#90
post #67

Earlier quoted context omitted.

https://academictorrents.com/browse.php?search=stuck_in_the_...

Interesting. You don’t have to be an academic to access these I guess?

They have magnet links and torrent files right there on the pages, so no.
Post reply on HN