Live data from Hacker News

ArchiveTeam has finished archiving all goo.gl short links

tracker.archiveteam.org

71–80 of 112 posts

Re: ArchiveTeam has finished archiving all goo.gl short links

#71
post #69
post #66

Earlier quoted context omitted.

Then why go to the trouble of archiving them, then upload them to a public archive site, only to then keep them secret? I'm sure pastebin is filled with people's AWS credentials, too, but you don't see them randomly denying access to listings

Because then you can access the archived destination if you already know the short URL. You just can't get a full list of potentially sensitive short URL/destination pairs.

You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks?

The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time

Picking one at random, it seems the super sekrit deets you're safeguarding include buyrussia21.co.kr which, yes, is for sure very, very secret

Re: ArchiveTeam has finished archiving all goo.gl short links

#72
post #71
post #69

Earlier quoted context omitted.

Because then you can access the archived destination if you already know the short URL. You just can't get a full list of potentially sensitive short URL/destination pairs.

You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyruss…

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Re: ArchiveTeam has finished archiving all goo.gl short links

#73
post #65
post #39

Earlier quoted context omitted.

The binary units like GiB, TiB, are technically supposed to be Gibibytes and Tebibytes. Thought it was a bit silly when they first popped up but now I find them adorkably endearing, and a good way to disambiguate something that's often left vague at your expense.

In my experience, nobody actually says "Tebibytes" out loud; it's just that silly. In writing, when the precision is necessary, the abbreviation "TiB" does see some actual use.

If that's the unit, I am saying it, but yes - everyone gives me weird looks every time and just assumes I am mispronouncing terabytes but yet does not correct me.

Re: ArchiveTeam has finished archiving all goo.gl short links

#74
post #67

Earlier quoted context omitted.

Academictorrents has monthly dumps of all reddit submissions and comments even after the API restrictions.

https://academictorrents.com/browse.php?search=stuck_in_the_...

Interesting. You don’t have to be an academic to access these I guess?

Re: ArchiveTeam has finished archiving all goo.gl short links

#75

Is there anyone archiving all of reddit? Or twitter? I mean even if their terms have changed to not allow it.

ArcticShift is a project with that goal. It picks up where PushShift left off when the API changes killed that project. https://github.com/ArthurHeitmann/arctic_shift

Thanks. I wonder if anyone does this for hacker news.

Re: ArchiveTeam has finished archiving all goo.gl short links

#76
post #71

Earlier quoted context omitted.

You are aware of which thread you're discussing this in, right? The one where a bunch of like-minded souls enumerated all the address space in a few weeks? The sibling link above that queries Wayback's warc index shows at least the first several are only 6 alnum wide so it's no wonder the ArchiveTeam got them in reasonable time Picking one at random, it seems the super sekrit deets you're safeguarding include buyruss…

i asked them why they did this. the answer surprisingly is because they fear if they release the full dumps they will get blocked because of the AI scraping wars.

Feels like a bit of a kick in the teeth that I contributed towards archiving something that I don’t even get access to. What happens if they disappear? The dataset is gone forever.

Re: ArchiveTeam has finished archiving all goo.gl short links

#77

I don't understand the page, it shows a list of data sets (I think?) up to 91 TiB in size The list of short links and their target URLs can't be 91 TiB in size can it? Does anyone know how this works?

The data is saved as a WARC file, which contains the entire HTTP request and response (compressed, of course). So it's much bigger than just a short -> long URL mapping.

Re: ArchiveTeam has finished archiving all goo.gl short links

#78
Can we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot measures over time on websites, in turn making it easier to crawl. Does something like this exist? It would be so awesome. It would also allow people to run a search engine at home.

Re: ArchiveTeam has finished archiving all goo.gl short links

#79
post #78

Can we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot m…

Why would I spend time and resources to feed a machine which wastes more resources to hallucinate fiction from data it ingested?

For digital preservation? We may discuss. For an LLM? Haha, no.

No, thank you.

Re: ArchiveTeam has finished archiving all goo.gl short links

#80
post #78

Can we build a blockchain/P2P-based web crawler that can create snapshots of the entire web with high integrity (peer verification)? The already-crawled pages would be exchanged through bulk transfer between peers. This would mean there is an "official" source of all web data. LLM people can use snapshots of this. This would hopefully reduce the amount of ill-behaved crawlers, so we will see less draconian anti-bot m…

> This would mean there is an "official" source of all web data. LLM people can use snapshots of this

that already exists, its called CommonCrawl:

https://commoncrawl.org/

Post reply on HN