Live data from Hacker News

2.1M of the oldest Usenet posts are now online for anyone to read

vice.com

191–200 of 320 posts

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#191
post #15

Earlier quoted context omitted.

It really seems to me that archive.org, or some similar organisation, should petition Google to donate their Usenet archive, which they clearly don't care about, to them. Or possibly buy it off them. It's a long shot since it would require Google to actually do something with a dataset they clearly don't care about, but they could get some good corporate PR points almost for free.

archive.org has various usenet archives. Some of the data is there, it's just not conveniently organized or searchable or viewable. eg: https://archive.org/details/usenet

Surely someone has had a go at bringing it all back online somewhere? If not, I might have a bash at it.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#193

Earlier quoted context omitted.

I am running it through a certain set of filters. From my SEO days I recalled that new websites are often penalized based on the certain keywords in search engines. Considering this is a new site, and there is 300 million plus posts and I am not able to read and moderate it, this is the best way I know of to deal with it. But perhaps you're right and I should get rid of it. I'll think about it. This is a valid commen…

Maybe rot13 the words you think you need to censor? That'd be in keeping with the usenet tradition at least from the mid-late 90s when I was reading/posting heavily. And maybe add a simple javascript ROT13 widget so people can easily reveal it? (There was a time in my life when I could read ROT13-ed things pretty accurately in my head.)

Double rot13 just to be sure.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#194

Earlier quoted context omitted.

An important part about usegroups is that it is from an era of dial-in connections. One can dial into your ISP, sync News (fetch new messages in groups one is subscribed to and send messages) and then go offline and read/answer the messages without paying per minute and without blocking ports for other dial-in customers. This also worked between different news servers in distinct networks with unreliable connections.…

I see. It seems somewhat like mailing lists to me. I go online, my client sends and receives messages depending on which groups I'm subscribed to, and messages bounce around in a distributed fashion, somewhat like emails bouncing between SMTP servers. I still don't quite get the user experience. What if you're subscribed to alt.binaries? Do you then sync all the binaries when you go online? That wouldn't make sense.…

Your client could sync the headers and the request individual messages as needed. But yes, your newsgroup provider had to have all the messages you wanted to see.

That's one of the reasons ISPs stopped offering them - they'd have to have the infrastructure for all of their customers to support something that perhaps only one person would actually read.

Universities would often not carry the complete alt.* heirachy for example (especially in later, post 2000 years where the binary groups started getting big).

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#195
post #31

Earlier quoted context omitted.

Yes. It's outside the WWW. Originally, it was outside the Internet. The basic concept is simply copying and forwarding email-like messages, grouped into a hierarchical set of newsgroup feeds. A host would subscribe to other hosts, and would allow other hosts to subscribe to their own feeds. And then just push/pull this through the network. Originally, this was done all with shell scripts dialing up remote hosts by mo…

So basically there was a fully synced copy of the entire newsgroup contents for the groups I'm subscribed to? Like it would fetch all the new messages that arrived in alt.atheism and store it on my disk? But that surely wouldn't work on alt.binaries. Was there something like a cutoff at a certain size and you'd have to explicitly request the full content for some specific messages?

I had a summer job in 2001 which included helping the newsmaster at one of the biggest ISP at the time in Finland. If I remember correctly, there was a dedicated satellite dish to get the full newsfeed from some specialised newsfeed operator. Binary groups were popular then and this was done to avoid the huge international traffic via the normal routes.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#196

Earlier quoted context omitted.

Why are many words censored? To take a completely random example (I just took one with many asterisks) where it makes the post completely illegible: https://www.usenetarchives.com/view.php?id=soc.sexuality.gen... For an archive this is a big no-no. Respect the source material! Otherwise, thank you for the time spent doing this.

I am running it through a certain set of filters. From my SEO days I recalled that new websites are often penalized based on the certain keywords in search engines. Considering this is a new site, and there is 300 million plus posts and I am not able to read and moderate it, this is the best way I know of to deal with it. But perhaps you're right and I should get rid of it. I'll think about it. This is a valid commen…

You should definitely get rid of whatever is being used currently. The first group I randomly clicked (alt.alien.visitors) was censoring the word "public" (and "sucks" and "pipe"), multiple times in the same post which, if it happens a lot, especially on innocuous words, is really going to spoil what is an excellent project.

Its not a bad idea to filter content though, and/or have a flag button on threads/posts. 300 million articles from 40 years of an obscure and anarchic corner of the internet are bound to contain posts that are either potentially illegal or which you otherwise don't necessarily want to be publishing.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#197
post #40

Oh excellent, I can now reread those massive threads on comp.arch about how NOTHING WILL WITHSTAND THE ATTACK OF THE KILLER MICROS! For example: https://www.usenetarchives.com/view.php?id=comp.arch&g=14965...

This is the first time reading those threads for me, being the young'un that I am. That said, the some of the posts in the thread you linked seem prescient. Supercomputers these days are a pile of commodity microprocessors clustered together with high-speed interconnects, just as predicted in that thread.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#198

Earlier quoted context omitted.

Yes I'm quite aware that if governments were after me they could probably dig up this stuff but thats not my concern. I'm focused on random internet users being able to dig up 10 year old posts which would be insanely difficult given how much crap there is on the internet and how much I have changed since then.

I remember reading corporations crawling the so called "deep web" and even the "dark nets", advertising their capabilities as superior for gathering new insights/intelligence a few years ago. Was just a short blip on my radar, because I haven't seen it mentioned afterwards. Anyways, what can be done probabably will be done, for whichever reason because it's getting cheaper and cheaper. And then the next leak happens.…

This really isn't how this works.

Stylometry is a thing, sure. But it's rarely accurate enough to use on bulk archives in an automated way, especially when archives come from different domains (eg, technical vs social writing).

It's useful for identifying possible alternate authors plays attributed to Shakespeare[1], but a lot less useful for trying to tie a reddit identity to a github profile.

Anyways, what can be done probabably will be done, for whichever reason because it's getting cheaper and cheaper.

Crawling and storing is cheap, but analysis is harder. Automated spell checking ("probabably" above not withstanding!) and auto-complete means stylometry is even less useful now.

Source: myself - I work in natural language processing, and have talked to people applying these techniques.

[1] https://www.frontiersin.org/articles/10.3389/fpsyg.2018.0028...

Post reply on HN