Live data from Hacker News

2.1M of the oldest Usenet posts are now online for anyone to read

vice.com

101–110 of 320 posts

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#102
post #12

I'm a bit surprised that a zip archive of only 2.1 million posts was almost 2GB. I wonder what a treemap of group size looks like.

Only 2GB means they aren't including alt.binaries.

So none of the uuencoded porn?

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#103
post #35

Can someone explain Usenet to me? As I understand it, it's not a website. Is it something outside of the WWW? Does it have its own protocol or does it use HTTP? Do you need a dedicated client software for it, similar to IRC? How do/did you register to it? Is it paid? Who hosts it?

Think something like Reddit + Bittorrent. Slight difference -- there would be no main "front page," -- only individual self-hosted hubs that could choose which "subreddits" (newsgroups) they would carry.

More like Reddit's database + Bittorrent. It was up to your Usenet client software to provide an interface on top of that, and there were several independently developed but fully developed clients to choose from.

That meant that unlike with Reddit, an idiotic interface redesign only ruined it for users of one client, and only until they got time to switch to another client.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#104
post #67

It's really striking the relatively high quality of post content. Jumping into a group related to an interest of mine all the discussions are very engaging. Assuming there's little in the way of moderation perhaps there's something to be said for small networks with a barrier to entry.

When Usenet originated, you needed to have an Internet-connected computer. In the late 80s, this was largely limited to university campuses, which meant that the process of educating new members on the social norms already in place largely only needed to happen for a few weeks in September. In September 1993, AOL offered Usenet access to any AOL subscriber. This gave rise to the term Eternal September.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#106

Earlier quoted context omitted.

I'm Gen Z too and for my whole life the advice has been "Don't put your real info on the internet". The only thing that comes up when you search my name is my gitlab profile and a technical blog. All of my cringy forum posts still exist but they are all under random usernames and would be very difficult to link back to me.

https://en.wikipedia.org/wiki/Stylometry

For any who are interested in witnessing a fictionalized example of applied Stylometry, I highly suggest the Star Wars Thrawn series by Timothy Zahn. Thrawn's specialty is making tactical inferences based on the physical art pieces created by a culture. It's a really interesting read, but don't take my word for it!

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#107
For you youngun's if you wanted "pictures" you'd have to go to alt.porn.*, choose an interesting title, get a bunch of posts with text data (wait for them to download), cut and paste them into a single text file then uudecode it to a JPEG of something that hopefully will match the title.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#108
post #37
post #27

Earlier quoted context omitted.

When you factor in that ZIP compresses by about half, that averages to around 2K per post -- one screenful of text on an 80x25 PC display -- which is unsurprising. It's probably distributed in terms of a few long, informative posts and a long tail of "AOL!" type posts, but still.

Well I just noticed that fully one percent of these posts are in alt.atheism. I wonder how much this influence by UI design and lexical ordering alone. I bet alt.atheism was always on the first page when you just wanted to flip through groups.

Actually, I doubt there was much influence there.

I hung out in talk.origins for a little bit. The usual practice of threads there was to have very, very deep threads with people quoting liberally (i.e., not bothering to trim messages, which is usual custom on Usenet) to reply with a short message. And talk.origins was consistently one of the highest post count groups. If alt.atheism was like talk.origins, then it's mostly the crowd of people it attracted rather than any accident of UI design.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#109
I had a brief occasion to use NNTP twelve or thirteen years ago, where my university was still using NNTP for course discussion boards.

It was awesome. You had a good choice of clients to read and post (instead of a half-baked webapp), and everything was incredibly snappy.

We have fallen so far. There's no money in building protocols and money in apps, so people build apps instead.

Re: 2.1M of the oldest Usenet posts are now online for anyone to read

#110

Folks, I am the guy behind this project. A friend of mine mentioned he saw the site mentioned on hacker news, so I came to check it out. If you have any questions for me, don't hesitate to ask, as time permits (and two little boys), I'll do my best to answer them.

Why are many words censored? To take a completely random example (I just took one with many asterisks) where it makes the post completely illegible: https://www.usenetarchives.com/view.php?id=soc.sexuality.gen... For an archive this is a big no-no. Respect the source material! Otherwise, thank you for the time spent doing this.

> For an archive this is a big no-no. Respect the source material!

Though I'm also curious, that's perhaps not the tone I would have used when asking. After all, better a censored archive than no archive.

I'm just speculating, but it may be the policy of usenetarchives.com, in order to accept their upload.

Censoring seems to be done around email addresses, names, and offensive words. Perhaps, this is done to reduce the chances of people later asking for the posts to be taken down entirely.

For example, I believe the takedown by Google of comp.lang.lisp and comp.lang.forth commented elsewhere was done because there was offensive content present. The Google support request that mentioned that reason was taken down, but it's what I remember.

Post reply on HN