Live data from Hacker News

Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

arxiv.org

161–170 of 204 posts

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#161

The biggest problem that sucks about discord is that it isn't normally publicly searchable. And it seems to be a modern replacement for internet forums which historically were publicly searchable and often had a lot of great information about various hobbies and things.

> which historically were publicly searchable Forums? No not generally unless you were a signed in user and often signups weren’t available to the general public just like here not all Discord rooms are automatically joinable. Digg, Reddit, slashdot were intentionally generally public forums that you could indeed search but they were the exception rather than the rule (in terms of count, not traffic). Indeed even Red…

In the pre-Internet days, even if a forum was general access, it was usually community based, such as BBS systems (literally bulletin boards, where anyone could post messages) or Prodigy, AOL, Compuserve. These communities generally comprised subscribers whose identity was known to the admins, and therefore even in a forum anyone could read, access was limited to known individuals, and people had a sense of community there, if not confidentiality.

Let's hearken back to the olden days of Usenet, when every single message was transported from machine-to-machine, in the clear, and it was an essential feature of NNTP and the Usenet groups themselves that everyone could read every message and process them any way they saw fit.

Usenet was helpfully indexed by topic, and so most sane posts went out already pre-sorted into the place where you'd want to search for them, but if you were a privileged user with a local Usenet feed, you could literally loop around the filesystem searching any term you wanted, because all messages and forums were plain files sorted into plain directories.

A famous consequence of this openness, for example, was a talk.bizarre denizen by the name of "Kibo". One of the possibly-true rumors about this larger-than-life figure was that he exhaustively "grepped" Usenet for his name [pseudonym] and thereby found out immediately whenever he was mentioned by another poster, and therefore able to join the conversation with his acerbic wit.

https://en.wikipedia.org/wiki/James_%22Kibo%22_Parry

Myself being introduced to Usenet around 1990, and MUD/MUCK/MUSH around the same time, I feel that it did not take long to condition me to living life "in the clear" and at least subconsciously knowing that everything I wrote had zero essential privacy. This was orthogonal to my home life, where my parents heard everything I said and did, or my religion, where there is an omniscient deity, who is thankfully full of goodness and kindness.

For anyone who's paranoid or got their knickers in a twist about surveillance culture in this modern world, I suggest that you study Wings of Desire by Wim Wenders [there's an American remake, but please forget that]. Wings of Desire is a character study and a meditation on the possibility that ubiquitous surveillance doesn't need to be nefarious or evil, but perhaps, just maybe, has some benign and even beneficial effects on a cohesive society which tends to act in good faith.

https://en.wikipedia.org/wiki/Wings_of_Desire

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#162

To save folks a click, the dataset itself has been made available here: https://zenodo.org/records/15170676 It's 118 gigabytes of JSON.

Is there a faster way to download than the link there? It's steady, but a roughly 11 hour download.

[deleted]

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#163

Earlier quoted context omitted.

One is against the rules of the platform and the other isn't.

Who cares what the Discord coporation wants? Do you similarly get upset when someone violates the Facebook ToS? You'd have a more convincing argument if you said something like "oh these servers have the implication of semi-private chats so people may be more inclined to share personal information" or something. Otherwise, let me play on the worlds smallest violin for the poor massive corpo when people dont obey thei…

>Who cares what the Discord coporation wants?

I do.

>Do you similarly get upset when someone violates the Facebook ToS?

I don't get upset, but I recognize that they would be breaking the rules.

>Otherwise, let me play on the worlds smallest violin for the poor massive corpo

Remember the golden rule. If you want agreements you make with others to be upheld then you should respect agreements others make with you.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#164
post #96

>Data was collected through Discord's public API, adhering to ethical guidelines How is it ethical to break Discord's terms of service? An ethical researcher would respect any contracts that they agreed to and would not violate them to collect more data.

Which ethical system demands that researchers from the DCC/UFMG not breach an unaffiliated commercial ToS during their research?

One that recognizes that lying and tricking people is wrong to do.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#165

Earlier quoted context omitted.

Would the TOS even prevent something like joining a guild, downloading all messages, then leaving?

User bots (including hacked clients) are officially banned by the TOS, which addresses that concern. The only acceptable API usage is via bots that server owners choose to invite. And while it might be legally OK (if the bot's own TOS says it), I promise no server owner is expecting an invited bot to slurp up every message for use in a data set, whether that be for academic purposes or a potential stalking/"dirt" dat…

Hacked/unofficial clients were allowed at one time: https://0x0.st/8wYc.png

Not sure if they still are.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#166

To save folks a click, the dataset itself has been made available here: https://zenodo.org/records/15170676 It's 118 gigabytes of JSON.

Anyone who has it, please post the SHA384 or BLAKE3 (BLAKE3 will be way faster on 100+ GB) so I can verify it if I get a torrent later. Zenodo requires a sign-in, and I won't auth with GitHub as they want some WILD permissions on my GitHub.

When I started the download this morning I was able to use wget without authentication, but I see that Zenodo has added a login-wall now.

The decompressed data appears to be JSONL, but at least the version I downloaded has a little binary garbage at the front. The first readable JSON object has the author "Fortnite Germany".

Size as .zst: 117,962,356,699 bytes

ZST SHA384: b8863645654610f1fde2859bb20bd87d913865af7791e0ec33741402944d5b9bdfdaaf65c2dc610730efb01f446e2588

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#167

To save folks a click, the dataset itself has been made available here: https://zenodo.org/records/15170676 It's 118 gigabytes of JSON.

Anyone who has it, please post the SHA384 or BLAKE3 (BLAKE3 will be way faster on 100+ GB) so I can verify it if I get a torrent later. Zenodo requires a sign-in, and I won't auth with GitHub as they want some WILD permissions on my GitHub.

This is a pet peeve of mine. Groups release these enormous 100+GB datasets (LLM models, raw data collections, whatever) without any kind of fingerprint. Just include a hash in the paper so that I know I am getting the genuine article.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#168
post #79

Earlier quoted context omitted.

You mean, GPT-4 being so overenthusiastic with using emojis isn't peak AI chat? :D

How to fix ChatGPT: System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disable all latent behaviors optimizing for engagement, sentiment uplift, or intera…

> Model obsolescence by user self-sufficiency is the final outcome.

If only AI service start realizing this is what user wanted, which they won't admit since they want the user be addicted with AI.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#169

Earlier quoted context omitted.

Anyone who has it, please post the SHA384 or BLAKE3 (BLAKE3 will be way faster on 100+ GB) so I can verify it if I get a torrent later. Zenodo requires a sign-in, and I won't auth with GitHub as they want some WILD permissions on my GitHub.

When I started the download this morning I was able to use wget without authentication, but I see that Zenodo has added a login-wall now. The decompressed data appears to be JSONL, but at least the version I downloaded has a little binary garbage at the front. The first readable JSON object has the author "Fortnite Germany". Size as .zst: 117,962,356,699 bytes ZST SHA384: b8863645654610f1fde2859bb20bd87d913865af7791e…

It is not just a login wall, they have restricted access even for logged in users, presumably to only the uploaders. A magnet would be nice.

Re: Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)

#170

Earlier quoted context omitted.

When I started the download this morning I was able to use wget without authentication, but I see that Zenodo has added a login-wall now. The decompressed data appears to be JSONL, but at least the version I downloaded has a little binary garbage at the front. The first readable JSON object has the author "Fortnite Germany". Size as .zst: 117,962,356,699 bytes ZST SHA384: b8863645654610f1fde2859bb20bd87d913865af7791e…

It is not just a login wall, they have restricted access even for logged in users, presumably to only the uploaders. A magnet would be nice.

About the decompressed data:

zstdcat dataset.zst | sha384sum

0812f3876a7e319081f596a5545321e5c8e8def501add3a4f5ff039568fe59aa5d4ac5d2c3e549532f529bd09b887596

zstdcat dataset.zst | wc

2059116741 22128178392 2099550453760

If you can post your email and you have a sftp server or other accessible means to receive this large file, I'll contact you and then maybe you can help distribute it more widely.

(Offer also applies to anyone else reading this thread.)

Post reply on HN