Live data from Hacker News

New accounts on HN more likely to use em-dashes

marginalia.nu

611–620 of 643 posts

Re: New accounts on HN more likely to use em-dashes

#611

Earlier quoted context omitted.

Another option instead of using identity is to use proof of work or hashcash such that anyone who thinks a comment is valuable can use some hash rate to upvote it. It doesn't matter how the content was generated, only that someone thought it was important, and you can independently verify this by checking how much hash effort went into hashing for that comment. This also does not require any identity either.

Advertisers are more willing to spend money to promote content than an individual is willing to do the same...

Using a real identity doesn't fix that problem either though: advertisers just pay real people in India to do ID checks.

Re: New accounts on HN more likely to use em-dashes

#612
post #129

Related: Show HN: Hacker News em dash user leaderboard pre-ChatGPT - https://news.ycombinator.com/item?id=45071722 - Aug 2025 (266 comments) ... which I'm proud to say originated here: https://news.ycombinator.com/item?id=45046883 .

You shared this with me via email and I had a great laugh.

I'm very disappointed to not have made the list—going to federal prison for 18 months didn't help my score.

Re: New accounts on HN more likely to use em-dashes

#613

Fwiw I did some more comparisons, looking for words disproportionately favored by noob comments: word noob new p-value ---------------------------- ai 14.93% 7.87% p=0.00016 actually 12.53% 5.34% p=1.1e-05 code 11.47% 6.04% p=0.00081 real 10.93% 2.95% p=2.6e-08 built 10.93% 2.11% p=2.1e-10 data 8.93% 3.51% p=6.1e-05 tools 7.6% 2.67% p=5.5e-05 agent 7.47% 2.95% p=0.00024 app 7.2% 3.09% p=0.00078 tool 6.8% 1.83% p=8.5e…

Worth pointing out that calculating p-values on a wide set of metrics and selecting for those under $threshold (called p-hacking) is not statistically sound - who cares, we are not an academic journal, but a pill of knowledge. The idea is, since data has a ~1/20 chance of having a p @OP have you considered calculating Cohen's effect size? p only tells us that, given the magnitude of the differences and the number of…

Yes, if OP did a full vocabulary comparison and took just those sub-threshold, it would be hacking. I'm not sure that's the case here, though? Given that (the post) OP started with em-dash, and probably didn't do repeated sampling, then it should be a pretty fair hypothesis that em-dash usage is a marker.

Your comment about pPerhaps Fisher's exact is more appropriate, on the per-word basis?

Re: New accounts on HN more likely to use em-dashes

#614

This feels like an existential threat to HN, and to the general concept of anonymous online discourse. Trust in the platform is foundational, and without it the whole thing falls down. Requiring proof of identity is the only solution I can think of, despite how unappealing it is. And even then, you'll still have people handing their account over to an LLM. I really struggle to imagine a way around it. It could be tha…

There's lots of alternatives. Others have mentioned invites and proof of work, and I'll mention a third alternative: a voucher system.

E.g. I make a new hackernews account, and say "just ask wikipedia, they will vouch for my new hackernews account". Then wikipedia checks if any of their accounts vouch for this new hackernews account. If a user with enough reputation on Wikipedia (e.g. your friends or one of your own wikipedia accounts) vouches for this new hackernews account then wikipedia tells hackernews "yes, that account is legit".

Hackernews knows the minimum amount possible about the new account. And while wikipedia knows something, they know WAY LESS than a full ID check. People can have multiple Wikipedia accounts.

And its a two way street; Wikipeida could ask hackernews about new accounts. Both sites would benefit from the collaboration.

Karma could actually become meaningful/useful for reputation checks.

The only unfortunate aspect is I'm not aware of any software tooling for such a system.

Re: New accounts on HN more likely to use em-dashes

#615
post #613

Earlier quoted context omitted.

Worth pointing out that calculating p-values on a wide set of metrics and selecting for those under $threshold (called p-hacking) is not statistically sound - who cares, we are not an academic journal, but a pill of knowledge. The idea is, since data has a ~1/20 chance of having a p @OP have you considered calculating Cohen's effect size? p only tells us that, given the magnitude of the differences and the number of…

Yes, if OP did a full vocabulary comparison and took just those sub-threshold, it would be hacking. I'm not sure that's the case here, though? Given that (the post) OP started with em-dash, and probably didn't do repeated sampling, then it should be a pretty fair hypothesis that em-dash usage is a marker. Your comment about p Perhaps Fisher's exact is more appropriate, on the per-word basis?

A Bonferroni correction would be suitable. I usually see it used in genome-wide association studies (GWAS) that check to see if a trait or phenotype is influenced by any single nucleotide polymorphisms (SNPs) in a genome. So it's doing multiple testing on a scale of ~1 million.

> One of the simplest approaches to correct for multiple testing is the Bonferroni correction. The Bonferroni correction adjusts the alpha value from α = 0.05 to α = (0.05/k) where k is the number of statistical tests conducted. For a typical GWAS using 500,000 SNPs, statistical significance of a SNP association would be set at 1e-7. This correction is the most conservative, as it assumes that each association test of the 500,000 is independent of all other tests – an assumption that is generally untrue due to linkage disequilibrium among GWAS markers.

https://journals.plos.org/ploscompbiol/article?id=10.1371/jo...

cf: https://en.wikipedia.org/wiki/Bonferroni_correction

Re: New accounts on HN more likely to use em-dashes

#616

Earlier quoted context omitted.

Another option instead of using identity is to use proof of work or hashcash such that anyone who thinks a comment is valuable can use some hash rate to upvote it. It doesn't matter how the content was generated, only that someone thought it was important, and you can independently verify this by checking how much hash effort went into hashing for that comment. This also does not require any identity either.

Advertisers are more willing to spend money to promote content than an individual is willing to do the same...

Having multiple different distribution channels can solve that problem. Advertisers cannot monopolize all distribution channels simultaneously because of the costs involved (it would be like someone trying to buy the whole economy).

Re: New accounts on HN more likely to use em-dashes

#617
post #461
post #248

Earlier quoted context omitted.

> four bland messages That's why. Boring, bland, etc. That account's M.O. is basically "write a paragraph that says nothing." Fwiw, I do think AI can be indistinguishable from dumb, boring people, but usually those kinds of people won't be on HN.

Oh we are on HN, just usually don't comment.

This made me laugh more that it should have. Thanks!

I feel like I'm certainly in that club as well.

Re: New accounts on HN more likely to use em-dashes

#618
post #11

Earlier quoted context omitted.

Karma aside, flooding the comments with a chosen narrative via army of bots seems like it's already happening. I suppose the bots can also do voting rings, but they don't necessarily need to.

> Karma aside, flooding the comments with a chosen narrative via army of bots seems like it's already happening. again with the conspiracy theories

Just because I’m crazy doesn’t mean they’re not watching me!

(Only half sarcastic)

Re: New accounts on HN more likely to use em-dashes

#619

Earlier quoted context omitted.

Interesting use of "Aunt Jemima" that nobody caught on, why did you use this particularly, afaik it doesn't exist anymore for being "racist"?

It would appear that you know wrong.

What do you mean by that? The brand was renamed several years ago:

https://www.pearlmillingcompany.com/our-history

"In June 2020, PepsiCo and The Quaker Oats Company made a commitment to change the name and image of Aunt Jemima, recognizing that they do not reflect our core values.

We want to thank everyone who has made us part of their family over the years, and look forward to starting a new chapter as the Pearl Milling Company."

Re: New accounts on HN more likely to use em-dashes

#620

Fwiw I did some more comparisons, looking for words disproportionately favored by noob comments: word noob new p-value ---------------------------- ai 14.93% 7.87% p=0.00016 actually 12.53% 5.34% p=1.1e-05 code 11.47% 6.04% p=0.00081 real 10.93% 2.95% p=2.6e-08 built 10.93% 2.11% p=2.1e-10 data 8.93% 3.51% p=6.1e-05 tools 7.6% 2.67% p=5.5e-05 agent 7.47% 2.95% p=0.00024 app 7.2% 3.09% p=0.00078 tool 6.8% 1.83% p=8.5e…

Worth pointing out that calculating p-values on a wide set of metrics and selecting for those under $threshold (called p-hacking) is not statistically sound - who cares, we are not an academic journal, but a pill of knowledge. The idea is, since data has a ~1/20 chance of having a p @OP have you considered calculating Cohen's effect size? p only tells us that, given the magnitude of the differences and the number of…

I think these term frequency comparisons are probably a pretty blunt tool, as some of the most well known AI indicators aren't words, but turns of phrase and sentence structure.

IMO a more interesting experiment would be to show comments to people (that haven't seen these conclusions), and have them assess whether they suspect them of being bots or AI authored, and then correlate that with account age.

Post reply on HN