Live data from Hacker News

New accounts on HN more likely to use em-dashes

marginalia.nu

411–420 of 643 posts

Re: New accounts on HN more likely to use em-dashes

#411

Earlier quoted context omitted.

I started making deliberate grammar and spelling mistakes in professional context. Not like I have a perfect writing anyway, but at least I could prove that it was self-written, not an auto-generated slop. (Could be self-written slop though :) This applies not only work-stuff itself also to the job-applications/cv/resume and cover-letters.

unrelated but I've never understood how to put a smiley at the end of parenthetical sentences (which comes up surprisingly often for me since I use smileys a lot and also like using parentheses). Just the smiley as an end parentheses (like this :) feels off but adding another parentheses (like this :) ) makes it look like it should be nested which causes problems since I also tend to nest parenthetical sentences (lik…

Post C++11 you can just do (like this:)), no extra space needed before the last parenthesis.

Re: New accounts on HN more likely to use em-dashes

#412

Fwiw I did some more comparisons, looking for words disproportionately favored by noob comments: word noob new p-value ---------------------------- ai 14.93% 7.87% p=0.00016 actually 12.53% 5.34% p=1.1e-05 code 11.47% 6.04% p=0.00081 real 10.93% 2.95% p=2.6e-08 built 10.93% 2.11% p=2.1e-10 data 8.93% 3.51% p=6.1e-05 tools 7.6% 2.67% p=5.5e-05 agent 7.47% 2.95% p=0.00024 app 7.2% 3.09% p=0.00078 tool 6.8% 1.83% p=8.5e…

Thank you marginalia_nu for article and this comment (word stats).

I got similar feeling. I'm new here, but got a feeling that some comments are like bot generated.

Such low p-values are proof that something is going on.

Hipotesis (after your recent word statistics): that some bots are "bumping up" AI related subjects. Maybe some companies using LLM tools want to promote some their products ;)

marginalia_nu respect for your work :)

Re: New accounts on HN more likely to use em-dashes

#414
post #50

I'm still salty that I can't use em-dashes anymore for fear of my writing being flagged as AI generated. Been using them for years—it's just `alt+shift+-` on a Mac keyboard and I find them more legible in many fonts compared to the simple dash on the typical numpad. It's so sad to me that good typographical conventions have been co-opted by the zeitgeist of LLMs.

> good typographical conventions

Here since 2010 in this account, I use em-dashes.

It's easy—and effective—to type using “Opt Shift -” on a Mac.

Oh yeah, left and right “curly quotes” as well, and the occasional …

> It's so sad

Don’t forget «’» — but ain’t nobody got time for that!

A few more to reclaim typography: https://howtotypeanything.com/alt-codes-on-mac/

Re: New accounts on HN more likely to use em-dashes

#415

Prior to the rise of LLM-written posts and the natural reaction of hair-trigger suspicion, I used to em and en dash fairly often in posts on HN. No reason really other than being a bit of a typography geek who happens to have always used dashes in casual writing instead of semicolons. So when I was setting up a modifier-key keyboard layer with AHK many years ago I put the em dash on modifier+dash just because I could…

Nice try

Re: New accounts on HN more likely to use em-dashes

#416

Fwiw I did some more comparisons, looking for words disproportionately favored by noob comments: word noob new p-value ---------------------------- ai 14.93% 7.87% p=0.00016 actually 12.53% 5.34% p=1.1e-05 code 11.47% 6.04% p=0.00081 real 10.93% 2.95% p=2.6e-08 built 10.93% 2.11% p=2.1e-10 data 8.93% 3.51% p=6.1e-05 tools 7.6% 2.67% p=5.5e-05 agent 7.47% 2.95% p=0.00024 app 7.2% 3.09% p=0.00078 tool 6.8% 1.83% p=8.5e…

Worth pointing out that calculating p-values on a wide set of metrics and selecting for those under $threshold (called p-hacking) is not statistically sound - who cares, we are not an academic journal, but a pill of knowledge.

The idea is, since data has a ~1/20 chance of having a p @OP have you considered calculating Cohen's effect size? p only tells us that, given the magnitude of the differences and the number of samples, we are "pretty sure" the difference is real. Cohen's `d` tells us how big the difference is on a "standard" scale.

Re: New accounts on HN more likely to use em-dashes

#417
post #325

Prior to the rise of LLM-written posts and the natural reaction of hair-trigger suspicion, I used to em and en dash fairly often in posts on HN. No reason really other than being a bit of a typography geek who happens to have always used dashes in casual writing instead of semicolons. So when I was setting up a modifier-key keyboard layer with AHK many years ago I put the em dash on modifier+dash just because I could…

I'm also increasingly aware that my own writing style and punctuation seem to line up with what might be associated with an AI, but some of the tells (em-dashes, spaces after periods, etc) seem like artifacts of when in history we learned to write. I wonder how much crossover there would be between a trained text analysis model looking for Gen-X authors and another looking for LLM's.

I worked on something like this in 2000-1. We were attempting to identify the native language and origin region of authors based on aberrant modes in second languages (as a simple case, a french person writing english might say "we are tuesday.") It was accurate and fast with the sota back then; I think you could one-shot a general purpose LLM today.

Re: New accounts on HN more likely to use em-dashes

#418

Earlier quoted context omitted.

My teenager recently asked me why I write like a chatbot, apparently unaware that some human beings prefer to write in complete sentences with attention to details like spelling, punctuation, grammar, and capitalization, and that LLMs were trained on this sort of writing. This makes me think of the fad where people on youtube will hold a microphone up in frame, because it somehow connotes authenticity. I'm sure some…

I started making deliberate grammar and spelling mistakes in professional context. Not like I have a perfect writing anyway, but at least I could prove that it was self-written, not an auto-generated slop. (Could be self-written slop though :) This applies not only work-stuff itself also to the job-applications/cv/resume and cover-letters.

I appreciate you including a few minor mistakes in this very post:

> I started making deliberate grammar and spelling mistakes in professional context[s]. Not like I have ~a~ perfect writing anyway, but at least I could prove that it was self-written, not an auto-generated slop. (Could be self-written slop though :)

> This applies not only [to] work-stuff itself also to the job-applications/cv/resume and cover-letters.

I conclude you are real.

Re: New accounts on HN more likely to use em-dashes

#419

Prior to the rise of LLM-written posts and the natural reaction of hair-trigger suspicion, I used to em and en dash fairly often in posts on HN. No reason really other than being a bit of a typography geek who happens to have always used dashes in casual writing instead of semicolons. So when I was setting up a modifier-key keyboard layer with AHK many years ago I put the em dash on modifier+dash just because I could…

I also used em-dash before LLMs, though I would not call myself a typography geek. But yesterday I wrote a birthday message to someone and replaced my em-dashes with minus signs, because I did not want them to think that my message is LLM generated..

Re: New accounts on HN more likely to use em-dashes

#420

Earlier quoted context omitted.

My teenager recently asked me why I write like a chatbot, apparently unaware that some human beings prefer to write in complete sentences with attention to details like spelling, punctuation, grammar, and capitalization, and that LLMs were trained on this sort of writing. This makes me think of the fad where people on youtube will hold a microphone up in frame, because it somehow connotes authenticity. I'm sure some…

I started making deliberate grammar and spelling mistakes in professional context. Not like I have a perfect writing anyway, but at least I could prove that it was self-written, not an auto-generated slop. (Could be self-written slop though :) This applies not only work-stuff itself also to the job-applications/cv/resume and cover-letters.

This only works as "proof" up until someone innovates an "authenticity" flag on the LLM output.
Post reply on HN