Live data from Hacker News

I miss using em dashes

bassi.li

121–130 of 167 posts

Re: I miss using em dashes

#121
post #78
post #52

Earlier quoted context omitted.

> Eventually, as models and their users both improve, we'll collectively realize that trying to reliably discriminate between AI and human writing is no different than reading tea leaves. We should judge content based on its intrinsic value, not its provenance. There are zillions of words produced every second, your time is the most valuable resource you have, and actually existing LLM output (as opposed to some theo…

Putting aside the claim that "LLM output [...] is almost always not worth reading"[1], the whole issue here is that this supposed ability of determining whether or not content is AI-generated doesn't exist. Is it really a valuable life skill to decide whether or not you want to read something based solely on its density of em dashes? Of course there are cases where you can tell that some text is almost certainly LLM…

> this supposed ability of determining whether or not content is AI-generated doesn't exist.

It seems like you’re just wrong here? Em dashes aside, the ‘style’ of llm generated text is pretty distinct, and is something many people are able to distinguish.

Re: I miss using em dashes

#122

I'm not sure why people let others change them. I keep punctuating like it's the 20th Century.

> I'm not sure why people let others change them. I keep punctuating like it's the 20th Century. In the 20th century, there were two spaces after an end of sentence period. (I still do that.)

I think Google Docs does this automatically.

Re: I miss using em dashes

#123
post #117
post #105

Earlier quoted context omitted.

This seems like begging the question to me. Why do you think it's not a good heuristic to be able to quickly spot the tell-tale signs of LLM involvement, before you've wasted time reading slop? Yes, there will be false positives. It's a heuristic after all.

Because the false positive rate is unacceptably high — we're talking about a standard, widely used character — and because if the heuristic becomes widespread enough to matter, then it will be trivially circumvented by bad actors anyway. Who is it helping if we collectively bully ourselves into excising a perfectly good punctuation mark from human language? If anything, I'd rather that renderers like Markdown just al…

Oh no, I'm not advocating ditching em-dahses. I love them -- the form I use, anyway.

I was just curious why you've decided paying attention to them is a bad heuristic. Sure, it can change once people instruct their LLMs not to use them, but still, for now, they sure seem to overuse them!

That and "let's unpack this". I swear, I'll forbid ChatGPT from using "unpack" ever again, in any context!

Re: I miss using em dashes

#124
post #120
post #117

Earlier quoted context omitted.

Because the false positive rate is unacceptably high — we're talking about a standard, widely used character — and because if the heuristic becomes widespread enough to matter, then it will be trivially circumvented by bad actors anyway. Who is it helping if we collectively bully ourselves into excising a perfectly good punctuation mark from human language? If anything, I'd rather that renderers like Markdown just al…

> the false positive rate is unacceptably high — we're talking about a standard, widely used character Citation needed. > Who is it helping if we collectively bully ourselves into excising a perfectly good punctuation mark from human language? Humans can adapt faster than LLM companies, at least for the moment. We need to be willing to play to our strengths. Who is it helping if we bully ourselves into ignoring a sim…

Citation needed.

https://en.wikipedia.org/wiki/Dash

Humans can adapt faster than LLM companies

No one said anything about LLM companies. If I were a spammer today, I'd just have my code replace dashes in LLM output with hyphens before posting it. As a human, I'm not going to suddenly stop using dashes because a handful of people are treating a silly meme as if it were a genuinely useful heuristic.

Re: I miss using em dashes

#125
post #119
post #98

Earlier quoted context omitted.

Suggestion that humans don't use symbols like em dash is based on the possible difficulty to use them. The reason is probably because many never used / heard of the Compose key. It's not hard to do it once you learn about it.

I don't think it's because they are mechanically difficult to use as you imply, but rather that they are subtler or more nuanced symbols than most people know how to use. A lot of people are comfortable using the dot, the comma, and maybe exclamation marks. AI-speech seems to strive for more formal writing by default.

May be, but mechanical difficulty adds to less usage too.

Re: I miss using em dashes

#126
post #19

I wouldn't worry about it. A year ago the red flag du jour was "delve"; this year it's em dashes; next year it'll be something else. In any case, this is a very online topic that I assume only a vocal minority are hung up on in the first place. If you picked a random person off the street and asked for their thoughts on em dashes, you'd probably get a blank stare. Eventually, as models and their users both improve, w…

Provenance matters because LLM writing is cheap compared to actually having to think about what to say.

I only have a limited amount of time to read. Skipping someone's Internet comment because it looks like spam often means I get to engage with something else.

Re: I miss using em dashes

#127

I'm not sure why people let others change them. I keep punctuating like it's the 20th Century.

I agree. Reminds me of a few years back, when I got a Red Hat baseball cap which is (obviously) red. I had people telling me "oh no you can't wear that, MAGA hats ruined that". To which I say, balderdash and poppycock. I refuse to let others' mistaken assumptions dictate my behavior. If someone sees me wearing a RHEL hat and hates me because they assume I'm a Trump supporter, that's their problem, not mine.

Would you wear an original swastika[1] on your baseball hat?

[1]https://en.m.wikipedia.org/wiki/Swastika

Which is to say, we all compromise.

I'd hate to lose my em- and en-dashes, but the original post seems to misuse en-dashes where hyphens belong (it could just be a font issue, but no matter).

Re: I miss using em dashes

#129
post #121
post #78

Earlier quoted context omitted.

Putting aside the claim that "LLM output [...] is almost always not worth reading"[1], the whole issue here is that this supposed ability of determining whether or not content is AI-generated doesn't exist. Is it really a valuable life skill to decide whether or not you want to read something based solely on its density of em dashes? Of course there are cases where you can tell that some text is almost certainly LLM…

> this supposed ability of determining whether or not content is AI-generated doesn't exist. It seems like you’re just wrong here? Em dashes aside, the ‘style’ of llm generated text is pretty distinct, and is something many people are able to distinguish.

No, I'm not wrong. Someone could easily write in the default output style of ChatGPT by hand (which will probably become increasingly common the longer that style remains in place), and someone could easily collaborate with ChatGPT on writing that looks nothing like what you're thinking.

If organizations like schools are going to rely on tools that claim to detect AI-generated text with a useful level of reliability, they better have zero false positives. But of course they can't, because unless the tool involves time travel that isn't possible. At best, such tools can detect non-ASCII punctuation marks and overly cliched/formulaic writing, neither of which is academic dishonesty.

Re: I miss using em dashes

#130

I'm calling it now: within a year, somehow a trend will appear where writers will perform their duties under voluntary video surveillance showing the human's face, fingers and screen contents (possibly by multiple cameras) to "prove" they did it on their own. This video will be too voluminous or intrusive to be viewed manually, so it will be analyzed by (you guessed it) AI to determine if the work was authentic. It w…

Or much simpler, use the integration with TPM/DRM-style chips to monitor HID input (not too fast too) and publish that to a blockchain as PoW (new NFTs?)
Post reply on HN