Live data from Hacker News

It's Not Just X. It's Y

mail.cyberneticforests.com

31–40 of 155 posts

Re: It's Not Just X. It's Y

#31
post #7

This is how early forms of "reasoning" in LLMs worked: just literally inserting words like "Wait...", "Hmm...", "Let me reconsider...", "But is it really..." into the token stream.

Is this not how current forms of reasoning work? It seems like the open models still output things like that, and the closed ones all just summarize their thinking instead to avoid distillation, but probably do the same thing internally.

Re: It's Not Just X. It's Y

#33

I like that these AI idioms exist. They're like watermarks for text. It's worth the cost of humans avoiding them. Companies will eventually train their models to be undetectable, but society would be better if they didn't.

I agree with the feeling. But if you agree with the analysis of the article, this cat & mouse game ultimately amounts to stop disclosing our reasoning threads through commonly accepted linguistic structures. That's quite a price to pay as a society...

Re: It's Not Just X. It's Y

#34

I like that these AI idioms exist. They're like watermarks for text. It's worth the cost of humans avoiding them. Companies will eventually train their models to be undetectable, but society would be better if they didn't.

Except that the entire point of the article is that they're not AI idioms. They're not "watermarks for text." They're legitimate language constructions that LLMs tend to overuse, but that real humans also use . Real humans do, in fact, say "align with" all the time, just as often as "corresponds." And you can pry my em dashes from my cold, dead hands.

The article is not God, just because it claims something doesn't mean we have to accept it.

For better or worse (and pretty much for worse), these usages have become AI idioms. Language evolves over time, things that used to be harmless become offensive, certain terms end up taking on the complete opposite meaning than their original meaning, and we are watching certain language patterns and idioms become watermarks for AI and while it sucks, it doesn't make it false.

Re: It's Not Just X. It's Y

#35

I like that these AI idioms exist. They're like watermarks for text. It's worth the cost of humans avoiding them. Companies will eventually train their models to be undetectable, but society would be better if they didn't.

Except that the entire point of the article is that they're not AI idioms. They're not "watermarks for text." They're legitimate language constructions that LLMs tend to overuse, but that real humans also use . Real humans do, in fact, say "align with" all the time, just as often as "corresponds." And you can pry my em dashes from my cold, dead hands.

What's worse is neurodivergent writing, including my own, often resemble AI output. Now it feels like I'm having to alter my own voice in online discussions just to specifically avoid being accused of pasting an AI response.

The "AI Detection" tools employed by schools also regularly flag writing from those with Autism, ADHD, and non-native English speakers as being AI generated as well.

So, naturally, I can't stand the phrase "write like AI" when these things tend to come up because no, there are no humans that "write like AI" it's the models that have stolen the literary devices from us and now have poisoned them.

Re: It's Not Just X. It's Y

#36

nice article, but i think as a non native english speaker, i always use the model in english for reasoning and then translate the output to my language. most of these considerations do not apply. because the translation step is taking out alot of these language artifacts

Do you manually translate or translate with an LLM? While reading, I was wondering how common these kinds of written tics are in languages outside English.

Re: It's Not Just X. It's Y

#37
post #23

Earlier quoted context omitted.

My point is I don't consider em dash vs hyphen to be a strong signal either way, humans and bots alike use both interchangeably.

A signal is not the same thing as a guarantee. Both of your points so far, i.e. your provided text & that bots often bother to replace em dashes to avoid detection, actually support that it is a signal though.

The stronger signal is the grammatical structure, not the specific glyph used.

Re: It's Not Just X. It's Y

#39
post #7

This is how early forms of "reasoning" in LLMs worked: just literally inserting words like "Wait...", "Hmm...", "Let me reconsider...", "But is it really..." into the token stream.

Is this not how current forms of reasoning work? It seems like the open models still output things like that, and the closed ones all just summarize their thinking instead to avoid distillation, but probably do the same thing internally.

I think the basic idea is the same (not being a frontier lab researcher I couldn’t say for sure), but there are different techniques, such as “reasoning tokens” that aren’t literally words, and more interesting structures than just sticking them into the stream.

Re: It's Not Just X. It's Y

#40
"So, if we publicly shame people whose text looks like it might have been written by a machine – because it mimics the language used for human reasoning – and people stop writing in ways that they internalize as "AI writing" out of fear of false detection, it sends a signal that your language for reasoning must be policed, or you too could be held up to public scrutiny."

This is honestly both terrifying and well articulated.

High praise to the blog author.

Post reply on HN