You're overthinking it. LLM's are massive scale regressors. Rhetorical techniques, cliches and memes (in the original Dawkins definition) spike the dataset and skew the data such that certain phrasings or rhetorical techniques will appear in the output in a similar ratio to what exists in the training set; however, the reality is that humans are able to change writing techniques depending on the nature of the conversation and the medium. LLM's don't have that, all of the scholastic papers, buzzfeed articles, and reddit posts factor into the output regardless of what it's writing. Meanwhile your average human can switch tone, style, and register depending on the situation and what they're writing.

I asked claude to take the above and rewrite it in it's own words and it came up with: "LLMs are essentially large-scale statistical regressors — they're built on huge volumes of text, so the rhetorical patterns, clichés, and memes (in the Dawkins sense of self-replicating cultural units) that saturate the training data show up in outputs at roughly the frequency they occur in that data.

The key limitation: a human writer naturally shifts register, tone, and style based on context — what they're writing, who it's for, the medium. An LLM doesn't really do that in the same way. Every academic paper, Buzzfeed listicle, and Reddit thread it was trained on bleeds into its output regardless of the actual writing task at hand, whereas a person adapts fluidly to the situation."

Even with the context, it still included an emdash, the rhetorical technique of threes, and "the key limitation". It's an inherent weakness of LLM's. There is no fixing it.