Live data from Hacker News

I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

marcusolang.substack.com

271–280 of 533 posts

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#271

Earlier quoted context omitted.

Each one of these has slightly different readings in my eyes.

Unlike the last variant, the first two imply there was some quantity of work and it was all completed. I don't really see the difference between the two though.

Well, option 1 implies that there was something else going on before the event described in the sentence. Option 2 is neutral about that.

Compare:

1. I did the work for that last week.

2. I proceeded to do the work for that last week.

Sentence 2 strikes me as questionably grammatical. It needs to be proceeding from something in the context.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#272

Always interesting (in an informative way) to see people "defending" em-dashes from my personal perspective. Before you get mad, let me explain: before ChatGPT, I only ever saw em-dashes when MS Word would sometimes turn a dash into a "longer dash" as I always thought of it. I have NEVER typed an em-dash, and I don't know how to do it on Windows or Android. I actually remember having issues with running a program tha…

I actually checked HN's comment data corpus to see if em dash usage rose after AI adoption became more widespread. I was kind of shocked to see that it did not.

Its overuse is definitely a marker of either AI or a poorly written body of text. In my opinion, if you have to rely on excessive parentheticals, then you are usually off restructuring your sentences to flow more clearly.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#273
post #25

A lot of training data was curated in Kenya[0]. I would imagine if LLM data was curated in Japan our LLMs would sound a lot like the authors of their most popular English text books. Maybe other common Japanese idioms would leak in to the training data, like "ね" or "でしょう", ChatGPT would say "Don't you agree?" at the end of every message. [0] https://www.theverge.com/features/23764584/ai-artificial-int...

This is a wild misunderstanding of LLMs. Data labeling has nothing to do with generating the astronomical text corpus used to train modern LLMs.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#274
post #80

> There were unspoken rules, commandments passed down from teacher to student, year after year. The first commandment? Thou shalt begin with a proverb or a powerful opening statement. “Haste makes waste,” we would write, before launching into a tale about rushing to the market and forgetting the money. The second? Thou shalt demonstrate a wide vocabulary. You didn’t just ‘walk’; you ‘strode purposefully’, ‘trudged we…

Hemingway was still a master of word choice. I recall an entire class spent on a few lines that conveyed a sense of heaviness to the scene. 'Plodding' was given a lot of attention.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#275
post #269

Earlier quoted context omitted.

Right, it is currently incapable of providing a straight answer without clearing it's throat selling the answer. It reminds me of those recipe blogs that just can't get to the fucking recipe. It's bad writing! But it's not bad technically, in a style-guide kind of way.

Sometimes I wonder if the throat-clearing is an indispensable part of getting to the "good bits" that follow. Like, do those extra tokens give it more "room to think" even if they're basically meaningless in themselves?

Isn’t that the point of the hidden chain of thought tokens, rather than the visible cruft?

I think the fluff, the emojis, the sycophancy is all symptomatic of the training process and human feedback.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#276
> I don't write like ChatGPT. ChatGPT, in its strange, disembodied, globally-sourced way, writes like me.

We will all soon write and talk like ChatGPT. Kids growing up asking ChatGPT for homework help, people use it for therapy, to resumes, for CVs, for their imaginary romantic "friends", asking every day questions from the search engine they'll get some LLM response. After some time you'll find yourself chatting with your relative or a coworker over coffee and instead of hearing, "lol, Jim, that's bullshit" you'll hear something like "you're absolutely right, here let me show you a bulleted list why this is the case...". Even more scarier, you'll soon hear yourself say that to someone, as well.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#277

Earlier quoted context omitted.

Obsession with short sentences and generally pushing extreme simplicity of structure and word choice has been terrible for English prose. It’s not been terrible because most people aren’t aided by such guidance (most are) but because the same people who can’t be trusted to wield a quill without the bumper-lanes installed see a sentence longer than ten words, or a semicolon, or god forbid literate and appropriate nuan…

"He walked up to Helen and asked, 'What are you doing?'" "He strode up to Helen and asked, 'What are you doing?'" "He sidled up to Helen and asked, 'What are you doing?'" "He tromped up to Helen and asked, 'What are you doing?'" Each of those sentences conveys as slightly different action. You can almost imagine the person's face has a different expression in each version. Yes, I hate it when amateurs just search/rep…

You forgot:

"He waddled up to Helen and asked, 'What are you doing?'"

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#278
post #80

> There were unspoken rules, commandments passed down from teacher to student, year after year. The first commandment? Thou shalt begin with a proverb or a powerful opening statement. “Haste makes waste,” we would write, before launching into a tale about rushing to the market and forgetting the money. The second? Thou shalt demonstrate a wide vocabulary. You didn’t just ‘walk’; you ‘strode purposefully’, ‘trudged we…

Our legal systems are based around being concise and succinct, relevant, and objectively unbiased.

I was raised to be respectful by "getting to the point, afap" to avoid wasting anybody's time.

But I've noticed that mostly only the members of the science and legal community exercising similar principles.

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#279
post #277

Earlier quoted context omitted.

"He walked up to Helen and asked, 'What are you doing?'" "He strode up to Helen and asked, 'What are you doing?'" "He sidled up to Helen and asked, 'What are you doing?'" "He tromped up to Helen and asked, 'What are you doing?'" Each of those sentences conveys as slightly different action. You can almost imagine the person's face has a different expression in each version. Yes, I hate it when amateurs just search/rep…

You forgot: "He waddled up to Helen and asked, 'What are you doing?'"

"He scrambled up to Helen and asked, 'What are you doing?'"

"He kick-flipped up to Helen and asked, 'What are you doing?'"

[edit] electric-slid! Pirouetted! Somersaulted!

Re: I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

#280
post #244

Earlier quoted context omitted.

> MASSIVE EM DASH Tangent: the thing I find most annoying about ChatGPT's use of em-dashes is that it never even uses them for the one thing they're best suited for. ChatGPT's em-dashes could almost always be replaced with a colon or a comma. But the true non-redundant-syntax use of em-dashes in English prose, is in the embedding into a sentence of self-interruptive 'joiner' sub-sentences that can themselves bear pun…

> These things are spoken entirely differently than — and on the page, they read entirely differently to — regular parenthetical-bearing sentences. > No, seriously, compare/contrast: "these things are spoken entirely differently than (and on the page, they read entirely differently to) regular parenthetical-bearing sentences." Those are spoken the same way, they read the same way, and they mean the same thing.

Is this maybe a thing like how only designers are aware of kerning? These read / sound very different to me, and to everyone I've brought up the subject with (who admittedly are in a certain bubble of people who either write professionally, or "do things" with their voices, or both.)

• The length of the verbal pause is different. (It's hard to quantify this, as it's relative to your speaking rate, which can fluctuate even within a sentence. But I can maybe describe it in terms of meter in poetry/songwriting: when allowed to, a parenthetical pause may be read to act as a one-syllable rest in the meter of a poem, often helpfully shifting the words in the parenthetical over to properly end-align a pair of rhyming [but otherwise misaligned] feet. An em-dash, on the other hand, acts as only a half-syllable rest; it therefore offsets the meter of the words in the subclause that follow, until the closing em-dash adds another half-syllable rest to set things right. This is in part why ChatGPT's favored sentences, consisting of "peer" clauses joined by a single em-dash, are somewhat grating to mentally read aloud; you end up "off" by a half-syllable after them, unless you can read ahead far enough to notice that there's no closing em-dash in the sentence, and so allow the em-dash-length pause to read as a semicolon-length pause instead.)

• The voicing of the last word before the opening parenthesis / first em-dash starts is different. (paren = slow down for last few words before the paren, then suddenly speed up, and override the word's normal tonal emphasis with a last-syllable-emphasized rising tone + de-voicing of vowels; em-dash = slow down and over-enunciate last few words before the em-dash, then read the last syllable before the em-dash louder with a overridden falling voiced tone)

• The speed at which, and vocal register with which, the aside / subclause is read is different. (parens = lowest register you can comfortably speak at, slightly quieter, slightly faster than you were delivering the toplevel sentence; em-dashes = delivery same speed or slower, first few syllables given overridden voiced emphasis with rising tone from low to normal, and last few syllables given overridden voiced emphasis with falling tone from normal to low)

• The voicing of the first words after the subclause ends is different. (closing paren = resume speaking precisely as if the parenthetical didn't happen; second em-dash = give a fast, flat-low nasally voiced performance of the first one or two syllables after the em-dash.)

To describe the overall effect of these tweaks:

A parenthetical should be heard as if embedded into the sentence very deliberately, but delivered as an aside / tangent, smaller and off-to-the-side, almost an "inlined footnote", trying to not distract from the point, nor to "blow the listener's stack" by losing the thread of the toplevel point in considering it.

An em-dash-enclosed interruptive subclause should read like the speaker has realized at the last moment that they have two related points to make; that they are seemingly proceeding, after a stutter, to finish the sentence with the subclause; but that they are then "backing up" and finishing the same sentence again with the toplevel clause. The verbalization should be able to be visualized as the outer sentence being "squashed in" to "make room" for the interruptive subclause; and the interruptive subclause "squashing at the edges" [tonally up or down, though usually down] to indicate its own "squeezed in" beginning and end edges.

Note that this isn't subjective/anecdotal descriptions from how I speak myself. These are actually my attempt to distill vocal coaching guidelines I've learned for:

• live sight-reading of teleprompter lines containing these elements, as a TV show host / news anchor

• default-assumed directorial expectations for lines containing elements like these, when giving screenplay readings as a [voice] actor (before any directorial "notes" come into play)

Post reply on HN