Live data from Hacker News

We need to tell people ChatGPT will lie to them, not debate linguistics

simonwillison.net

351–360 of 485 posts

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#351
post #347

Earlier quoted context omitted.

Many gendered languages use male gender as the default when it's ambiguous or unknown, so that would be expected behavior. And yes, it's absolutely possible to make it switch to whatever gender you want, and generally to pretend to be whatever you want it to be. The identity comes from the conversation context (including all the hidden messages), not from the LLM.

It's the default when speaking about groups of people or people whose identity isn't known. But if someone speaks about themselves then they have to choose the right one. This is also something that language learners have to pay attention to - to not sound weird.

The right one in this case is "doesn't have a gender", so what does that correspond to in your language? In mine, that would be neuter, and GPT-4 seems to prefer that.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#352

Earlier quoted context omitted.

Better than getting extincted, I suppose.

Oh man, I am going to be a huge nerd now. Pre-Brian Herbert and Kevin Anderson, the Butlerian Jihad came about because people became lazy under AI, lazy of mind , and were eventually enslaved by those who controlled the AI. Not much more was said about it. One could have, yes, militant robots, Exterminate! Exterminate! out of it, or you could posit a more Huxley-like dystopia, one of convenience. Control the AI, cont…

> Pre-Brian Herbert and Kevin Anderson, the Butlerian Jihad came about because people became lazy under AI, lazy of mind, and were eventually enslaved by those who controlled the AI.

Oh I love this, basically Wall-E is a better Dune than the latest books? Genius!

Anyway, I read Dune many years ago and don't remember this, is this explained in the novels or is it coming from some other sources?

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#353

Earlier quoted context omitted.

I think good annotation is much harder than obtaining data. But now Reddit, Twitter and Quora will realise what kind goldmine their data is they might close easy access to it.

None of the GPT models rely on annotated or classified data. It's unsupervised.

The final training steps of gpt 3.5 did, and we assume the same for 4. The RLHF step.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#354
post #108
post #69

Earlier quoted context omitted.

When it speaks my language - it uses masculine gender when speaking about itself.

That's interesting, what language is that? Also, is it possible to give a prompt to make ChatGPT switch to feminine gender?

It's Latvian. But, I guess it acts the same when speaking other gendered languages.

Both "chatbot" and "large language model" are masculine - so this is why it picked the masculine gender, I guess.

I was asking it about its training dataset and when it said "I've been trained on..." it picked the masculine form of the word "trained". This is why machine translation is a hard problem to solve. Things like this can easily get lost in translation.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#355

Earlier quoted context omitted.

You could reasonably describe it as "human language emulator" back when people were using GPT-2 and the likes to compose text. But what we have today doesn't just emulate human language - it accepts tasks in that language, including such tasks that require reasoning to perform, and then carries them out. Granted, the only thing it can really "do" is produce text, but that already covers quite a lot of tasks - and the…

Interesting perspective. I'm still learning about what it really is, and I'm having trouble marrying the thoughts of a parent commenter with yours: > ... does what it is engineered to do pretty well, which is, generate text that is representative of its training data following on from input tokens. It can't reason ... versus > ... doesn't just emulate human language - it accepts tasks in that language, including such…

Nobody can definitely answer this question because we don't know what exactly is going on inside the model of that size. We can only speculate based on the observed behavior.

But in this case, I didn't imply that it's "reasoning beyond the domain of language", in a sense that language is exactly what it uses to reason. If you force it to perform tasks without intermediate or final outputs that are meaningful text, the result is far worse. Conversely, if you tell it to "think out loud", the results are significantly better for most tasks. Here's one example from GPT-4 where the "thinking" effectively becomes a self-prompt for the corresponding SQL query: https://gist.github.com/int19h/4f5b98bcb9fab124d308efc19e530....

Or here's an even more interesting example where GPT-4 does this kind of "thinking out loud" unprompted: https://gist.github.com/int19h/8251bd00b7a4858a69cf3922ae674...

I think the real point of disagreement is whether this constitutes actual reasoning or "merely completing tokens". If you showed the transcript of a chat with GPT-4 solving a multi-step task to a random person off the street, I have no doubt that they'd describe it as reasoning. Beyond that, one can pick the definition of "reason" that best fits their interpretation - there is no shortage of them, just as there is no shortage of definitions for "intelligence", "consciousness" etc.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#356
post #347

Earlier quoted context omitted.

It's the default when speaking about groups of people or people whose identity isn't known. But if someone speaks about themselves then they have to choose the right one. This is also something that language learners have to pay attention to - to not sound weird.

The right one in this case is "doesn't have a gender", so what does that correspond to in your language? In mine, that would be neuter, and GPT-4 seems to prefer that.

My language has only masculine and feminine genders. There's no neuter gender. And there isn't anything like "singular they" either.

"Chatbot" and "language model" are both masculine so, I guess that's why ChatGPT uses "masculine it" when speaking about itself.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#357
post #95
post #67

Earlier quoted context omitted.

Nope - it speaks as a man would speak in my language.

How do non-binary persons speak differently in your language?

They have to pick either masculine or feminine gender. There's no other option.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#358

The part that's concerning about ChatGPT is that a computer program that is "confidently wrong" is basically indistinguishable from what dumb people think smart people are like. This means people are going to believe ChatGPT's lies unless they are repeatedly told not to trust it just like they believe the lies of individuals whose intelligence is roughly equivalent to ChatGPT's. Based on my understanding of the appro…

Your characterization of “dumb people” as somehow being more prone to misinformation is inaccurate and disrespectful. Highly intelligent people are as prone to irrational thinking, and some research suggests even more prone. Go look at some of the most awful personalities on TV or in history, often they are quite intelligent. If you want to school yourself on just how dumb smart people are I suggest going through the back catalog of the “you are not so smart” podcast.

Based on my understanding of the approach behind ChatGPT, it is probably very close to a local maximum in terms of intelligence so we don't have to worry about the fearmongering spread by the "AI safety" people any time soon if AI research continues to follow this paradigm.

ChatGPT is extremely poorly understood. People see it as a text completion engine but with the size of the model and the depth it has it is more accurate in my understanding to see it as a pattern combination and completion engine. The fascinating part is that the human brain is exclusively about patterns, combining and completing them, and those patterns are transferred between generations through language (sight or hearing not required). GPT acquires its patterns in a similar way. A GPT approach may therefore in theory be able to capture all the patterns a human mind can. And maybe not, but I get the impression nobody knows. Yet plenty of smart people have no problem making confident statements either way, which ties back to the beginning of this comment and ironically is exactly what GPT is accused of.

Is GPT4 at its ceiling of capability, or is it a path to AGI? I don’t know, and I believe nobody can know. After all, nobody truly understands how these models do what they do, not really. The precautionary principle therefore should apply and we should be wary of training these models further.

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#359
post #48
post #41

Lying is an intentional act where misleading is the purpose. LLMs don’t have a purpose. They are methods. Saying LLMs lie is like saying a recipe will give you food poisoning. It misses a step (cooking in the analogy). That step is part of critical thinking. A person using an LLM’s output is the one who lies. The LLM “generates falsehoods”, “creates false responses”, “models inaccurate details”, “writes authoritative…

Everything you said here is true. I still don't think this is the right way to explain it to the wider public. We have an epidemic of misunderstanding right now: people are being exposed to ChatGPT with no guidance at all, so they start using it, it answers their questions convincingly, they form a mental model that it's an infallible "AI" and quickly start falling into traps. I want them to understand that it can't…

My problem is that in communicating that LLMs can’t be trusted by stating that they are lying we introduce a partial falsehood of our own. I agree to some extent that we need a snappy way to communicate this (and that the proposed wording I’ve given probably doesn’t get there fully yet comparative to “lying”. “Lying” definitely is a good description when analyzed through the lens of simplicity, but I hope it’s not the best we can do. Hallucinating is worse (as you noted in the article.

I like the “epidemic of misunderstanding” idea, but I’m reminded of that Bezos quote about being misunderstood. Perhaps we need to apply some of that approach here.

> JEFF BEZOS: Well, I would say one thing that I have learned within the first couple of years of starting the company is that in inventing and pioneering requires a willingness to be misunderstood for long periods of time.

https://hbr.org/podcast/2013/01/jeff-bezos-on-leading-for-th... (Also other similar quotes in other interviews)

Re: We need to tell people ChatGPT will lie to them, not debate linguistics

#360

The part that's concerning about ChatGPT is that a computer program that is "confidently wrong" is basically indistinguishable from what dumb people think smart people are like. This means people are going to believe ChatGPT's lies unless they are repeatedly told not to trust it just like they believe the lies of individuals whose intelligence is roughly equivalent to ChatGPT's. Based on my understanding of the appro…

It looks like it'll be competition for phony experts and politicians. I think it is easier to improve AI algorithms than deal human with human liars.
Post reply on HN