Live data from Hacker News

Lawyer cites fake cases invented by ChatGPT, judge is not amused

simonwillison.net

141–150 of 319 posts

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#141

Earlier quoted context omitted.

Whether a statement is true or false doesn’t depend on the mechanism generating the statement. We should hold these models (or more realistically, their creators) to the same standard as humans. What do we do with a human that generates plausible-sounding sentences without regard for their truth? Let’s hold the creators of these models accountable, and everything will be better.

> What do we do with a human that generates plausible-sounding sentences without regard for their truth? Elect them as leaders?

ChatGPT is perfect for generating company mission statements, political rhetoric, and other forms of BS.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#142
post #7

No, it did not “double-check”—that’s not something it can do! And stating that the cases “can be found on legal research databases” is a flat out lie. What’s harder is explaining why ChatGPT would lie in this way. What possible reason could LLM companies have for shipping a model that does this? It did this because it's copying how humans talk, not what humans do. Humans say "I double checked" when asked to verify so…

ChatGPT did not lie; it cannot lie. It was given a sequence of words and tasked with producing a subsequent sequence of words that satisfy with high probability the constraints of the model. It did that admirably. It's not its fault, or in my opinion OpenAI's fault, that the output is being misunderstood and misused by people who can't be bothered understanding it and project their own ideas of how it should function…

This harks back to around 1999 when people would often blame computers for mistakes in their math, documents, reports, sworn filings, and so on. Then, a thousand different permutations of "computers don't make mistakes" or "computers are never wrong" became popular sayings.

Large Language Models (LLMs) are never wrong, and they do not make mistakes. They are not fact machines. Their purpose is to abstract knowledge and to produce plausible language.

GPT-4 is actually quite good at handling facts, yet it still hallucinates facts that are not common knowledge, such as legal ones. GPT-3.5, the original ChatGPT and the non-premium version, is less effective with even slightly obscure facts, like determining if a renowned person is a member of a particular organization.

This is why we can't always have nice things. This is why AI must be carefully aligned to make it safe. Sooner or later, a lawyer might consider the plausible language produced by LLMs to be factual. Then, a politician might do the same, followed by a teacher, a therapist, a historian, or even a doctor. I thought the warnings about its tendency to hallucinate speech were clear — those warnings displayed the first time you open ChatGPT. To most people, I believe they were.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#143

Earlier quoted context omitted.

ChatGPT did not lie; it cannot lie. It was given a sequence of words and tasked with producing a subsequent sequence of words that satisfy with high probability the constraints of the model. It did that admirably. It's not its fault, or in my opinion OpenAI's fault, that the output is being misunderstood and misused by people who can't be bothered understanding it and project their own ideas of how it should function…

Whether a statement is true or false doesn’t depend on the mechanism generating the statement. We should hold these models (or more realistically, their creators) to the same standard as humans. What do we do with a human that generates plausible-sounding sentences without regard for their truth? Let’s hold the creators of these models accountable, and everything will be better.

That standard is completely impossible to reach based on the way these models function. They’re algorithms predicting words.

We treat people and organizations who gather data and try to make accurate predictions with extremely high leniency. It’s common sense not to expect omnipotence.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#144
post #7

No, it did not “double-check”—that’s not something it can do! And stating that the cases “can be found on legal research databases” is a flat out lie. What’s harder is explaining why ChatGPT would lie in this way. What possible reason could LLM companies have for shipping a model that does this? It did this because it's copying how humans talk, not what humans do. Humans say "I double checked" when asked to verify so…

GPT4 can double-check to an extent. I gave it a sequence of 67 letter As and asked it to count them. It said "100", I said "recount": 98, recount, 69, recount, 67, recount, 67, recount, 67, recount, 67. It converged to the correct count and stayed there. This is quite a different scenario though, tangential to your [correct] point.

But would GPT4 actually check something it had not checked the first time? Remember, telling the truth is not a consideration for it (and probably isn't even modeled), just saying something that would typically be said in similar circumstances.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#145

Earlier quoted context omitted.

Whether a statement is true or false doesn’t depend on the mechanism generating the statement. We should hold these models (or more realistically, their creators) to the same standard as humans. What do we do with a human that generates plausible-sounding sentences without regard for their truth? Let’s hold the creators of these models accountable, and everything will be better.

>Let’s hold the creators of these models accountable, and everything will be better. Shall we hold Adobe responsible for people photoshopping their ex's face into porn as well?

I don’t think the marketing around photoshop and chatgpt are similar.

And that matters. Just like with self-driving cars, as soon as we hold the companies accountable to their claims and marketing, they start bringing the hidden footnotes to the fore.

Tesla’s FSD then suddenly becomes a level 2 ADAS as admitted by the company lawyers. ChatGPT becomes a fiction generator with some resemblance to reality. Then I think we’ll all be better off.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#146

Earlier quoted context omitted.

https://news.ycombinator.com/item?id=35963163 ( "Texas professor fails entire class from graduating- claiming they used ChatGTP [sic]", 277 comments) https://news.ycombinator.com/item?id=35980121 ( "Texas professor failed half of class after ChatGPT claimed it wrote their papers ", 22 comments)

My main takeaway is that failing the second half of the class and misspelling ChatGPT leads to > 10x engagement.

My main takeway is that the guy who registers chatgtp.com is going to make a lot of money by providing bogus answers to frivolous questions :-)

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#147

This is why it is very important to have the prompts fill in relevant fragments from a quality corpus. That people think these models “tell the truth” or “hallucinate” is only half the story. It’s like expecting your language center to know all the facts your visual consciousness contains, or your visual consciousness to be able to talk in full sentences. It’s only when all models are working well together the truth…

> That people think these models “tell the truth” or “hallucinate” is only half the story.

A meta-problem here is in choosing to use descriptive phrases like tell the truth and hallucinate, which are human conditions that further anthropomorphize technology with no agency, making it more difficult for layman society to defend against its inherent fallibility.

  UX = P_Success*Benefit - P_Failure*Cost
It's been well over a decade since I learned of this deviously simple relationship from UX expert Johnny Lee, and yet with every new generation of tech that has hit the market since, it's never surprising how the hype cycle results in a brazen dismissal of the latter half.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#148
post #14

Earlier quoted context omitted.

Yeah, that was my conclusion too: What’s a common response to the question “are you sure you are right?”—it’s “yes, I double-checked”. I bet GPT-3’s training data has huge numbers of examples of dialogue like this.

They should RLHF this behaviour out. Asking people to be aware of limitations is in similar vein as asking them to read ToC

If the model could tell when it was wrong it would be GPT-6 or 7. I think the best 4 could do is maybe it can detect when things enter the realm of the factual or mathematical etc and use a external service for that part

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#149

Earlier quoted context omitted.

Lying implies intent, and knowing what the truth is. Saying something you believe to be true, but is wrong, is generally not considered a lie but a mistake. A better description of what ChatGPT does is described well by one definition of bullshit: > bullshit is speech intended to persuade without regard for truth. The liar cares about the truth and attempts to hide it; the bullshitter doesn't care if what they say is…

> Lying implies intent, and knowing what the truth is. Saying something you believe to be true, but is wrong, is generally not considered a lie but a mistake. Those are the semantics of lying. But "X like a duck" is about ignoring semantics, and focusing not on intent or any other subtletly, but only on the outward results (whether something has the external trappings of a duck). So, if it produces things that look l…

A person who is mistaken looks like they're lying. That doesn't mean they're actually lying.

That's the thing people are trying to point out. You can't look at something that looks like it's lying and conclude that it's lying, because intent is an intrinsic part of what it means to lie.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#150

Hilarious. It’s important to remember: 1) ChatGPT is not a research tool 2) It sort of resembles one and will absolutely act like one if you ask it to, and it it may even produce useful results! But… 3) You have to independently verify any factual statement it makes and also 4) In my experience the longer the chat session, the more likely it is to hallucinate, reiterate, and double down on previous output

This is completely true but completely in conflict with how many very large companies advertise it. I’m a paid GitHub Copilot user and recently started using their chat tool. It lies constantly and convincingly, so often that I’m starting to wonder if it wastes more time than it saves. It’s simply not capable of reliably doing its job. This is on a “Tesla autopilot” level of misrepresenting a product but on a larger…

Where does Github misrepresent their Chat beta? On their marketing website?
Post reply on HN