Live data from Hacker News

Lawyer cites fake cases invented by ChatGPT, judge is not amused

simonwillison.net

281–290 of 319 posts

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#281
post #142

Earlier quoted context omitted.

This harks back to around 1999 when people would often blame computers for mistakes in their math, documents, reports, sworn filings, and so on. Then, a thousand different permutations of "computers don't make mistakes" or "computers are never wrong" became popular sayings. Large Language Models (LLMs) are never wrong, and they do not make mistakes. They are not fact machines. Their purpose is to abstract knowledge a…

Doctors, lawyers, historians, and anyone else shouldn't use chatgpt for their work.

Why not? They should use it, with sufficient understanding of what it is. Doctors should not use it to diagnose a patient, but could use it to get some additional ideas for a list of symptoms. Lawyers should obviously not write court documents with it or cite it in court, but they could use it to get some ideas for case law. It's a hallucinating idea generator.

I write very technical articles and use GPT-4 for "fact-checking". It's not perfect, but as a domain expert of what I write, I can sift out what it gets wrong, and still benefit from what it gets right. It has both - suggested some ridiculous edits to my articles, and found some very difficult to spot mistakes, like where a reader might misinterpret something from my language. And that is tremendously valuable.

Doctors, historians, lawyers, and everyone should be open to using LLMs correctly. Which isn't some arcane esoteric way. The first time we visit ChatGPT, it gives a list of limitations and what it shouldn't be used for. Just don't use it for these things, understand its limitations, and then I think it's fine to use it in professional contexts.

Also, GPT-4 and 3.5 now is very different from the original ChatGPT that wasn't a significant departure from GPT-3. GPT-3 hallucinated everything that could resemble a fact more than an abstract idea. What we have now with GPT-4 is much more aligned. It probably wouldn't produce what vanilla ChatGPT produced for this lawyer. But the same principles of reasonable use apply. The user must be the final discriminator that decides whether the output is good or not.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#282
post #263
post #207

Earlier quoted context omitted.

> Large Language Models (LLMs) are never wrong, and they do not make mistakes. I call B.S. If LLMs never made mistakes we wouldn't train them. Any random initialization would work.

I've noticed that there's a lot of shallow fulmination on HN recently. People say things like "I call bullshit", or "I don't believe this for a second", or even call others demeaning things. My brother (and I say this with empathy), no one is here to hear your vehement judgement. If you have anything of substance to contribute, there are a million different ways to express it with kindness and constructively. As for…

Hiding a backhanded dig ("anything of substance"?!) is worse. Just call bullshit in return. I wear big britches.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#283

Earlier quoted context omitted.

Whether a statement is true or false doesn’t depend on the mechanism generating the statement. We should hold these models (or more realistically, their creators) to the same standard as humans. What do we do with a human that generates plausible-sounding sentences without regard for their truth? Let’s hold the creators of these models accountable, and everything will be better.

No. What does this even mean? How would you make this actionable? LLM's are not "fact retrieval machines", and open AI is not presenting chat GPT as a legal case database. In fact they already have many disclaimers stating that GPT may provide information that is incorrect. If humans in their infinite stupidity choose to disregard these warnings, that's on them. Regulation is not the answer.

LLM's are not fact retrieval machines you say? but https://openai.com/product claims:

"GPT-4 can follow complex instructions in natural language and solve difficult problems with accuracy."

"use cases like long form content creation, extended conversations, and document search and analysis."

and that's why we need regulations. In US, one needs FDA approval before claiming that a drug can treat some disease, the food preparation industry is regulated, vehicles are regulated and so on. Given existing LLMs marketing, this should have the same warnings, probably similar to "dietary supplements":

"This statement has not been evaluated by the AI Administration. This product is designed to generate plausible-looking text, and is not intended to provide accurate information"

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#284

Earlier quoted context omitted.

I think ChatGPT and Photoshop are both "designed for" the creation of novel things. In Photoshop, though, the intent is clearly up to the user. If you edit that photo, you know you're editing the photo. That's fairly different than ChatGPT where you ask a question and this product has been trained to answer you in a highly-confident way that makes it sound like it actually knows more than it does.

If we’re moving past the marketing questions/concerns, I’m not sure I agree. For me, for now, ChatGPT remains a tool/resource, like: Google, Wikipedia, Photoshop, Adaptive Cruise Control, and Tesla FSD, (e.g. for the record despite mentioning FSD, I don’t think anyone should ever take a nap while operating a vehicle with any currently available technology). Did I miss when OpenAI marketed ChatGPT as a truthful resour…

I don't think it "legal matters" or not is important.

OpenAI is marketing ChatGPT as accurate tool, and yet a lot of times it is not accurate at all. It's like.. imagine Wikipedia clone which claims earth is flat cheese, or a Cruise Control which crashes your car every 100th use. Would you call this "just another tool"? Or would it be "dangerously broken thing that you should stay away from unless you really know what you are doing"?

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#285
> What’s much harder though is actually getting it to double-down on fleshing those out.

Now, it is. When ChatGPT first became public though, those were the Wild West days where you could get it to tell you anything, including all sorts of unethical things. And it would quite often double-down on "facts" it hallucinated. With current GPT-3.5 and GPT-4, the alignment is still a challenging problem, but it's in a much better place. I think it's unlikely a conversation with GPT-4 would have gone the way it did for this lawyer.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#286
post #43

Earlier quoted context omitted.

Or maybe validated one or two, and then assumed they must all be correct.

I often want good data, so I validate everything. ChatGPT tends to only give a limited number of results in the response.

GPT-4 has made some substantial improvements recently in references. A few weeks ago using it would get things sort-of correct (the authors would be nearly the same, with one or two just added in or removed, maybe journal number a bit off, etc) and sometimes hallucinate more wholesale, or at least enough that searching couldn't find it as easily. Now, it gets things pretty much perfectly correct everytime as far as I can tell. It gives links that work, and exactly the right reference, every time, even without plugins. However, with plugins now, it doesn't even matter so much any more since it can browse anyway. Very impressive nonetheless.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#287

Earlier quoted context omitted.

My main takeaway is that failing the second half of the class and misspelling ChatGPT leads to > 10x engagement.

Err, out of abundance of caution, the misspelling of "ChatGPT" which I [sic]'d is original to the Texas A&M professor, who repeated the misspelling multiple times in his email/rant. The HN poster quoted the professor literally, and I am thus transitively [sic]'ing the professor – not the HN poster. I am not mocking an HN poster's typo.

That's interesting. Unlike Reddit, maintaining an article's actual title isn't a priority on this site, and moderation is only too happy to change it at their whims. I'm surprised that the spelling wasn't corrected by mods out of pedantry.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#288
post #282
post #263

Earlier quoted context omitted.

I've noticed that there's a lot of shallow fulmination on HN recently. People say things like "I call bullshit", or "I don't believe this for a second", or even call others demeaning things. My brother (and I say this with empathy), no one is here to hear your vehement judgement. If you have anything of substance to contribute, there are a million different ways to express it with kindness and constructively. As for…

Hiding a backhanded dig ("anything of substance"?!) is worse. Just call bullshit in return. I wear big britches.

No dig intended.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#289
post #283

Earlier quoted context omitted.

No. What does this even mean? How would you make this actionable? LLM's are not "fact retrieval machines", and open AI is not presenting chat GPT as a legal case database. In fact they already have many disclaimers stating that GPT may provide information that is incorrect. If humans in their infinite stupidity choose to disregard these warnings, that's on them. Regulation is not the answer.

LLM's are not fact retrieval machines you say? but https://openai.com/product claims: "GPT-4 can follow complex instructions in natural language and solve difficult problems with accuracy." "use cases like long form content creation, extended conversations, and document search and analysis." and that's why we need regulations. In US, one needs FDA approval before claiming that a drug can treat some disease, the food…

[deleted]

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#290
post #280
post #142

Earlier quoted context omitted.

This harks back to around 1999 when people would often blame computers for mistakes in their math, documents, reports, sworn filings, and so on. Then, a thousand different permutations of "computers don't make mistakes" or "computers are never wrong" became popular sayings. Large Language Models (LLMs) are never wrong, and they do not make mistakes. They are not fact machines. Their purpose is to abstract knowledge a…

I just went to ChatGPT page, and was presented with the text: "ChatGPT: get instant answers, find creative inspiration, and learn something new. Use ChatGPT for free today." If something claims to give you answers, and those answers are incorrect, that something is wrong. Does not matter what it is -- model, human, dictionary, book. Claiming that their purpose is "to produce plausible language" is just wrong.. no one…

When you first use it, a dialog says “ChatGPT can provide inaccurate information about people, places, or facts.” The same is said right under the input window. In the blog post first announcing ChatGPT last year, the first limitation listed is about this.

Even if the ChatGPT product page does not specifically say that GPT can hallucinate facts, that message is communicated to the user several times.

About the purpose, that is what it is. It’s not clearly communicated to non-technical people, you are right. To those familiar to the AI semantic space, LLM already tells the purpose is to generate plausible language. All the other notices, warnings, and cautions point casual users to this as well, though.

I don’t know… I can see people believing what ChatGPT says are facts. I definitely see the problem. But at the same time, I can’t fault ChatGPT for this misalignment. It is clearly communicated to the users that facts presented by GPT are not to be trusted.

Post reply on HN