Live data from Hacker News

Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

arstechnica.com

171–180 of 212 posts

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#171

Earlier quoted context omitted.

ChatGPT 3.5 or GPT 4? Almost every negative comment about LLMs is by someone using an older, weaker model and making generalisations. Here’s GPT 4 giving me a riddle: https://chat.openai.com/share/1753ce5a-d44d-44ac-bc97-599a26...

> But was it GPT4 I keep seeing this cop-out, which ignores that it's fundamentally the same architecture, and has the same flaws. More wallpaper to hide the cracks better makes it an even worse tool for these use cases because all it does is fool more people into thinking it has capabilities that it fundamentally doesn't.

To me it's a bit like someone making the claim "humans are flawed, and we should think critically about the things they say", and someone responding with "well which human are you talking about? Because Einstein is orders of magnitude above the Walmart checkout guy".

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#172

Earlier quoted context omitted.

Look up 'strict liability'.

That's an exception, not the general rule.

What of it? It exists in some cases, so the mens rea requirement is not universal.

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#173

Earlier quoted context omitted.

Maybe, but it's surprisingly good in the face of all the non-version-indicating complaints about how terrible people think it is. Mostly I doubt that the lawyer was using GPT4, because the lawyer sounds like the kind of person who would be ignorant of the significance of the difference.

The kind of person too lazy to check the output of a computer program before submitting it to a court of law is the type of person too cheap to pay $20 for the good version of the program. Think: Lionel Hutz.

No, checking was done!

"Oops, I'd better remove that comma".

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#174

Every citation created by chat gpt will be hallucinated. It knows what they look like. It doesn’t know what they are. It doesn’t actually “know” anything. It is a statistical engine for generating the next reasonable looking word.

I asked an LLM "during the day, what colour is the sky?", and it told me the sky is typically blue, and that this has to do with sunlight scattering as it enters our atmosphere.

I think it's worth using "hallucinate" to more precisely refer to inaccurate content. I'd think of this in the same sense as a false-positive from some test result.

I like the term "bullshit" when used to mean "made without regard to the truth".

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#175
post #19

I asked ChatGPT to tell me a riddle. It was “What is always hungry, needs to be fed, and makes your hands red?” (Or something like that) I asked for a hint about 5 times and it kept giving more legitimate sounding hints. Finally I gave up and asked for the answer to the riddle, and it spit out a random fruit which made no sense as the answer to the riddle. I then repeated the riddle and asked ChatGPT what the answer…

> Or, learn how to say “I don’t know” This is the correct answer. It is like a sad salesman who is out of his depth, but decides to keep bullshiting!

I'd rather a confidence score for each response. The last thing I need is another reason for the AI to ignore the question or feel the need to explain why it was ignoring it.

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#176
post #43

Prior discussion on the same matter from a different link: https://news.ycombinator.com/item?id=36095352

Thanks! Macroexpanded:

Lawyer cites fake cases invented by ChatGPT, judge is not amused - https://news.ycombinator.com/item?id=36097900 - May 2023 (304 comments)

A man sued Avianca Airline – his lawyer used ChatGPT - https://news.ycombinator.com/item?id=36095352 - May 2023 (127 comments)

ChatGPT-Authored Legal Filing “Replete with Citations to Non-Existent Cases" - https://news.ycombinator.com/item?id=36092509 - May 2023 (71 comments)

Presumably also related from earlier today:

Mandatory Certification Regarding Generative Artificial Intelligence - https://news.ycombinator.com/item?id=36131942 - May 2023 (35 comments)

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#177

Earlier quoted context omitted.

It's still fundamentally the same, hallucinates just the same, and anthropomorphizes itself as a confident, knowledgeable, intelligent being just the same. A newer, better, faster, more capable car still isn't an airplane, even if it go fast enough to spend several seconds in the air.

Sure, and 40 year olds have the same capabilities as 4 year olds, because "same architecture" or "fundamentally the same". And putting random weights inside the GPT-4 model architecture should behave "fundamentally the same" as the trained GPT-4 weights, because it's "same architecture". Forget this "training" stuff.

It's not a person, it's a machine. And it's one that will still produce hallucinations that embarrassingly prove that it has no notion of intelligence, and do so confidently. That it does so less than it's sibling is entirely irrelevant.

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#178

Earlier quoted context omitted.

> Is anyone really believing a lawyer is that stupid? There's over a million lawyers in the United States. You'd expect at least one of them to be a 1-in-a-million level of bad, or 4.7 standard deviations below the mean assuming a Gaussian distribution of competency. An average person would normally never come across that lawyer in their lifetime, but media will find that lawyer and amplify their mistakes to everyone…

There's no reason to believe it's a Gaussian distribution around the mean. Given that there are admission tests, you'd rather hope it's only the tail end of a Gaussian distribution, with the cutoff being what's required to pass the bar.

There's going to be a distribution around the mean of the proctoring of those tests. There may even be outright corruption and bribery going on at the tail end.

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#179

Earlier quoted context omitted.

Sure, and 40 year olds have the same capabilities as 4 year olds, because "same architecture" or "fundamentally the same". And putting random weights inside the GPT-4 model architecture should behave "fundamentally the same" as the trained GPT-4 weights, because it's "same architecture". Forget this "training" stuff.

It's not a person, it's a machine. And it's one that will still produce hallucinations that embarrassingly prove that it has no notion of intelligence, and do so confidently. That it does so less than it's sibling is entirely irrelevant.

Nope, wrong. The number of errors and the magnitude of each is very relevant.

Re: Lawyer cited 6 fake cases made up by ChatGPT; judge calls it “unprecedented”

#180
post #56

Earlier quoted context omitted.

I've had a few moments with ChatGPT that are great anecdotes similar to your own: - Asked it to generate a MadLib for me to play that was no more than a paragraph long. It produced something that was several paragraphs wrong. I told it "no. That's X paragraphs. I asked for one that is only 1 paragraph long" and it would respond "I'm sorry for the misunderstanding. Let me try again" and then would make the same mistak…

There's definitely a potential for a D&D DM with an LLM, but you'd need a lot of careful prompting and processing to handle the token limits today's models have. Simply put: a d&d game has more story and state than the 30,000-ish words an LLM can think about at once. I think there's a lot of interesting opportunities there.

That’s the whole point of these "agents", and things like LangChain or LlamaIndex.

Haven’t gotten around to that part yet, it seems it could help.

Post reply on HN