Live data from Hacker News

Lawyer cites fake cases invented by ChatGPT, judge is not amused

simonwillison.net

261–270 of 319 posts

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#261
post #142

Earlier quoted context omitted.

ChatGPT did not lie; it cannot lie. It was given a sequence of words and tasked with producing a subsequent sequence of words that satisfy with high probability the constraints of the model. It did that admirably. It's not its fault, or in my opinion OpenAI's fault, that the output is being misunderstood and misused by people who can't be bothered understanding it and project their own ideas of how it should function…

This harks back to around 1999 when people would often blame computers for mistakes in their math, documents, reports, sworn filings, and so on. Then, a thousand different permutations of "computers don't make mistakes" or "computers are never wrong" became popular sayings. Large Language Models (LLMs) are never wrong, and they do not make mistakes. They are not fact machines. Their purpose is to abstract knowledge a…

TBH, I think the answer to this is to fill the knowledge gap. Exactly how is the difficult part. How do you make "The moon is made of rock" be more likely than "The moon is made of cheese" when there is significantly more data (input corpus) to support the latter?

Extrapolating that a bit, future LLMs and training exercises should be ingesting textbooks and databases of information (legal, medical, etc). They should be slurping publicly available information from social media and forums (with the caveat that perhaps these should always be presented in the training set with disclaimers about source / validity / toxicity).

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#262
post #258

Earlier quoted context omitted.

That's a different error context I think. It's a mistake if the model produces nonsense, because it's designed to produce realistic text. It's not a mistake if it produces non-factual information that looks realistic. And it fundamentally cannot always produce factual information, it doesn't have that capacity (but then, neither do humans and with the ability to source information this statement may well be obsolete…

Indeed, I intended to imply that a model cannot err in the same way a computer cannot. This parallels the concept that any tool is incapable of making mistakes. The notion of a mistake is contingent upon human folly, or more broadly, within the conceptual realm of humanity, not machines. LLMs may generate false statements, but this stems from their primary function - to conjure plausible language, not factual stateme…

I conceptually agree with you that a fool blames his tools.

However! If LLMs produced only lies no one would use them! Clearly truthiness is a desired property of an LLM the way sufficient hardness is of a bolt. Therefore, I maintain that an LLM can be wrong because truthiness is its primary function.

A craftsman really can just own a shitty hammer. He shouldn't use it. But the hammer can inherently suck at being a hammer.

Consider that some cars really are lemons.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#263
post #207
post #142

Earlier quoted context omitted.

This harks back to around 1999 when people would often blame computers for mistakes in their math, documents, reports, sworn filings, and so on. Then, a thousand different permutations of "computers don't make mistakes" or "computers are never wrong" became popular sayings. Large Language Models (LLMs) are never wrong, and they do not make mistakes. They are not fact machines. Their purpose is to abstract knowledge a…

> Large Language Models (LLMs) are never wrong, and they do not make mistakes. I call B.S. If LLMs never made mistakes we wouldn't train them. Any random initialization would work.

I've noticed that there's a lot of shallow fulmination on HN recently. People say things like "I call bullshit", or "I don't believe this for a second", or even call others demeaning things.

My brother (and I say this with empathy), no one is here to hear your vehement judgement. If you have anything of substance to contribute, there are a million different ways to express it with kindness and constructively.

As for RLHF, it is used to align the LLM, not to make it more factual. You cannot make a language model know more facts than what it comes out of training with. You can only align it to give more friendly and helpful output to its users. And to an extent, the LLM can be steered away from outputting false information. But RLHF will never be comprehensive enough to eliminate all hallucination, and that's not its purpose.

LLMs are made to produce plausible text, not facts. They are fantastic (to varying degrees) at speaking about the facts they know, but that is not their primary function.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#264
And so it starts, I suspect we'll be seeing this issue more and more since it's easy to just get GPT to spit out some text. I believe that the true beneficiaries of LLMs are those that are experts in their fields. They can just read the output and deal with the inaccuracies.

Does anyone know if training an LLM with just one type of data, law in this case, creates a more accurate output?

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#265
post #262
post #258

Earlier quoted context omitted.

Indeed, I intended to imply that a model cannot err in the same way a computer cannot. This parallels the concept that any tool is incapable of making mistakes. The notion of a mistake is contingent upon human folly, or more broadly, within the conceptual realm of humanity, not machines. LLMs may generate false statements, but this stems from their primary function - to conjure plausible language, not factual stateme…

I conceptually agree with you that a fool blames his tools. However! If LLMs produced only lies no one would use them! Clearly truthiness is a desired property of an LLM the way sufficient hardness is of a bolt. Therefore, I maintain that an LLM can be wrong because truthiness is its primary function. A craftsman really can just own a shitty hammer. He shouldn't use it. But the hammer can inherently suck at being a h…

I agree for the most part, but I wish to underscore the primary function inherent in each tool. For a LLM, it is to generate plausible language. For a bolt, it is to offer structural integrity. For a car, it is to provide mobility. Should these tools fail to do what they were designed for, we can rightfully deem them as defective.

GPT was not primarily made to produce factual statements. While factual accuracy certainly constitutes a desirable design aspiration, and undeniably makes the LLM more useful, it should not be expected. Automobile designers, for example, strive to ensure safety during high-speed collisions, a feature that almost invariably benefits the user. However, if someone uses their car to demolish their house, this is probably not going to leave them satisfied. And I don't think we can say the car is a lemon for this.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#266
post #262
post #258

Earlier quoted context omitted.

Indeed, I intended to imply that a model cannot err in the same way a computer cannot. This parallels the concept that any tool is incapable of making mistakes. The notion of a mistake is contingent upon human folly, or more broadly, within the conceptual realm of humanity, not machines. LLMs may generate false statements, but this stems from their primary function - to conjure plausible language, not factual stateme…

I conceptually agree with you that a fool blames his tools. However! If LLMs produced only lies no one would use them! Clearly truthiness is a desired property of an LLM the way sufficient hardness is of a bolt. Therefore, I maintain that an LLM can be wrong because truthiness is its primary function. A craftsman really can just own a shitty hammer. He shouldn't use it. But the hammer can inherently suck at being a h…

> If LLMs produced only lies no one would use them!

Yes, they would. Producing lies (well, fiction) is an important LLM application domain.

> Clearly truthiness is a desired property of an LLM

Correct. But truthiness is not truthfulness.

“Truthiness refers to the quality of seeming to be true but not necessarily or actually true according to known facts.”

https://www.merriam-webster.com/words-at-play/truthiness-mea...

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#268

Genuine question: why have these models all been trained to sound so confident? Is it not possible to have rewarded models that announced their own ignorance? Or is even that question belying an "intelligence" view of these models that isn't accurate?

I believe the GPT4 "paper" mentions that the RHLF part destroyed the models ability to gauge it's confidence.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#269

Earlier quoted context omitted.

GPT4 can double-check to an extent. I gave it a sequence of 67 letter As and asked it to count them. It said "100", I said "recount": 98, recount, 69, recount, 67, recount, 67, recount, 67, recount, 67. It converged to the correct count and stayed there. This is quite a different scenario though, tangential to your [correct] point.

The example of asking it things like counting or sequences isn't a great one because it's been solved by asking it to "translate" to code and then run the code. I took this up as a challenge a while back with a similar line of reasoning on Reddit (that it couldn't do such a thing) and ended up implementing it in my AI web shell thing. heavy-magpie|> I am feeling excited. system=> History has been loaded. pastel-matur…

Oh god, it's even worse at naming things than people are.

Re: Lawyer cites fake cases invented by ChatGPT, judge is not amused

#270
post #265
post #262

Earlier quoted context omitted.

I conceptually agree with you that a fool blames his tools. However! If LLMs produced only lies no one would use them! Clearly truthiness is a desired property of an LLM the way sufficient hardness is of a bolt. Therefore, I maintain that an LLM can be wrong because truthiness is its primary function. A craftsman really can just own a shitty hammer. He shouldn't use it. But the hammer can inherently suck at being a h…

I agree for the most part, but I wish to underscore the primary function inherent in each tool. For a LLM, it is to generate plausible language. For a bolt, it is to offer structural integrity. For a car, it is to provide mobility. Should these tools fail to do what they were designed for, we can rightfully deem them as defective. GPT was not primarily made to produce factual statements. While factual accuracy certai…

> For a LLM, it is to generate plausible language.

LLMs are not being sold as delivering only plausible language. Is it the craftsman's fault when the salesman lies?

Post reply on HN