ChatGPT provides false information about people, and OpenAI can't correct it
81–90 of 92 posts
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#82Earlier quoted context omitted.
People use it to get facts, and trust it more than Google. I have some tech illiterate boss that asked me to do stuff some way because "ChatGPT said so", instead of trusting me, an experience professional. It wasn't like this with a Google search, so why now ? Natural language has a big impact on how the product is perceived
We've seen numerous stories at this point about lawyers trusting AI to generate case documents that turned out to have false cites - AI generated scientific papers are being published. Doctors are using AI. Law enforcement is using AI. Everyone is using it and a lot of people are using it with the assumption that it's intelligent and factual. That it works like the computer from Star Trek. People on this very forum w…
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#83Earlier quoted context omitted.
How is LLM+RAG any different from a fine-tuned LLM without RAG? They're both trained on data, and have the capability to hallucinate.
Theoretically some people think RAG sounds more feasible to make factually accurate. After all, if you've trained an LLM on a masses of unchecked data you've scraped from the internet, your training data probably includes "Joe Biden is the president" and "Donald Trump is the president" and "Barrack Obama is the president" and "Emmanuel Macron est le président" and so on. It would be understandable if an LLM was confu…
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#84Earlier quoted context omitted.
If LLM providers want access to the EU market they will need to find a way to comply with GDPR, and if OpenAI cannot find a way to do it then a different LLM provider will.
> the GDPR requires information about individuals is accurate Given that you can make LLMs say pretty much whatever you want using the right prompts, this seems impossible. LLMs are not a search engine, and based on conversational context might say Emmanuel Macron is the president of France or a baby giraffe.
Can the LLM provide personal data of an individual who is covered by GDPR? Then the LLM is subject to GDPR.
Can this individual exercise their rights with regards to the data that the LLM returns about them? Arguably they can indeed exercise the right of access by means of the right prompts, but can the individual rectify errors or erase such data? If not, then the provider of the LLM is violating GDPR.
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#85Earlier quoted context omitted.
> the GDPR requires information about individuals is accurate Given that you can make LLMs say pretty much whatever you want using the right prompts, this seems impossible. LLMs are not a search engine, and based on conversational context might say Emmanuel Macron is the president of France or a baby giraffe.
You said it: based on context. This is not about what you can make an LLM say when being manipulated in a convoluted way to provide an inaccurate response. It's about what an LLM will say in the context of a prompt requesting personal data related to an individual who is covered by GDPR. Can the LLM provide personal data of an individual who is covered by GDPR? Then the LLM is subject to GDPR. Can this individual exe…
> ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle.
We also know that LLMs don't know the current date, and therefore can make calculation errors (which is made worse by their poor math performance as a language token generator). So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old.
There is an incalculable number of ways for LLMs to output incorrect information. In an effort to comply with strict regulation the preprompt contextual limit is going to be exceeded.
Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#86Does the GDPR require information that is not asserted to be factual to actually be factual? If I have a random number generator producing arbitrary strings, am I required to ensure that the strings do not contain untrue statements about individuals?
When people treat the information as factual and the company doesn't do enough to clarify that, then yes.
The fact that LLMs hallucinate is certainly no secret, even the linked article says OpenAI openly admits that they can't avoid it right now.
What would constitute enough?
Would it be enough to place a statement placed onscreen at the start of every conversation to say that information may not be accurate and that if the information was significant then it should be independently verified?
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#87Earlier quoted context omitted.
You said it: based on context. This is not about what you can make an LLM say when being manipulated in a convoluted way to provide an inaccurate response. It's about what an LLM will say in the context of a prompt requesting personal data related to an individual who is covered by GDPR. Can the LLM provide personal data of an individual who is covered by GDPR? Then the LLM is subject to GDPR. Can this individual exe…
Who is the judge of the degree of contextual convolution? Must the LLM remain strictly factual when you simply append "ELI5" to a prompt? > ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle. We also know that LLMs don't know the current date, and therefore can make calculation errors (which is made worse by their poor math performance as a language token generator). So on one hand it mi…
That is not personal data under GDPR.
«So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old.»
Or it might say that Macron was born on 14th July 1977, which is incorrect. The claimed impossibility to correct a date of birth returned by the LLM is the trigger of the GDPR complaint that the article refers to.
«Also this creates a situation where all but the most powerful LLMs (and LLM providers) will be non-complaint»
Only under the premise that it is somehow inevitable to feed personal data of living individuals to an LLM for training, and that the only way to correct mistaken data or to stop an LLM from providing such data is "more power".
I reject the premise, not the least because, firstly, OpenAI (the most "powerful" provider) is claiming it is impossible. All that says is that OpenAI's platform was not originally designed with that problem in mind and that, as that of now, they are unwilling to redesign it from scratch only because some guy complained in Austria. It's basically a speedrun of Microsoft claiming Internet Explorer was an essential component of Windows 98.
Meanwhile, LLMs and other AI models are an active area of research. If OpenAI truly cannot stop their LLMs from returning personal data protected by GDPR, and honestly has no way to allow data holders to exercise their rights of deletion or correction, you can be sure that some startup will disrupt the LLM market by finding a way to do it without needing to out-compete OpenAI neither in hardware nor on training corpus size.
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#88Glad that people are realizing the ChatGPT/AI hype that is driving all companies nuts by creating unnecessary peer pressure to integrate some kind of GenAI in their products even if it doesn’t make sense.
They've always been around, they've just been dismissed as "Luddites" who fear an inevitable future. But of course, the Luddites are always proven correct in hindsight.
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#89Earlier quoted context omitted.
They've always been around, they've just been dismissed as "Luddites" who fear an inevitable future. But of course, the Luddites are always proven correct in hindsight.
when has your claim "the Luddites are always proven correct in hindsight." been true?
Contrary to popular propaganda, the original Luddites weren't opposed to technology, they were opposed to the effect of technological progress on the working class. They knew that automation and mass production would be used to devalue labor, flood the market with inferior products, and that all of the benefits and profit from the industrial revolution would go to the corporations, at the expense of their quality of life. And they were correct.
And modern day "Luddites" were correct about the centralization and commoditization of the web, social media, crypto and NFTs, Elon Musk and autonomous vehicles, and will be proven right about AI. Tech has nothing left to offer but grift upon grift.
https://www.newyorker.com/books/page-turner/rethinking-the-l...
https://news.ycombinator.com/item?id=37664682
https://librarianshipwreck.wordpress.com/2018/01/18/why-the-...
Re: ChatGPT provides false information about people, and OpenAI can't correct it
#90Earlier quoted context omitted.
Who is the judge of the degree of contextual convolution? Must the LLM remain strictly factual when you simply append "ELI5" to a prompt? > ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle. We also know that LLMs don't know the current date, and therefore can make calculation errors (which is made worse by their poor math performance as a language token generator). So on one hand it mi…
«> ELI5 how is France governed? > ...and Macron is the lion, the king of the jungle.» That is not personal data under GDPR. «So on one hand it might say Macron was born December 1st 1977 (which is correct), but if you ask how old he is some LLMs might say 45 years old.» Or it might say that Macron was born on 14th July 1977, which is incorrect. The claimed impossibility to correct a date of birth returned by the LLM…
Indeed if a startup can find a way to scrub PII of living people from 20 billion pages of text (and prevent LLMs from ever hallucinating) they would be quite a valuable company, in the LLM dev space and numerous other ventures. Until then the EU might have to go without access to language models.