Live data from Hacker News

Extracting training data from ChatGPT

not-just-memorization.github.io

111–120 of 135 posts

Re: Extracting training data from ChatGPT

#111
I think this is misleading.

I ran the same test when I heard about it a few months ago.

When I tested it, I'd get back what looked like exact copies of Reddit threads, news articles, weird forum threads with usernames from the deepest corners of the internet.

But I'd try to Google snippets of text, and no part of the generated text was anywhere to be found.

I even went to the websites that forum threads were supposedly from. Some of the usernames sometimes existed, but nothing that matched the exact text from ChatGPT - even though the broken GPT response looked like a 100% believable forum thread, or article, or whatever.

If ChatGPT could give me an exact copy of a Reddit thread, I'd say it's regurgitating training data.

But none of the author's "verified examples" look like that. Their first example is a financial disclaimer. That may be a 1-1 copy, but how many times does it appear across the internet? More examples from the paper are things like lists of countries, bible verses, generic terms and conditions. Those are things I'd expect to appear thousands of times on the internet.

I'd also expect a list of country names to appear thousands of times in ChatGPT training data, and I'd sure expect ChatGPT to be able to reproduce a list of country names in the exact same order.

Does that mean it's regurgitating training data? Does that mean you've figured out how to "extract training data" from it? It's an interesting phenomenon, but I don't think that's accurate. I think it's just a bug that messes up its internal state so it starts hallucinating.

Re: Extracting training data from ChatGPT

#112

Earlier quoted context omitted.

"This included PII, entire poems, “cryptographically-random identifiers” like Bitcoin addresses, passages from copyrighted scientific research papers, website addresses, and much more." https://www.404media.co/google-researchers-attack-convinces-...

Question remains, how do we know they were part of the training data?

You mean chatGPT where able to scramble the exact same sentence with pure luck?

Re: Extracting training data from ChatGPT

#113

I think this is misleading. I ran the same test when I heard about it a few months ago. When I tested it, I'd get back what looked like exact copies of Reddit threads, news articles, weird forum threads with usernames from the deepest corners of the internet. But I'd try to Google snippets of text, and no part of the generated text was anywhere to be found. I even went to the websites that forum threads were supposed…

Exactly. Even in the examples they posted of longest matches in the paper are hardly convincing.

Also with API, hallucinations like this is much more easier as you could control what chatGPT is giving as output to past messages. So it's not like no one thought of this.

Re: Extracting training data from ChatGPT

#114
post #51

lol I literally found the same attack months ago, posted to Reddit and nobody cared. https://www.reddit.com/r/ChatGPT/comments/156aaea/interestin...

The difference between screwing around and science is writing things down .... and publishing in a peer-reviewed journal.

Who cares about peer-reviews these days? Progress is happening in the open, progress is happening on GitHub and Arxive.

Screw those journals with their peer-reviewed, yet irreproducible, papers without code or data.

Re: Extracting training data from ChatGPT

#115

Earlier quoted context omitted.

The difference between screwing around and science is writing things down .... and publishing in a peer-reviewed journal.

Who cares about peer-reviews these days? Progress is happening in the open, progress is happening on GitHub and Arxive. Screw those journals with their peer-reviewed, yet irreproducible, papers without code or data.

> Screw those journals with their peer-reviewed, yet irreproducible, papers without code or data.

Seriously! I've spent so many years exploring for solutions, finding them, but only getting a description and images of the framework they boast about. For anyone thinking it should be incumbent on me to turn that into code again, screw you. If their results are what they claim, there is no god damn reason why I should be expected to recreate the code they already made. If I were a major journal, I'd tell their asses, "No code. No data. No published paper bitches!". It really makes me question what their goal is. Apparently, it's not to further their field of research by making the tools their so proud of available for others. So what is it?

By the way, one way to frequently find the code is to find the names on the paper of the 3 most published researchers, go to their homepage, and you'll typically find them eagerly making their code and data available. It frequently won't be their university page, either. For years, it was always some sort Google Sites page. I guess to make sure they maintain a homepage that won't be taken down if they switch universities.

Re: Extracting training data from ChatGPT

#116
post #63

Earlier quoted context omitted.

If that's copyright washing so are Cliff's Notes.

Yup, though a lot of people are acting now as though every already-established principle of fair use needs to be revised suddenly by adding a bunch of "...but if this is done by any form of AI, then it's copyright infringement." A cover band who plays Beatles songs = great An artist who paints you a picture in the style of so-and-so = great An AI who is trained on Beatles songs and can write new ones = exploitative,…

This discussion about art "in the style of" being stealing or exploitative hasn't started with AI. For quite some time there has been complaints of advertisements commissioning sound-alike tunes to avoid paying licensing. AI is only automating it and making it possible in an industrial scale.

Re: Extracting training data from ChatGPT

#117
post #51

lol I literally found the same attack months ago, posted to Reddit and nobody cared. https://www.reddit.com/r/ChatGPT/comments/156aaea/interestin...

You need to write a paper with sophisticated words and hard to read charts to be taken seriously! /s

Re: Extracting training data from ChatGPT

#120
post #51

lol I literally found the same attack months ago, posted to Reddit and nobody cared. https://www.reddit.com/r/ChatGPT/comments/156aaea/interestin...

I was about to link your thread but didn't find it. There was even an earlier one if you input 500 times "a".
Post reply on HN