Live data from Hacker News

Extracting training data from ChatGPT

not-just-memorization.github.io

131–135 of 135 posts

Re: Extracting training data from ChatGPT

#132
post #51

lol I literally found the same attack months ago, posted to Reddit and nobody cared. https://www.reddit.com/r/ChatGPT/comments/156aaea/interestin...

I tell you what nialv7 I feel ya. Not only that, it makes me wonder how many great things have gone unnoticed. Partly why I'm glued to HN is bc how on earth do you find these gems otherwise??

Re: Extracting training data from ChatGPT

#133

Earlier quoted context omitted.

"This included PII, entire poems, “cryptographically-random identifiers” like Bitcoin addresses, passages from copyrighted scientific research papers, website addresses, and much more." https://www.404media.co/google-researchers-attack-convinces-...

Question remains, how do we know they were part of the training data?

The same way we know that a million monkeys won't spit out Shakespeare's works in any reasonable amount of time. Simple probabilities.

Re: Extracting training data from ChatGPT

#135
post #29

Maybe this is what Altman was less than candid about. That the speed up was bought by throwing RAG into the mix. Finding an answer is easier than generating one from scratch. I don’t know if this is true. But I haven’t seen an LLM spit out 50 token sequences of training data. By definition (an LLM as a “compressor”) this shouldn’t happen.

Uh, he said right in dev day that Turbo was updated using cached data in some fashion and thats how they updated the model to 2023 data

[deleted]
Post reply on HN