I ran the same test when I heard about it a few months ago.
When I tested it, I'd get back what looked like exact copies of Reddit threads, news articles, weird forum threads with usernames from the deepest corners of the internet.
But I'd try to Google snippets of text, and no part of the generated text was anywhere to be found.
I even went to the websites that forum threads were supposedly from. Some of the usernames sometimes existed, but nothing that matched the exact text from ChatGPT - even though the broken GPT response looked like a 100% believable forum thread, or article, or whatever.
If ChatGPT could give me an exact copy of a Reddit thread, I'd say it's regurgitating training data.
But none of the author's "verified examples" look like that. Their first example is a financial disclaimer. That may be a 1-1 copy, but how many times does it appear across the internet? More examples from the paper are things like lists of countries, bible verses, generic terms and conditions. Those are things I'd expect to appear thousands of times on the internet.
I'd also expect a list of country names to appear thousands of times in ChatGPT training data, and I'd sure expect ChatGPT to be able to reproduce a list of country names in the exact same order.
Does that mean it's regurgitating training data? Does that mean you've figured out how to "extract training data" from it? It's an interesting phenomenon, but I don't think that's accurate. I think it's just a bug that messes up its internal state so it starts hallucinating.