Live data from Hacker News

Mathematicians want proof OpenAI didn't use their work

theverge.com

31–40 of 93 posts

Re: Mathematicians want proof OpenAI didn't use their work

#32

Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.

Multiple OpenAI staff have publicly said they cannot do that as accessing specific user settings without their consent (or legal requirement) violates their internal privacy policy.

However, the mathematicians could easily declare whether they had the toggle on or off. Yet curiously, they will not say!

Re: Mathematicians want proof OpenAI didn't use their work

#34

If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.

This is exactly the issue. What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

Andreas opted out on June 29th[0]. As discussed elsewhere on HN[1]

[0] https://mathstodon.xyz/@andreasthom/117240535270608201

[1] https://news.ycombinator.com/item?id=49638353

Re: Mathematicians want proof OpenAI didn't use their work

#35

If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.

Only under gross negligence would it be unprovable: Did you use a model whose training set included user data? Did the transcripts of any of the agents include a tool call whose result including user data?

>Did you use a model whose training set included user data?

OpenAI, as with all AI companies, openly admits that it trains on user data unless the user opts out. But the mathematicians have not said whether or not they opted out.

>Did the transcripts of any of the agents include a tool call whose result including user data?

They have already explicitly denied this.

Re: Mathematicians want proof OpenAI didn't use their work

#37

If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.

This is exactly the issue. What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

One of them said they turned it off in June, back in the original mastodon thread.

Re: Mathematicians want proof OpenAI didn't use their work

#38

Earlier quoted context omitted.

This is exactly the issue. What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

Andreas opted out on June 29th[0]. As discussed elsewhere on HN[1] [0] https://mathstodon.xyz/@andreasthom/117240535270608201 [1] https://news.ycombinator.com/item?id=49638353

Then none of his work after June 29 will be included in the training data.

What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.

Re: Mathematicians want proof OpenAI didn't use their work

#39

Earlier quoted context omitted.

This is exactly the issue. What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

One of them said they turned it off in June, back in the original mastodon thread.

Then none of his work after June will be included in the training data.

What I was referring to is the fact that neither Levent Alpöge nor Tristan Buckmaster will answer this question.

Re: Mathematicians want proof OpenAI didn't use their work

#40

If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.

What? Can they just look if they fetched certain documents/conversations and feeded them into the training loop?

You seem to be confusing context-fetching with training.
Post reply on HN