Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

221–230 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#221

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.

I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#222

Earlier quoted context omitted.

I refer you to this: https://news.ycombinator.com/item?id=49643556 Quoting: > "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

Old Facebook trick - likely resetting that box each time the app is updated.

"We've made some updates to improve the security and privacy experience" => "We've changed some of the options available and reset everyone to defaults"

Re: More questions about whether researchers can trust OpenAI with unpublished math

#224
post #191

Earlier quoted context omitted.

There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem". I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

Whoa that's a slippery slope! Next you'll want model runners to cite the data their models were trained on

In fact we should though.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#225

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

Literally every famous open math problem has had >1 mathematicians ask ChatGPT to solve it. Probably greater than >1000 if you count randos. There is no math problem that OpenAI/Anthropic can solve that didn’t have users already try it in Chat/Claude.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#226

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#228

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.…

I think Google does train on anything you put into docs if you aren't careful with the Gemini integration?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#229

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

Imagine OpenAI Astra model weights were made public because the datacenter they use had T&C that allows them to make them public

Would that be ok in your mind?

Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)

Nobody would care if they provided published research that author made public same as a google search would offer that.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#230
People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources.

It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.

Post reply on HN