Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

241–250 of 848 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#241

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training.

From their docs[1] (archive copy is at [2]):

> You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new conversations will not be used to train our models.

> For a linked teen account, a parent or guardian may manage whether conversations can be used to improve our models through Parental controls.

> Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.

[1] https://help.openai.com/en/articles/5722486-how-your-data-is...

[2] https://web.archive.org/web/20260910151242/https://help.open...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#242

But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?

If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.

Sam Altman has provided people with more generous usage than Claude or Gemini.

There is no doubt ChatGPT is the most generous LLM provider!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#243

Earlier quoted context omitted.

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.…

That’s … Googles entire reason for making these “you don’t pay with money” tools. Did you not understand that?

What? I don't understand how you even came up with this idea, much less consider it so obvious to condescend about it. Do you have even a single example of a research project that got scooped because the Google Docs team forwarded their private documents to someone?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#244
post #207

Earlier quoted context omitted.

There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem". I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

But who gets credit then? Every mathematician who's work was read by an LLM during training? By that logic, we should put every published mathematician's name on the authorship of this paper. Sure, this guy should be higher up the list, but everyone's name should be on it by standard academic convention. But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument t…

Does every mathematician get cited in every maths paper? I think it's pretty clear who should be cited.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#245

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

Imagine OpenAI Astra model weights were made public because the datacenter they use had T&C that allows them to make them public Would that be ok in your mind? Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did) Nobody would care if they provided published research that author made public same as a google search would offer that.

  Would that be ok in your mind?
It would in my mind. Hopefully companies have looked through the agreement.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#246
Big Tech will slurp up every piece of data it can about you and sell it to anyone it can, all to make you the target of someone else's goals, whether that is an advertiser, an employer, law enforcement, a stalker, or the government.

You will not be able to opt out unless you completely isolate yourself from society, tough shit.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#247

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

> If it is found that OpenAI and other labs are not respecting the training opt out

HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#248
post #216
post #160

Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.

What does "with no substantive human input" mean? All of the training data is human input, isn't it?

They also train on synthetically generated data.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#249
Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.

My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#250
post #241

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training. From their docs[1] (archive copy is at [2]): > You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new…

I mean, how else would those buttons work? It's explicitly feedback data. And "this is good" or "this is bad" is empty if divorced from what "this" actually is.
Post reply on HN