Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

251–260 of 850 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#251

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

You know, I don't think I've ever accused anyone of being a shill. I've thought about it maybe a few times (daringfireball). This is going to be as close as I get. I don't know the facts in this case but I cannot believe the argument being made here with a straight face. Is it common sense that tools you use and pay for steal your work and profit off of it at your expense and without recognition? If this isn't the textbook definition victim blaming, I don't know what is.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#252
post #241

Earlier quoted context omitted.

To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training. From their docs[1] (archive copy is at [2]): > You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new…

I mean, how else would those buttons work? It's explicitly feedback data. And "this is good" or "this is bad" is empty if divorced from what "this" actually is.

If the buttons are incompatible with the absence of the feature, I'd expect the buttons not to exist when the feature is disabled. Anything else seems like a straight up footgun. I guess it'd also be acceptable to pop up a scary warning box asking "are you sure?"

Re: More questions about whether researchers can trust OpenAI with unpublished math

#253
post #241

Earlier quoted context omitted.

To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training. From their docs[1] (archive copy is at [2]): > You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new…

I mean, how else would those buttons work? It's explicitly feedback data. And "this is good" or "this is bad" is empty if divorced from what "this" actually is.

It could go into personalization / memory. Or they could be A/B testing some system prompt tuning and consider the thumbs up / thumbs down as statistical feedback on the particular flags that are enabled for your account.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#254

I think this is stupid, for three reasons: 1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods. 2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether the…

If the mathematicians are using ChatGPT, then they themselves are benefiting from the work of other ChatGPT users, so ChatGPT using their work is not wrong!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#255

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

In general, a lot of moral invariants that natural selection has rendered as "intuitive" to us are no longer intuitive or possible. These natural brakes are not braking.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#256
post #228

Earlier quoted context omitted.

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.…

I think Google does train on anything you put into docs if you aren't careful with the Gemini integration?

Yes, this is a problem with modern AI systems in general. It's not just OpenAI, and if you know any artists you know this is why they're pretty vehemently opposed to all AI.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#257

Earlier quoted context omitted.

I mean, how else would those buttons work? It's explicitly feedback data. And "this is good" or "this is bad" is empty if divorced from what "this" actually is.

If the buttons are incompatible with the absence of the feature, I'd expect the buttons not to exist when the feature is disabled. Anything else seems like a straight up footgun. I guess it'd also be acceptable to pop up a scary warning box asking "are you sure?"

> If the buttons are incompatible with the absence of the feature, I'd expect the buttons not to exist when the feature is disabled. Anything else seems like a straight up footgun.

It's called a "dark pattern." They want you to shoot yourself in the foot, so they'll do their best to aim your gun at your foot and put your finger on the trigger. And then when you do, because you don't have perfect understanding or execution, they'll say "your fault!"

Re: More questions about whether researchers can trust OpenAI with unpublished math

#258
post #241

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

To add to this, merely using the thumbs up/down button in a chat could share your entire conversation with them for model training. From their docs[1] (archive copy is at [2]): > You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new…

"may" = will, unless they screw up

Re: More questions about whether researchers can trust OpenAI with unpublished math

#259

Earlier quoted context omitted.

They've done similar things with similar risks repeatedly.

Example?

You can read their Wikipedia page [1].

[1] https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_is...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#260
And I have been proclaiming a cry of “your data for analytical purposes is being stolen” (you can’t opt out of analytical purposes) and people perhaps astroturfers straw man back to “just turn off training bro”.

Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.

Post reply on HN