Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

261–270 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#261

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing. Or at least: tell me I should be careful/worry about those particular thi…

No, people need to take responsibility for their actions. We don't need more over regulation.

Verify, don't trust.

You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#262

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

Why are we being so charitable to trillion dollar organizations here? If OAI keeps re-enabling the train model toggle on every app update to codex, does it also fall under "common sense opsec" to re-disable this toggle every time?

Arguably you expect that unless you are explicit about providing permissions to these labs to use your data for training then your data is yours, and not theirs. Especially on a paid account (let alone an enterprise one). Why is the opt-out supposed to be "common sense opsec" rather than the opt-in should be common sense regulation?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#263
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

Yes but what about “analytical purposes” what does that cover and can you turn it off? I have found out you cannot. It’s the Trojan backdoor to your data.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#264

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

This is such a naive position though. The most successful companies of the last decade have precisely been ... selling usage data. Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.

Yeah but marketing companies are aggressively fingerprinting and stalking you to sell you snacks from japan, or oscilloscopes because they figured out you work in a lab, etc. Not to fuck you over by stealing your livelihood (which is what is happening to these mathematicians). It's on a whole new scale.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#265

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

It tells me that AI companies are just another mechanism to extract and extort value from the masses for the rich.

Just another rich man’s trick

Perhaps the last one before they destroy that world and try to hide away as people forget and history is rewritten again. I don’t think they’ll succeed this time.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#266
Both things can be true:

1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.

2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.

The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#267

Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript wou…

Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#269
post #261

Earlier quoted context omitted.

We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing. Or at least: tell me I should be careful/worry about those particular thi…

No, people need to take responsibility for their actions. We don't need more over regulation. Verify, don't trust. You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.

So I should verify that my data isn't just shared for product improvement but also to take credit from me?

I should verify with wireshark and other software that my LG TV isn't listening to me and selling my data.

I should make sure that whatever product I buy I spend the time to go over every setting page in case there is a switch (defaulted on) that says "I authorize the sell of my data".

I should make sure to look at every ingredients on the back of each box of food product to make sure it will not kill me.

I should document myself on the undisclosed growing practices (because no packaging here) of the vegetables and fruit I am buying and make sure that I equal PhD researchers on the dangers of the pesticides used by the specific company I am buying from.

I should make sure myself that the battery in any device is up to standard and will not blow me and my living place by researching the factory that made it and buying testing equipment.

I should make sure to educate myself on how my retirement 401k investment strategy works otherwise, I may not have proper retirement.

... I could go on and on; it's infinite.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#270
post #261

Earlier quoted context omitted.

We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing. Or at least: tell me I should be careful/worry about those particular thi…

No, people need to take responsibility for their actions. We don't need more over regulation. Verify, don't trust. You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.

i'm sure you consult your attorney every time you agree to terms and conditions
Post reply on HN