Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

181–190 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#181
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).

People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:

  1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)

  2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.

  3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.

  4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).

edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#182
I think this is stupid, for three reasons:

1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

Re: More questions about whether researchers can trust OpenAI with unpublished math

#184
post #33

It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.

I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What…

> and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.

Do you have any evidence of this? They don't dispute the timeline, but they never said they knew what Levant/Buckmaster were doing.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#188
Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#190

I think this is stupid, for three reasons: 1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods. 2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether the…

> in this case too they didn't actually have the solution

Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.

I would almost expect training to overweight conversations with novel scientific and mathematical implications.

> the only reason they can't definitively say no is that for privacy reasons

They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.

Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.

> that opted-out user data was used for training

Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.

Post reply on HN