Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

31–40 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#31
post #4

TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.

I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?

Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]

If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]

[1] https://archive.is/EcwD8 [2] https://archive.is/yZdAF

Re: More questions about whether researchers can trust OpenAI with unpublished math

#32

Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.

Yeah but that will not prevent the stealing, it will only make the fight easier afterwards.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#33
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#35
post #5

This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point". They don't even claim to have had a proof, only to have been working on it.

I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capt…

But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#37
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?

The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.

That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#38
post #34

I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?

Given OpenAI's well documented history of unethical behaviour it seems adorably naive to think they actually do that in general, or that they wouldn't pull this particular data separately to generate these proofs.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#39

people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.

As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken. That's still a pretty big marker of competence in my eyes. The point of controversy seems to be who gets credit

The question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA.

It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.

Post reply on HN