Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

451–460 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#451

Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims. This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers. All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal…

> but right now there's no credible evidence, only claims. since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried. you're being naive

OpenAI has said that their models were definitely not trained on any of Buckmaster's sessions after July 3rd (from https://archive.ph/75WcF); likely they found that's when he switched the "allow training" setting off.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#452
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

So, at my company (and most companies I think), we use confidential in-house versions of the AI software. We don't want any confidential information leaking into the public realm. Are these scientists doing that, or are they just using the public version of the software?

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#454
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

The irony is that OpenAI got into this trouble only because they tried to play "nice". They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list. They wanted to give Buckmaster a chance to be the one solved N-S problem.

While this behavior is highly questionable, if OpenAI just published the final result without notifying Buckmaster first and simply cited his previous researches, there would be no ground for anyone to accuse OpenAI for anything. Their self-perceived "generosity" backfired dearly and I'm sure they'll never make the same mistake again. There is probably a policy forbidding any OpenAI employee to contact external researchers like that now.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#455
Isn’t that the whole spiel of these things, you run all kind of text and other data through it and it kind of remembers it and learns from it then it spouts it back out like a human would. Makes sense to me that a training run based on conversations that were fed into the system by users is results in the model learning from these so the model will spit the knowledge back out again, just in a way that’s not directly attributable to the original content (which is the most important step as otherwise it would just be plagiarism). I guess that’s why OpenAI can get better and better as well so fast, people work with it and teach it how to do things by giving it feedback and iterating with it, and all that goes back into the training loop. And training data about millennium prize problems is probably quite spars. Wonder if anyone has tried injecting nonsense science into the training data (e.g. work out a fantasy science theory with names and all kinds of stuff) to see if the model will regurgitate it in a couple of months for other users.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#456
post #428

Earlier quoted context omitted.

> but I can believe it to be accidental What accident is it when the system is designed to function that way?

Their claim is that training on their solution is "unlikely but possible". Consider this scenario. Has a google crawler read my new novel, which I may or may not have posted on my blog, page by page, as I wrote it? Can you, without knowledge of what I have actually done, claim that the google crawler has not seen the novel? Without any evidence that I have posted the novel online, it might be tempting to say that the…

I'm not taking them at their word, sorry. Genuinely, there is no reason to.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#459

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems

You forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers.

Everyone in this thread seems to have made up their mind about OpenAI's guilt though.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#460
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

The irony is that OpenAI got into this trouble only because they tried to play "nice". They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list. They wanted to give Buckmaster a chance to be the one solved N-S problem. While this behavior is highly questionable, if OpenAI just published the final result without notifying Buckmaster first and simply ci…

The reality would be the same. They probably used prior session history between the research and Astra to train the internal model, and used it to front run-the researcher.

This is the biggest self-own in the history of software. If you can relate to Pixar, OpenAI is Chick Hicks celebrating at the end of the Piston Cup and wondering why he's getting booed.

The lack of self-awareness is something to behold, and says a lot about their corporate values.

Post reply on HN