Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

741–750 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#741

Earlier quoted context omitted.

Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritu…

It's just conflict of interest. OpenAI is trying to get billions and billions and there's so much at stake. You spend millions trying to preempt two guys. It just makes you seem like a big bully. People would get angry even if it was esports or football.

Hearing "rumors" and just trying to overtake them and then asking to collaborate instead of starting out offering the resources beforehand. Just sounds like strong arming. Just doesn't sit right with me.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#742
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

The irony is that OpenAI got into this trouble only because they tried to play "nice". They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list. They wanted to give Buckmaster a chance to be the one solved N-S problem. While this behavior is highly questionable, if OpenAI just published the final result without notifying Buckmaster first and simply ci…

If they tried to play nice they would have offered the compute upon hearing the rumors, and not just "authorship" after or close to getting a result. It's just a PR stunt.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#743
post #441

Earlier quoted context omitted.

So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.

The claim that OpenAI somehow used the mathematicians' ideas to leapfrog them seems unsupported at this time and IMHO it was irresponsible to bring it up because credulous people will immediately believe that narrative. And from my perspective, if some math folks typing in a few questions to OpenAI provides sufficient training data for OpenAI to solve a big problem... that's amazing! A few conversations/prompts out o…

The (unprovable, yes, without OpenAI being willingly transparent) argument is that openAI constructed a prompt to scoop them using some inside knowledge about the approach, which they allude to in the announcement.

In the transcripts, Brubeck is very cagey and evasive about the prompt, when it was supplied, and its contents.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#744
post #699
post #658

Earlier quoted context omitted.

> They intentionally left Buckmaster and Alpöge out of the citations. No, they asked if they could do a joint publish.

No, they asked one guy to do a joint publish conditioned on leaving the other collaborator out, with veiled threats. The joint publish part smells awfully like admission of guilt given there’s absolutely no reason to do it if you believe you independently arrived at the result using only public prior work. The leaving out collaborator part is outright academic malpractice. Disclosure: I was an academic once.

To add: with a requirement that he rewrite the proof to credit OpenAI.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#745
post #739

Earlier quoted context omitted.

> I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well. Because the training data is millions of hours human efforts being distilled into a cascading hierarchy of enrichment by interested parties without providing attribution or compensation?

> without providing attribution or compensation? many teachers also taught many students over the course of history, and very few would eventually pay any compensation or even attribute their financial (or career) outcomes to the teachers. What made model training different?

Huh? In your example these many teachers were paid for teaching these students and were able to make a living off of teaching without the students compensating or attributing their financial (or career) outcomes to the teachers while now we have a system where we are expected to pay a monthly amount to a corporation that has inhaled all human knowledge without any financial compensation to the people who created, managed or maintained this knowledge. The effective difference being that our knowledge, which used to be a means of income, has now become a subscription cost.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#746
post #739

Earlier quoted context omitted.

> I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well. Because the training data is millions of hours human efforts being distilled into a cascading hierarchy of enrichment by interested parties without providing attribution or compensation?

> without providing attribution or compensation? many teachers also taught many students over the course of history, and very few would eventually pay any compensation or even attribute their financial (or career) outcomes to the teachers. What made model training different?

Because the model is owned by a for profit corporation, ran and owned by total psychos and the (presumably) competent teacher is a friendly uncle?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#747
post #706

All of these accusations could be true. But there's also no way for a company to casually claim "No, we did not train on your data", without verifying all the knobs the user might have turned to enable or disable data sharing. I just don't understand getting the pitchforks out because a company did not give an answer immediately. And the effect such data entering training would have affected the output is even less c…

At the very least, a company shrugging and saying it’s impossible to know whether academic plagiarism had occurred is a claim that needs to be justified, not taken at face value.

And even if so, it should be on the company to design systems to avoid academic plagiarism and offer the right transparency. It shouldn’t suffice to say “we don’t know what went into the model, when, or how” —- that’s a solvable problem that an accountable company can satisfy.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#748

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training." This is the third day of total hysteria that is based on nothing of substance. Move on folks.

Where did OpenAI say this?

And it still leaves open the question of the prompt itself, which can just as easily encode information about the same knowledge.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#749

Earlier quoted context omitted.

> but right now there's no credible evidence, only claims. since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried. you're being naive

OpenAI has said that their models were definitely not trained on any of Buckmaster's sessions after July 3rd (from https://archive.ph/75WcF ); likely they found that's when he switched the "allow training" setting off.

Very strangely, they said something directly contradictory. initially that it was impossible to rule out whether bucmkaster’s conversations went into training data. Now they claim the opposite with full confidence.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#750

Most scientific breakthroughs are simply a continuation of previous work. I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did. Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforti…

Please stop with this psychoanalyis and mind reading with AI and human fear. It’s a thought terminating cliche at this point.

In this case, it’s much simpler and more human. Largely between two humans — buckmaster and Bubeck. The interesting question is what the role of contribution and credit for research in the ai world.

The capabilities of AI aren’t even in question in this case.

Post reply on HN