Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

771–780 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#771

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

We were also curious and we looked further into this. We've determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. This goes beyond what we said earlier, when we were less sure. If prompts were submitted earlier than that and training was not opted out, there may be a chance they made their way into our training pipeline i…

The authors had supposedly worked on it for a year, though.

And why aim straight for scooping other researchers upon hearing rumours about their success? Normal, ethically acting, researchers would never do that.

And how about existence of non-sofic groups, which is actually the topic here?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#772
post #486

Earlier quoted context omitted.

No. 1. Buckmaster contacted OpenAI first. Not the other way. 2. Giving the $1M bounty to a human mathematician for the effort and giving him credit would be excellent PR. They had already burned much more than $1M for the generation. Adding him as author also costs nothing. Purely pragmatical. 3. “As long as he removed Alpöge” part itself is against academic honesty by all means. 4. Buckmaster rejected fame and $1M o…

There are some mixed up things in your post, maybe double check next time, especially before quoting anyone, as you really undermine your point even if you're directionally right. > Buckmaster rejected fame and $1M only because doing (3) would be wrong I doubt Buckmaster would have accepted the offer to "write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it" ev…

You’re right. I used quotes when I was really paraphrasing.

Here is the actual paragraph from the statement:

> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

Context matters in communication. In that context I understand that dialog more like: we’re powerful and you are not, do the smart thing and play along, if not I don’t have to play nice. He presented a very good “offer that he can’t refuse”. But that’s my interpretation.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#773

Earlier quoted context omitted.

The problem here is that OAI (and others) pretend or claim that this is uncharted legal territory, where in fact it is very simple. We have a machine that is fed data, and produces new data as a result. If that new data depends (in any way) on the fed data, then from a legal viewpoint it is derived from that data. Whether they anthropomorphize the operation performed by the machine does not matter. They can anthropom…

> If that new data depends (in any way) on the fed data, then from a legal viewpoint it is derived from that data. The "in any way" part is either so broad it makes everything derivative, or not, in which case things are no longer simple. If everything is derivative then it seizes to be meaningful. The words I write are derivative, I literally copied them from someone else, yet my sentences as a whole can be fully no…

> If everything is derivative then it seizes to be meaningful.

That's why we tolerate it for humans, and also because we cannot prove it. But yes, if you go too far in this, you will see legal consequences.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#774

Earlier quoted context omitted.

Would this be unethical if it was a human who heard rumors about a solution then attacked the problem, solved it and published first? Often knowing of the mere existence of a solution carries a lot of information--you would know the problem is accessible, you would expect clues in recent progress (the two Spanish researchers in this case), you would probably have a sense if the solution is a counterexample or positiv…

The problem with your counter-hypothetical is that not only is it unrealistic, it's utterly impossible. No human would be able to do in such a short timeframe what the LLM did. Part of what makes the OpenAI move so egregious is how bullying it was. It was the big guy coming along with their nearly infinite resources and squashing the little guy who's devoted a good chunk of his career to the problem.

Actually my hypothetical is completely realistic as I've been involved in such scenarios. It's unrealistic maybe for a millennium problem to come in on a rumor and still front-run but not at all for the many other problems we work on and which manifest our ethical code. If you're saying ethical rules change depending on the prize be clear about it, because I can see arguments that they change to favor either side.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#775

Earlier quoted context omitted.

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

My university has an agreement with Microsoft copilot. We can log into copilot in many ways, and it's only if you log in the correct way that you get the "Enterprise Data Protection" copilot version, with a green shield symbol. There are many ways to go wrong here!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#776
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

The irony is that OpenAI got into this trouble only because they tried to play "nice". They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list. They wanted to give Buckmaster a chance to be the one solved N-S problem. While this behavior is highly questionable, if OpenAI just published the final result without notifying Buckmaster first and simply ci…

> They told Buckmaster that he could publish the final result as the author as long as he removed Alpöge from the author list.

How is that "nice"?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#777
post #632

Earlier quoted context omitted.

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable. I suspect the main reason the community is not receiving it well is largely…

If you’re making decisions of ethical and material importance based on rumors, I’d be surprised if anything ethically palatable did happen.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#778

Earlier quoted context omitted.

> no human has broad enough knowledge and enough time to try them all. The other part is, humans don’t really want to fund other humans doing this. Very few want to be a math major; and of those that do, fewer complete a grad degree; and for those that do get grad degrees, there’s scant few research jobs; and for those who do get jobs there’s hardly any research funding to go around. There does seem to be unlimited m…

I’m assuming the reported 22 million dollars worth of tokens used to solve this particular problem is far more than what humans have paid to solve it previously. So I think you’re correct.

22x more to be exact

Re: More questions about whether researchers can trust OpenAI with unpublished math

#779

https://x.com/markchen90/status/2097400166554993041?s=20 that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

Mark isn't saying the toggle does nothing. He's saying that if you leave it on, your data can be used to help train our models. If you opt out, we don't train on your data.

Given how much PII is fed through these systems, would it being opt-out by default not be violating the GDPR by a failure to require explicit consent (or otherwise provide the legal basis for processing)? If a court decides as much, I imagine it would mean that all data harvested this way must be extracted from the models, and all instances where it would have been shared would have to be identified, which would really be something.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#780

Earlier quoted context omitted.

> We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. We're scared of big tech companies concentrating ridiculous amounts of power, destroying the communities that support and guide scientific research, without even thinking about the dangers and possible consequences, because a PR stunt is more important in the short term.

If this were true, it should be stated more clearly, than most of the criticism which seems to aim to minimize the capabilities of these models. Way more often I see >AI is a scam and steals human insight and doesn't produce anything original vs >AI is too capable/powerful and will concentrate power even more than it does already due to its capabilities The latter is rarer because it requires admitting that AI is use…

> if this were true, it should be stated more clearly

stated more clearly by who? people in social media? I don't know what your feed shows you, but if you focus on what the visible people in the math community is (and have been) saying is precisely what I said.

Post reply on HN