Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

611–620 of 848 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#612
Who cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal.

I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need.

I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems would be solved under the axioms that again, are human defined, under brute force. The real question is, are those axioms the bottom level, and if they are not, who is going to set the new aximons and can we understand them.

I've no doubt there's useful breakthroughs that will happen, but I think it should be remembered that the method being used is still a heuristic brute force approach is being very narrowly applied against axioms and math and physics which humans described in the first place, and almost undoubtably has errors and/or is not complete.

Its a great example of the power of LLMs but its not 'we've solved science now just pour more tokens in'

Re: More questions about whether researchers can trust OpenAI with unpublished math

#613

Earlier quoted context omitted.

We checked and determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo. See: https://…

You guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having these worries.

Option 1 is opting out manually. Option 2 is business / enterprise plans, which opt out by default.

Any ideas of things we could do to make it clearer?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#614

Earlier quoted context omitted.

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

I heard they also tried to strong-arm them into removing the name of their collaborator who happened to work at a different company (Anthropic)... I haven't looked into it myself, but if true, that seems incredibly scummy.

That's also incorrect.

My understanding is that they asked the independent researcher to improve OpenAI's AI generated proof and be the lead author of the paper to publish OpenAI's result.

This is the paper where they did not want the Anthropic employee collaborating. Not their work.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#615

[dead]

How about the fact that it almost certainly did not happen? I read today that OpenAI after investigation was able to categorically rule out that usage data from before beginning of July could have affected the system that was used.

Well, it takes a lot of faith to give that "almost certain" levels of credence.

What do you think about the results of people investigating themselves for wrongdoing as a general matter?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#616
post #411

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

If they have solved hundreds of open problems in math, why are they publishing results for the ones other mathematicians happen to be working on at the same time? Why not the others?

Well I'm sure if they find a millennium prize problem that no mathematician has worked on recently they will get right on publishing that.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#617

Who cares. the biggest thing about this is that its still brute force in a verifiable domain, and that it was still a human set goal. I also don't believe it much practical use, unless I'm mistaken, approximations of Navier stokes have been available for a long time to whatever precision you need. I'm not a complete disbeliever by any stretch , and also a complete amateur, but it was inevitable that these problems wo…

Just to point out re: navier stokes - what was being proven was not a solver or approximations for it, but showing specific circumstances under which it actually returns incorrect (or numerically unusable) answers. Which had been suspected but wasn't known for certain

Re: More questions about whether researchers can trust OpenAI with unpublished math

#618
OpenAI's ethical and reputational own-goal aside, my big takeaway is that it seems that:

if I'm using Codex to develop some new algorithm (in any space), OpenAI appears to be training its model on my code sessions

anyone using that model (OpenAI or a competitor) might be able to receive from the model a solution that is similar or the same as the one I developed, emerging from the training data

Re: More questions about whether researchers can trust OpenAI with unpublished math

#619
post #441

Earlier quoted context omitted.

So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.

The claim that OpenAI somehow used the mathematicians' ideas to leapfrog them seems unsupported at this time and IMHO it was irresponsible to bring it up because credulous people will immediately believe that narrative. And from my perspective, if some math folks typing in a few questions to OpenAI provides sufficient training data for OpenAI to solve a big problem... that's amazing! A few conversations/prompts out o…

I'm curious how many other 300 billion output tokens OpenAI has "paid for" that have resulted in no breakthroughs.

Either they had a pretty good idea that investing this type of money in that compute on a model in training would lead to these specific results, or they gambled with other people's money.

I want to hear about the gambles and expenditures they don't brag about. In America's energy economy, there's finite resources to expend.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#620
This is just mental illness at this point. I don't blame the mathematicians that have found a way to get attention from the mainstream press for once, but we should not fall for it here.

1. No one but OpenAI has produced a proof of NS so these accusations of plagiarism are pretty embarrassing. It reminds me of the line from the Social Network: "If they invented Facebook then why didn't they invent Facebook?". If these people proved NS before OpenAI where is their proof?

2. If they plagiarised Andreas Thom then why was his initial response to praise the proof and talk about how different it was from his own attempt? It's only now that it is clear that no one bothers checking these things that suddenly his story changes.

Post reply on HN