Live data from Hacker News

On the Navier–Stokes Millennium Prize Problem

openai.com

181–190 of 1001 posts

Re: On the Navier–Stokes Millennium Prize Problem

#181
post #130
post #58

Earlier quoted context omitted.

Buckmaster: > "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer." OpenAI (i.e. this OP): > "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helpe…

Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data? This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there ma…

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

(I work at OpenAI.)

Re: On the Navier–Stokes Millennium Prize Problem

#182

Earlier quoted context omitted.

No, its a pure outrage. I defended OpenAI til today. I now affirm they must be totally destroyed, burned utterly to the ground.

Why the rage? Weather an individual or a company found the solution (stolen or not) they both used AI to come get the solution. We have AGI and the intelligence abundance is going to be amazing for everyone in the future.

> Why the rage?

I think it's the dishonesty, the threats of "destroying the career" of one of the mathematicians, and the request that one of the authors disavow *the other individual he was working with for the last 1-2 years* so he could claim the Clay prize as part of OpenAI.

It doesn't surprise me that OpenAI's team were surprised he'd turn it down; it shows that they just assume everyone else is as slimy as they are.

Re: On the Navier–Stokes Millennium Prize Problem

#183

Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.

Not only that, but it used 10k agents coherently over 88 hours to come up with the proof. This is a significant advance.

Re: On the Navier–Stokes Millennium Prize Problem

#184

It seems like some other mathematicians (not affiliated with openAI) have also (or close to) done this. A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774

The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys: > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin…

I´m waiting on the other side version, because I know there is no justifiable way to talk to a person like they did.

Sociopathic behaviour.

Re: On the Navier–Stokes Millennium Prize Problem

#185

Earlier quoted context omitted.

341k lines of lean without comments

Had no idea this was what lean looked like- that's mind blowing. I'm not even sure how someone would critique this if they wanted to

The point of lean proofs (as it stands) is simply one bit of information: that a given mathematical statement is indeed true.

It's a way to be absolutely certain (modulo bugs in the lean kernel) that a proof you came up for a statement is indeed correct. It is really not meant to be analyzed, much less now that they are fully llm written.

Re: On the Navier–Stokes Millennium Prize Problem

#186
post #130
post #58

Earlier quoted context omitted.

Buckmaster: > "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer." OpenAI (i.e. this OP): > "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helpe…

Why can't they rule it out? Is even OpenAI unable to track the provenance of all of their training data? This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there ma…

A careful reading of "we cannot rule out that de-identified data derived from their usage of our products helped improve our models" could be saying that yes they trained on it but they don't know if that training data resulted in an "improvement" to the model. That is, they can't rule out that the only reason the model found this solution was because it had been trained on this approach.

The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".

Re: On the Navier–Stokes Millennium Prize Problem

#187
post #111

Earlier quoted context omitted.

I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?

They are highly capable, no doubt about that, but: 1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence. 2) If the threats are to be believed, it is concerning how far they are willing to go…

"2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would believe their products and credibility would take by themselves but here we are."

Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.

And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.

Re: On the Navier–Stokes Millennium Prize Problem

#189

It seems like some other mathematicians (not affiliated with openAI) have also (or close to) done this. A statement was posted about the surrounding events by one of the them: https://cims.nyu.edu/%7Etristanb/statement.pdf Also Terrence Tao's post: https://mathstodon.xyz/@tao/117233528517340774

The allegations of contamination (using Tristan and Levent's work) aren't very well evidenced, but this behavior by OpenAI (from the authors' statement) makes them seem like the bad guys: > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin…

(To help people keep track: that's OpenAI (allegedly) threatening Tristan Buckmaster (NYU) to remove Levent Alpöge as a co-author. Alpöge is a well-known[0] Anthropic mathematician).

[0] https://hn.algolia.com/?query=Alpöge

(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)

Re: On the Navier–Stokes Millennium Prize Problem

#190

Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.

The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”. [1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/

Compute will always be the bottleneck even if this were true.
Post reply on HN