Live data from Hacker News

On the Navier–Stokes Millennium Prize Problem

openai.com

941–950 of 1001 posts

Re: On the Navier–Stokes Millennium Prize Problem

#941

Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.

If there is any truth to this timeline, then presumably it just means an additional 2 weeks of RL training on Astra.

"Twice as capable in mathematics" just means they found some problems that Astra couldn't solve, or make progress on (who knows how they chose to define "capable", or "twice as" for that matter), then put those 2 weeks of training in to focus on those gaps.

At this point, focused on their IPO, the best way to interpret OpenAI press releases is "what is the least this can mean, without being an actual lie". They are not shy - if there was a more impressive claim they could make, they would have made it.

Re: On the Navier–Stokes Millennium Prize Problem

#942

Earlier quoted context omitted.

> Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra. Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?

I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.

An Astra sized model takes months to train - they are not saying 2 weeks to train from scratch. The only only interpretation of this "2 weeks" claim that is consistent with reality is that they mean 2 weeks of additional training on top of whatever their starting point was, so it's more like this:

|--- N months of base model training -->|--- X months of post-training -->(Astra?)|-- 2 weeks more training (on Navier-Stokes adjacent material, perhaps)--> this "new" model

Re: On the Navier–Stokes Millennium Prize Problem

#943

"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?

Means the authors of the paper are "authors".

Re: On the Navier–Stokes Millennium Prize Problem

#945
post #233

Earlier quoted context omitted.

It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question If OpenAI is training mode…

If you're alleging that they don't actually have a highly capable model and the work they're attributing to it was actually plagiarized from human mathematicians, well, that would be big if true, but I'd be inclined to take the other side of that bet. With most previous splashy AI results, others have subsequently used the model to do other things around the same difficulty level. Also, it would still be necessary to…

You don't need to speculate here, since much of the story is not being disputed.

The recent ground-breaking work on Navier-Stokes "blow-ups" was done over a period of years by mathemtaticians Diego C´ordoba and Luis Martınez-Zoroa.

NYU professor Tristan Buckmaster and Anthropic employee (& mathematician) Levent Alpoge took the above work as a starting point, and over a year with LLM assistance developed a blow-up proof under certain conditions.

Buckmaster: "We used several LLMs throughout: Anthropic’s Claude, OpenAI’s Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments."

Buckmaster says he thinks that Martınez-Zoroa, whose work this all builds on, deserves the Fields Medal for his work.

OpenAI claim that on Sept 1st they heard a rumor the problem has been solved (which happened on August 15th), and then decided to re-solve it themselves using a 2-week old model, then later reached out to Prof. Buckmaster and Levant to come to some agreement to co-publish.

OpenAI: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ". In other words, not only did they deliberately choose to tackle a problem they heard had already been solved (in turns out only partially solved), but they may have done so using a model that was aware of the successful way to attack the problem.

It seems there are three potential scandals here:

1) OpenAI by their own admission chose to try to scoop mathematicians who they had heard had already completed a proof

2) OpenAI may have used a model that had seen "de-identified" messages indicating the direction to take

3) An OpenAI employee essentially threatened to "ruin the career" of the NYU professor who had been working on this if he did not cooperate with them

The direct plagiarism possibility, 2), while it should be a warning to anyone using OpenAI's models, doesn't need to be true for OpenAI to have benefited from the researcher's work. It's enough that they heard Navier-Stokes had been solved and could then go out with their swarm of 10,000 agents and $20M of compute to hunt out the latest research and brute force it.

Magnus Carlson once said that if he wanted to cheat all it would take would be for someone to indicate to him (a wink from someone in the audience perhaps) when a position warranted more time to be spent on it (because there was something important to be found if he did). It seems that, at absolute minimum, this is what OpenAI did here, although in context of math this is not cheating - the "wink" was a rumor, originating from who knows where, that a proof existed (but had not yet been published) and therefore there was potential to rush in and scoop rights to publish or co-publish.

Re: On the Navier–Stokes Millennium Prize Problem

#946

Is this truly the beginning of the AGI era? Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution. If this is what anyone calls 'slop' then slop has no meaning. I'm all for it on the use case of solving mathematical breakthroughs!

No AGI here, it's just extreme brute forcing. AI doesn't understand fluids dynamics, it just slops it's way to the solution.

Re: On the Navier–Stokes Millennium Prize Problem

#947
post #630

Earlier quoted context omitted.

Most experimental physics and other natural sciences are strongly driven by their theoretical siblings, i.e. in particle research nothing gets built without a solid theoretical foundation of what you expect to find (or where you expect existing theories to break down), the same is true in other areas, no one is doing an experiment in quantum physics before they have a solid theoretical understanding of the effects th…

Most of high energy theoretical physics is very non-rigorous or even hand-wavy. I think AI isn’t there yet for such problems.

Agreed! Many people are saying AI isn't really intelligent yet because it can't come up with genuinely new things. Maybe finding a rigorous formulation of QFT / high-energy physics would be a great test for whether they are!

Re: On the Navier–Stokes Millennium Prize Problem

#948

Is this truly the beginning of the AGI era? Running agents and prompting excessively to produce 'slopcode' to solve mathematical problems and generate a solution. If this is what anyone calls 'slop' then slop has no meaning. I'm all for it on the use case of solving mathematical breakthroughs!

No AGI here, it's just extreme brute forcing. AI doesn't understand fluids dynamics, it just slops it's way to the solution.

Which is still something. When we didn't have a solution, a solution is more than zero. But it's not understanding, either by the LLM or by us. That is, a brute-force solution doesn't explain anything, even if it proves something. It doesn't open new doors for further understanding and further discovery.

Re: On the Navier–Stokes Millennium Prize Problem

#949

"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?

How could they confirm or deny this?
Post reply on HN