Live data from Hacker News

On the Navier–Stokes Millennium Prize Problem

openai.com

881–890 of 1001 posts

Re: On the Navier–Stokes Millennium Prize Problem

#882

Terence Tao has some observations that seem to be directed at this, https://mathstodon.xyz/@tao/117237320796901560 > "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising researc…

Related ongoing thread:

Tao: Open math problems being non-renewably mined by AI - https://news.ycombinator.com/item?id=49616968 - Sept 2026 (276 comments)

Re: On the Navier–Stokes Millennium Prize Problem

#883

the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side. https://x.com/rynorhn/status/2097223532438487463

That other researcher was working on a smaller related problem. He was also using LLMs to do it, so either way most of the credit goes to the LLM here.

"It was splendid! Waiter, share my regards with the oven."

Re: On the Navier–Stokes Millennium Prize Problem

#884
post #253

Earlier quoted context omitted.

>we did not read any private chats The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?

If they opted out of training, then we definitely did not train on them. If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof. Reasons for my doubt: - I know…

> If they opted out of training, then we definitely did not train on them.

are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.

Re: On the Navier–Stokes Millennium Prize Problem

#885
post #851
post #735

Earlier quoted context omitted.

Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will. Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or…

How can you afford to burn $50k on an already solved problem?

Those are two separate questions:

> How can I afford?

Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.

> Already solved

That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.

Re: On the Navier–Stokes Millennium Prize Problem

#886

"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?

I've been saying it for a while now, but no one gives a fuck. Let me repeat it again.

THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.

"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.

I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.

Re: On the Navier–Stokes Millennium Prize Problem

#887

Earlier quoted context omitted.

> If you need privacy, then you are going to have to pay full price for those tokens (API). At this point, how can we even trust that they aren't accidentally training on those tokens too?

It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.

> It'd be corporate suicide for them to be caught violating zero-data-retention commitments

Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.

Re: On the Navier–Stokes Millennium Prize Problem

#888
post #377

Earlier quoted context omitted.

If humans can figure out to optimize to circumvent bottlenecks, I have no doubt each new bottleneck will also get routed around, just now automated.

We are not in an everything-has-an-API world yet, and it'll for sure take some time to get there.

I'd argue we've been in an "everything-has-an-API" world for a long time now — it's just that discoverability of said APIs is still crap.

Re: On the Navier–Stokes Millennium Prize Problem

#889

Earlier quoted context omitted.

Why is it unlikely?

Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user. It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes. We also don't know if the authors unintentionally provided data to OpenAI through alterna…

> It's unknowable

Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.

> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes

It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).

Re: On the Navier–Stokes Millennium Prize Problem

#890

Earlier quoted context omitted.

I came here to say this, like... wow. I'm pretty sure at at least a few of the places I've worked that would be grounds for immediate termination.

They probably should have added the disclaimer: opinions are my own..

Yeah that isn't a magic get-out clause. I don't think I would have been immediately fired for this from anywhere I work at, but that's partly because I live in the UK.

Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.

Post reply on HN