Is this the one that was allegedly based on someone else's actual work & prompts? https://news.ycombinator.com/item?id=49605915 https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd... https://cims.nyu.edu/~tristanb/statement.pdf
Yes, that was the allegation last night. I work at OpenAI, though not on the team that did this, and my understanding is: - we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular in…
On the Navier–Stokes Millennium Prize Problem
881–890 of 1001 posts
Re: On the Navier–Stokes Millennium Prize Problem
#882Buried under the drama is the fact that OpenAI is claiming that an internal model they’ve been training for less than two weeks is more than twice as capable in mathematics as Astra, which was only made public a week ago. Even if this improvement is limited to mathematics, that is an astounding feat.
The implication from their last couple of published articles[1][2] is that they think they’ve achieved “recursive self improvement”. [1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
Re: On the Navier–Stokes Millennium Prize Problem
#883"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Re: On the Navier–Stokes Millennium Prize Problem
#884Earlier quoted context omitted.
I worked at OpenAI previously, but don't know any of the people involved in this. My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet". It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that…
They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.
Re: On the Navier–Stokes Millennium Prize Problem
#885Re: On the Navier–Stokes Millennium Prize Problem
#886Terence Tao has some observations that seem to be directed at this, https://mathstodon.xyz/@tao/117237320796901560 > "We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising researc…
Tao: Open math problems being non-renewably mined by AI - https://news.ycombinator.com/item?id=49616968 - Sept 2026 (276 comments)
Re: On the Navier–Stokes Millennium Prize Problem
#887the context here is super important, for those who haven't seen it yet. OAI maybe just trained on a real researchers solution and then celebrated having scored the goal unassisted save for the brief commentary at the bottom of this blog post. Here's the other side. https://x.com/rynorhn/status/2097223532438487463
That other researcher was working on a smaller related problem. He was also using LLMs to do it, so either way most of the credit goes to the LLM here.
Re: On the Navier–Stokes Millennium Prize Problem
#888Earlier quoted context omitted.
>we did not read any private chats The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
If they opted out of training, then we definitely did not train on them. If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof. Reasons for my doubt: - I know…
are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.
Re: On the Navier–Stokes Millennium Prize Problem
#889Earlier quoted context omitted.
Not a counterpoint per se, but I burned $50k recently on a much more modest math problem (result already known, just thought I had a sketch of a more interesting proof), and the LLM thought it had proved it within those bounds but had instead subtly fucked up the Lean definition. Take from that what you will. Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or…
How can you afford to burn $50k on an already solved problem?
> How can I afford?
Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.
> Already solved
That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.
Re: On the Navier–Stokes Millennium Prize Problem
#890"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.