Live data from Hacker News

The Navier–Stokes Millennium Prize Problem

simonwillison.net

41–50 of 232 posts

Re: The Navier–Stokes Millennium Prize Problem

#41
post #36
post #20

Can’t wait to see the human verifying the results and then figure out that the AI model actually cheated and the results are not correct.

the proofs were verified in Lean, so unlikely.

As long as the proofs itself are correct. How long they were this time?

Edit: at least ~600,000 lines

https://stanfordtechreview.com/articles/openai-buckmaster-na...

Re: The Navier–Stokes Millennium Prize Problem

#42
post #13

I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect. That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating…

OAI's side of the story is that they discovered the approaches (and indeed solved problems - Euler equations vs NS equations) differed. They then offered Buckmaster lead authorship of OAI's proof, without Alpöge. But they never demanded that Alpöge be stripped of coauthorship on resolving the regularity of the Euler equations. At least that's the claim.

https://xcancel.com/SebastienBubeck/status/20973794116915163...

Re: The Navier–Stokes Millennium Prize Problem

#43
>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?

But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."

Re: The Navier–Stokes Millennium Prize Problem

#44
post #18

It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing). What I cannot reconcile is the timeline and the concern in this specific case. I don…

I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.

Re: The Navier–Stokes Millennium Prize Problem

#45
post #37

I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messa…

If they have zero-retention, then it is not possible. So what they are saying is that they don't have zero retention.

But why is this news? This is clearly described in their ToS and in the Settings to improve their models.

Re: The Navier–Stokes Millennium Prize Problem

#46
post #15

LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence»,…

Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.

However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.

Re: The Navier–Stokes Millennium Prize Problem

#47
post #24

> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ... I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years. I think giving someone hope that an answer exists might as…

> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.

If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.

Re: The Navier–Stokes Millennium Prize Problem

#48
post #29
post #22

Earlier quoted context omitted.

Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.

If that is the case, why on earth would you use it in any professional setting?

Depends on your profession? I sometimes work on open source code as part of my professional duties. Nothing that goes on there is necessary to keep private.

Re: The Navier–Stokes Millennium Prize Problem

#49
post #2

> My two favourite hypothetical questions regarding this used to be: > If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe h…

LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.

Re: The Navier–Stokes Millennium Prize Problem

#50
post #29
post #22

Earlier quoted context omitted.

Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.

If that is the case, why on earth would you use it in any professional setting?

Aren't business consulting firms even worse? They are explicitly for business and there are known cases where they shared confidential information of one of their customers with another one.
Post reply on HN