Live data from Hacker News

The Navier–Stokes Millennium Prize Problem

simonwillison.net

211–220 of 234 posts

Re: The Navier–Stokes Millennium Prize Problem

#211
post #197

Earlier quoted context omitted.

It seems completely trivial to feed sessions to their own LLM and ask it to look for various things in them, from detecting problematic use cases to finding interesting mathematical work.

let's say that they ask a single question for each session they get. they are immediately doubling the compute they need in processing and then post-processing the same session twice. nothing trivial about it. not saying they cannot feed "their own LLM" saying it isn't trivial especially at scale. if you do not trust me try it without the "at scale" part.

It's trivial and a solved issue for the companies developing frontier AI models. Obviously, you don't even need AI for searching every prompt every user has ever written to find interesting topics, but you can create automated summaries and use AI on them if you want. There is no "scaling issue" here for companies who are used to processing almost everything that has ever been written anyway.

I didn't want to insinuate that it's trivial for small companies or individuals to do big data mining at that scale, sorry if I made that impression.

Re: The Navier–Stokes Millennium Prize Problem

#212
post #24

> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ... I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years. I think giving someone hope that an answer exists might as…

You see this in many things. Look at climbing for instance, or running. Once the first v17 had been done, suddenly many people did them. Or running with a sub 2 hour marathon.

Re: The Navier–Stokes Millennium Prize Problem

#213
post #143

Earlier quoted context omitted.

1) Because AI models are 1000x better about following a problem to its conclusion than coming up with a genuinely new idea. 2) Because if we accept the facts ChatGPT only came up with its "new" idea after being told exactly what the new idea was by a mathematician (OpenAI doesn't dispute this btw). And OpenAIs story comes down to the usual "We didn't look at it, trust me bro", which is made more hard to believe becau…

I don’t know much about NS problem or Euler problem anyway. Can’t tell how significant it is to go from Tristan approach to OpenAI’s results. It is 10k agents after all with a model that is 2x-3x more powerful than Astra. So claims some intuition is his alone and the model can’t find it independently is Tristan’s belief not the reality, which I don’t find particularly convincing. I do think that Levent guy isn’t inde…

> Ultimately, the model’s capability is the real surprise factor here, the fact they could get a solution after all in 3 days, that is the cause of drama ...

And there we go, a beside-the-point answer that tries to bring it back to the PR claim. And of course, an unsupported PR claim: if the initial insight really did come from a human, that would indicate the model's capability either isn't there or didn't matter.

And sorry to say, but competitive adversarial rewards for humans + access to the answer always ends in peeking. In kindergarten. In universities. In Engineering firms. In Finance. No difference other than the justification afterwards. Humans peek in competitive settings. They just do. The fix is not "trust me bro", the fix is making sure they have no ability to do it in the first place.

And OpenAI's answer to the "did you peek at the answer?" is trust me, bro ... as per usual.

If OpenAI indeed played fair, then this is a serious "dick move", and not at all how academia is supposed to work. AND if they played fair OpenAI effectively attacked one of their customers here.

If OpenAI didn't play fair then this is stealing from their customers. Which at a different level is their whole business model (get training data from customers, who even pay for that, train better model with that data, charge more, repeat).

Re: The Navier–Stokes Millennium Prize Problem

#214
post #18

It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing). What I cannot reconcile is the timeline and the concern in this specific case. I don…

The breakthrough was on August 15th, but Tristan and Levent had been working towards it (with the help of various models) for the best part of a year.

I personally doubt that their work influenced the OpenAI result - OpenAI themselves say "While unlikely, we cannot rule out that..." - but that "we cannot rule out" is exactly the problem.

If even OpenAI "cannot rule out" the influence of their usage of ChatGPT on this layer result then my discomfort at not understanding how my own usage of ChatGPT affects its training is magnified.

Re: The Navier–Stokes Millennium Prize Problem

#215

I think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models I mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messa…

They could explain how they use data from their users "to improve our models".

In the absence of that, how can we make informed decisions about how to use their tools?

(And "you can opt out" is a weak answer IMO, because I can't guarantee that everyone I email about something has also opted out.)

Re: The Navier–Stokes Millennium Prize Problem

#216
post #197

Earlier quoted context omitted.

let's say that they ask a single question for each session they get. they are immediately doubling the compute they need in processing and then post-processing the same session twice. nothing trivial about it. not saying they cannot feed "their own LLM" saying it isn't trivial especially at scale. if you do not trust me try it without the "at scale" part.

It's trivial and a solved issue for the companies developing frontier AI models. Obviously, you don't even need AI for searching every prompt every user has ever written to find interesting topics, but you can create automated summaries and use AI on them if you want. There is no "scaling issue" here for companies who are used to processing almost everything that has ever been written anyway. I didn't want to insinua…

I not only think it is not trivial, I know it is not solved.

Re: The Navier–Stokes Millennium Prize Problem

#217
post #195
post #182

Earlier quoted context omitted.

Surely this has nothing to do with specifically table tennis, nor that it's amateur. A better example would be mixed boys/girls sports classes in school, where the boys deliberately hold back as to not injure/scare the girls. It's a pretty obvious and human thing not to go out and completely destroy a much weaker opponent. We're social animals after all. There also may be an element of energy conservation, there's ob…

I think the difference is that, in the table tennis example, it happens unconsciously and you can't help it. But you're right, it likely happens in most sports. The mixed boys/girls classes example is, as you said, deliberate. And I too remember holding back on purpose in such situations when I was a kid. With the arcade, no one was deliberately holding back. I'm sure of it.

Is it that you’re holding back against weaker opponents, or that you’re more excited and engaged when you’re playing against someone that forces you closer to the edge of your ability?

Re: The Navier–Stokes Millennium Prize Problem

#218
post #142

Earlier quoted context omitted.

Why would I trust any of them?

It's either that, not use an LLM, or local-host an LLM. If you can local-host one then great, this is a compelling reason to do that too. A lot of people aren't in that position.

Everyone is in a position not to use an LLM. They may not want to inconvenience themselves on someone else’s desires though. I think at that point you should evaluate your working relationship.

I don’t use them. I have evaluated them and the trade off is too detrimental to me and society in the long term.

Re: The Navier–Stokes Millennium Prize Problem

#219
post #216

Earlier quoted context omitted.

It's trivial and a solved issue for the companies developing frontier AI models. Obviously, you don't even need AI for searching every prompt every user has ever written to find interesting topics, but you can create automated summaries and use AI on them if you want. There is no "scaling issue" here for companies who are used to processing almost everything that has ever been written anyway. I didn't want to insinua…

I not only think it is not trivial, I know it is not solved.

I don't trust your judgment, it's in my opinion even hilarious given that we're talking about companies worth almost a trillion dollar (4 trillion in the case of Google). Be that as it may, it was nice chatting with you!

Re: The Navier–Stokes Millennium Prize Problem

#220
post #207
post #75

Earlier quoted context omitted.

It's extremely unlikely that there is anything interesting or novel in the optimisation of a small SaaS SQL

But what if they added "be interesting and novel" to their prompt? Did you consider that?

I did not.

Having spent considerable amounts of time undoing the effects of overconfident junior programmers who decided to be interesting and novel on small business code bases, I guess if we're going to replace juniors with AI we may as well ask for the full experience.

Post reply on HN