Live data from Hacker News

On the Navier–Stokes Millennium Prize Problem

openai.com

691–700 of 1001 posts

Re: On the Navier–Stokes Millennium Prize Problem

#691

Earlier quoted context omitted.

> But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. Just because something is legal and permitted by terms of service doesn't mean it's morally right.

>Just because something is legal and permitted by terms of service doesn't mean it's morally right. What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on? Or were they supposed to not try to so…

> What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data?

Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.

Re: On the Navier–Stokes Millennium Prize Problem

#692

Earlier quoted context omitted.

"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a num…

All of those statements sound true, based on what I've heard. - "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input - it's all true that a team worked on this, a bunch of compute was burned, and the problem was…

I'm confused, your employer very directly stated that they are unable to confirm that the model was not trained on the conversations.

Re: On the Navier–Stokes Millennium Prize Problem

#693

So what took an autonomous agentic system using a significantly more powerful internal model, totaling multi-millions of dollars of compute in training and inference, was likely to already be solved by a team of a few humans with an orders of magnitude smaller LLM budget, had OpenAI not been foaming at the mouth to jump the shark and claim “AI solves Millenium Problem.” Also it sounds like the human research effort s…

Totally. Anthropic is like a village cottage shop who was just like chilling until big bad OpenAI came in

anthropic has nothing to do with it

Re: On the Navier–Stokes Millennium Prize Problem

#694

Earlier quoted context omitted.

All of those statements sound true, based on what I've heard. - "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input - it's all true that a team worked on this, a bunch of compute was burned, and the problem was…

>> I'm not sure how any of this provides evidence that OpenAI took any of their work. Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).

> OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published.

I don’t think they’re too concerned about appeasing you, enraged_camel.

For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.

Re: On the Navier–Stokes Millennium Prize Problem

#695
post #254

Earlier quoted context omitted.

>we did not read any private chats The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?

If they opted out of training, then we definitely did not train on them. If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof. Reasons for my doubt: - I know…

But you can say (with the cooperation of the parties involved, of course) if any of the preliminary work that the other researchers did was part of the dataset. It is possible to be more transparent than you are being.

Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.

Re: On the Navier–Stokes Millennium Prize Problem

#696
post #596

Earlier quoted context omitted.

It's denial and coping. Most people i see show this tendency around AI which is also why it 's easy to be far ahead of most population nowadays

I'd say that believing to be "far ahead" is much deeper kind of coping.

how i wish so..

Re: On the Navier–Stokes Millennium Prize Problem

#697

Earlier quoted context omitted.

Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too. If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".

what counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?

It's quite well explained here[1], which is linked from the Privacy section of their plan overview[2].

Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.

You'd have to take their word, but that goes for anything in life.

[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...

[2]: https://chatgpt.com/pricing/

Re: On the Navier–Stokes Millennium Prize Problem

#698
post #585
post #560

Earlier quoted context omitted.

At the scale at which these models are now, regardless of whether they are proprietary or open weight or list their training datasets, there are hundreds of billions of works that have gone into trillions of parameters, each one providing tiny perturbations in some tiny fraction of the weights. It is probably impossible to attribute provenance to any specific input (which is also why the courts' finding of Fair Use i…

There are two different things: - was item X in the training data - did the inclusion of X in the training data lead to Y I understand why the second is hard, but why is the first one hard?

Yep, the last part in my post was suggesting some ways we could determine if "item X was in the training data" (as well as some potential blockers for that from OpenAI's perspective.)

Re: On the Navier–Stokes Millennium Prize Problem

#699

Earlier quoted context omitted.

>If anyone has counterpoints to this I'd love to hear them! Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement. To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah Cha…

The amount of goalpost moving is insane. "Yeah it can solve Millenium problems, but can it do it with nothing more than a one sentence prompt?" Also there are proofs where the only human steering was "keep going".

Parent comment is claiming creativity and genius beyond human experts, so why not ask for a fully unassisted AI novel result? Having access to the entire corpus of human knowledge, what else such amazing entity would require to solve a hard problem by its own?

Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.

>Also there are proofs where the only human steering was "keep going".

Which ones?

Re: On the Navier–Stokes Millennium Prize Problem

#700
post #46
post #15

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." I think they should be able to unravel whether or not any sessions by Tristan or Levent went into the training data for this model.

If they could then it wouldn't be de-identified data...

The researcher could share a string from one of their conversations and OpenAI can confirm whether it exists in their training data.

Or OpenAI could just look at their code and say what it does (maybe have their AI do it if they're having so much trouble with this?)

Post reply on HN