Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

821–830 of 861 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#821
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

[deleted]

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#822
post #631
post #606

Earlier quoted context omitted.

OpenAI says > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models That basically means, we don’t know, and we hope the model didn’t look up user conversations, and the best thing we can do is hope. That’s seriously disgusting. I can understand why on a technical level why perhaps it is impossible to answer what exactly the model had access to,…

Every other paper in existence has been ingested with 99% of writers not knowing it will be retroactively used for training. But session data which is disclosed as being used in terms of service is disgusting? If the work is duplicative/derivative then the preprints they put in sessions can be shown by the users and we can see.

What's disgusting is that they refuse to acknowledge what they and only they know: were these chats fed in to the new model that found these results?

Nobody would be disgusted at following stated policy, that's doesn't make sense, without first objecting to the policy. But you're bringing up distractions from the actual concerns: did OpenAI use the private chats and why won't they confirm or deny it?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#823
post #724

Earlier quoted context omitted.

> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work > If true, that's generous and beyond the level of generosity one should expect "We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so gen…

The two proofs are structurally very different, and don’t even prove the same conjecture. It’s becoming very clear that OpenAI did not steal anything here.

If it was "very clear", OpenAI wouldn't be threatening the researcher with "it's very bad for your career" or state "we can't tell you if it was trained on user input".

Additionally, according to the researcher, OpenAI's proof follows the same approach they used, and which was largely unused in academia, but OpenAI claimes they arrived at it immediately.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#824
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

This is hard to argue without a fine understanding of how much insight OpenAI had about the stab at the problem from the "public rumor" alone.

If there was any sort of coarse insight that "they're trying to solve it this way", then both those things can be true:

- The massive amount of compute from OpenAI re-discovering Buckmaster's work solely from the coarse insight (and solving the rest as well).

- OpenAI still acknowledging they basically scooped the coarse insight using compute, and they're willing to credit Buckmaster.

Buckmaster says "Concretely, what Levent and I did was to take the Cordoba and Martinez-Zoroa program, which achieved blowup results with rough forcing, and, with a great deal of help from LLMs, push it to smooth forcing and to the incompressible Euler equations".

Could that simply be the prompt they used at OpenAI? How "stolen" would the proof be in that case?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#825
Of course OpenAI used these guys' training data. This is part of their ratchet strategy: if you can monopolise the creation of new knowledge, you win. The moat they are building is user data, which a billion people are now throwing at them every day. They are solving difficult math problems using it. The only way to get the this data and these techniques back out is by distillation - expensive, slow work.

This moat ensures that Anthropic and OpenAI will remain at the frontier - no one else can train a super-intelligent model because they lack the trove of user data.

My main takeaway from reading this is - why do we always fall for this same trap? Every single time we allow a software company to build a monopoly in pretty much the same way. Microsoft, Oracle, AWS - They have a cool tech and we let them run away with it.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#826

Of course OpenAI used these guys' training data. This is part of their ratchet strategy: if you can monopolise the creation of new knowledge, you win. The moat they are building is user data, which a billion people are now throwing at them every day. They are solving difficult math problems using it. The only way to get the this data and these techniques back out is by distillation - expensive, slow work. This moat e…

There is no trap you fall into that you can get yourself out of, because you have no say in the matter.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#827

Earlier quoted context omitted.

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves. You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training). Then…

You can just spin up deep research agents that ingest many sources at once to produce reports that don't replicate any one source too much. Since agents compare against sources they provide across-source analysis - what is the distribution of positions on this topic, is it debated or settled. Not truth, just summarizing, but I think this would be very useful for training.

Besides reporting on search sources you can also run the same queries on multiple LLMs closed book mode, and judge their distribution as well. It helps a lot if models are more aware of their knowledge holes. Scale it up for billions of topics if you have the pockets, the DR data is copyright free.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#828

Earlier quoted context omitted.

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

That was.. obvious? How are you shocked? Honestly, how insane must the suspension of disbelief on this site be, that anyone here is shocked?

You’d be surprised how many blind spots this site has.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#829

Earlier quoted context omitted.

I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context. As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus. So is…

The chats the professor had are not generic knowledge. And yes of course you still need to cite Newton and Leibniz depending on what result you want to mention. What’s allowed to be not cited are not status quo, the term you’re looking for is “folklore” results aka results that have been around so long that 1) nobody knows who came up with them or 2) everyone knows who came up with them. The second point: if you say…

Even if AI used the result, AI pushed it to the finish line while Levent and Tristan did not. But I understand the approach was different, the information leak was only that it was "doable".

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#830
post #794

Earlier quoted context omitted.

if you assume the choice of route is a uniformly distributed random variable, yes. but this assumption does not seem consistent with "Almost nobody else I know of was working on it", from Tristan's quote. nor with "It is not the direction one arrives at in a few days by giving a model the problem statement".

It doesn't really make sense what he said. Everyone expects N-S to have blow-up. If you are going to try to show blow-up, C and D are strictly easier than A and B. So Tristan was not clear about what detail of "route" he is talking about, because what he literally said can not be it.

yeah, it looks like i don't have enough knowledge of this subject to discuss this. but i'd be interested in reading what more knowledgeable people have to say about it -- Tristan's work, how different from others' it was (they claim no one else was following the same path), and how it compares with the AI's
Post reply on HN