Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

561–570 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#561

Earlier quoted context omitted.

From the company that likely committed federal crimes by hacking Hugginface

One thing that wasn't obvious to me or adults around me when I was younger: most laws define whatever acts a law punishes as individuals commiting to it, not as situations manifesting anyhow. It's not a murder just because someome died hit by a bullet you fired, but you have to have personally decided to kill that person leading to their death[1][2]. OpenAI's LLMs are not humans, and neither is the company. So by thi…

> OpenAI's LLMs are not humans

Neither are guns. Which is why we punish the person shooting the gun and not the gun.

Industrial equipment, which is how i would classify LLMs, hurting people is nothing new. The relevant questions are:

- did someone intend it to happen?

- was someone negligent in taking reasonable steps to prevent something foreseeable?

The justice system doesn't punish people for legitimate accidents. e.g. if you are shooting at a shooting range, take all reasonable precautions, but someone was hiding behind the target, you are probably not guilty even if you shoot the guy.

As far as openAI goes, the logic is the same. The question is, was it intentional, was it unintentional but reasonable precautions weren't taken or was it truly an accident?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#562

Earlier quoted context omitted.

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

[dead]

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#563

Earlier quoted context omitted.

OAI doesn't need to mention Buckmaster's name directly in a prompt. They just need to select a basket of sessions that is guaranteed to contain Buckmaster's and then direct the LLM to attack only a specific method/angle. This is trivial to do while maintaining plausible deniability about not using his work.

what reason do we have to believe that they did this? both things were proved by AI, isn't it logical that they could have very similar approaches? it is common that multiple people essentially simultaneously prove/invent the same thing I see zero evidence of wrongdoing

This depends on what "proved by AI" meant.

Was that a one shot prompt? or something guided by human, step by step?

If that's the later, it won't use the same approach when not guided by the same human.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#564
post #429
post #403

Earlier quoted context omitted.

it's very possible they only had to use the massive compute budget because they were trying to plagiarize his work before he published it though, e.g. autonomously do things in ~7 days what he had likely been thinking about for ~1 year.

This doesn't really make sense. You don't need massive amounts of computing to plagiarize something. The most nefarious explanation seems to be that they got wind it was possible to solve NS via LLMs and perhaps a small nudge in the right direction.

wasn't it reported elsewhere that they used the equivalent of $22M (street) in Astra tokens? obviously it's not the same when you own the machinery but still.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#565

Earlier quoted context omitted.

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good…

I agree for most of the people in the story except Buckmaster himself seems to have been supplying real ideas.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#566
post #531

Earlier quoted context omitted.

It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training. Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this. OpenAI have petabytes of data, all anonymized. It could take months to say…

Frankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.

I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced somehow

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#569
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

its what all LLM users have all been doing, profiting off others' IP through a number cruncher, while relinquishing their own

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#570

Earlier quoted context omitted.

I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.

It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data? What surprises me is they're not more boldly/plainly lying about it.

I think they are pretty clear that they train on some prompts, given they sell the ability to be excluded
Post reply on HN