Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

481–490 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#481

Earlier quoted context omitted.

> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained

I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point. Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.

Either are possible.

They have been collaborating on this solution for a year, and Astra was trained in February this year so it’s entirely possible the direction of their research was in the training corpus.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#482
post #232

- Tristan and his co-author (Harvard/Anthropic) developed a theoretical framework and validated it using Codex and Claude. - OpenAI did related research around similar timeframe. - Tristan claimed OpenAI offered a proposal that included dropping the Anthropic-affiliated co-author. - Sebastian (a prominent OpenAI researcher involved) denied these claims. - Tristan have no concrete evidence that OpenAI accessed their s…

Sebastian did not deny those claims. He affirmed them, and apologized.

It’s true that Tristan has no concrete evidence OAI accessed their session. It’s impossible for him to have that without OAI’s say so.

Why are you framing things so pro OAI?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#483

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

If their solutions are significantly different, as OpenAI claims, would it still be considered plagiarism?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#484

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

Here’s a broader question – how many other academics contributed to Buckmaster’s result, by way of sharing the logs of their own (failed?) attempts into OpenAI’s training data set? How should he and OpenAI go about crediting all of them?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#485

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

> did Tristan opt out

That “opt-out” thing is a dark pattern. It’s not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You can’t take back what you’ve already shared. Also I don’t think it covers all the cases that they use your data. It’s really an opt-in button for voluntarily giving away your data for training.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#486
post #451

Earlier quoted context omitted.

> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained

That doesn't make it fine. We should not excuse this behaviour just because its rampant already, especially when it comes to such a serious prize

Sure, but unless you’ve got some exceptionally deep pockets, congress has seemingly no interest in turning the fact that it’s ethically bankrupt into any practical recourse.

Ai companies got where they are by stealing all of the intellectual property from human history. It seems entirely likely that their goal is to purloin everything produced going forward as well.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#487

I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…

I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…

Everybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty.

And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#488

Is this the future we're headed towards? Where I'll be afraid to use google docs in case Google identifies value in whatever I'm writing about and snipes it if it my docs make it into the next round of model training?

Yes… “future”… I have really bad news for you regarding all your personal data stored on Google’s servers.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#489

Earlier quoted context omitted.

By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.

Just replace the model with a human student. "Training" on textbooks => fine "Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.

Presumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO).

The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work.

Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism.

Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#490
post #17

> It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. T…

the part where they didn't want the Anthropic person credited even though they deserve credit is also particularly scummy. Corporate greed over common decency.
Post reply on HN