Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

441–450 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#441

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#442

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

That was.. obvious? How are you shocked? Honestly, how insane must the suspension of disbelief on this site be, that anyone here is shocked?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#443

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable

I have terrible news about how literally every leading AI model was trained

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#444

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

It’s also very shortsighted to stiff a customer like that. Why would I trust them with my data and ideas?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#445

Earlier quoted context omitted.

No, they use that wording because they are mocking this tweet from an Anthropic employee: https://xcancel.com/_sholtodouglas/status/209721833169057800... . I don't know why.

That is even worse. Back to 5th grade I guess

Look at the President and the US government. This is society now.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#446
post #415

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

Does opting out matter? "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI. https://x.com/markchen90/status/2097400166554993041

My understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#447

Earlier quoted context omitted.

if they do not deny training on them, they can't deny plagiarism.

By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.

Just replace the model with a human student.

"Training" on textbooks => fine

"Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#448
post #441

Earlier quoted context omitted.

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.

Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#449

Earlier quoted context omitted.

if they do not deny training on them, they can't deny plagiarism.

By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.

Isn’t that one of the most salient and straightforward argument against LLMs?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#450
post #21

>I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. If you think these companies are not training on your prompts you are incredibly naive. These models were built by stealing and pirating literally…

My guess is that if you opt out of your data being included then they honour that instruction, but I wouldn't bet my business on it.
Post reply on HN