Earlier quoted context omitted.
If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained
Navier-Stokes – Tristan Buckmaster [pdf]
451–460 of 862 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#452From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
Opt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on.
So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#453Earlier quoted context omitted.
If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained
Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#454Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems. "Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."
It's part of their TOS that they can train on users' private chats.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#455IMO the parsimonious answer seems to be OpenAI has a pretty good model (because it did finish) and stole someones work... and threatened them over it. TBH all OpenAI need to do is solve another millennial problem and none of it would matter - people expect them to behave heinously regardless - but if they have generalized superhuman math model... well I guess they're allowed io.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#456Earlier quoted context omitted.
if they do not deny training on them, they can't deny plagiarism.
By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.
as bad academic conduct you may steal someone else's unpublished work, work on it yourself for a bit, and then publish it as your own work. and then threaten the original author!
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#457Earlier quoted context omitted.
Or, given how dedicated he appears to be to the company, a promotion and a raise.
Dark take, but I really hope not. If anything it would be a good opportunity to buy some goodwill by washing themselves from all the alleged shadiness so far.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#458Earlier quoted context omitted.
Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#459Earlier quoted context omitted.
It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course. Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were g…
They are not training a whole model in a matter of days
Model editing to remove PII that slipped through, all sorts of things of that sort.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#460From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.