Earlier quoted context omitted.
Anthropic too, and evidently others. Amazon were recently confirmed to be doing the same thing: https://www.404media.co/we-tracked-a-shipment-of-rare-books-...
And that’s bad, right?
Navier-Stokes – Tristan Buckmaster [pdf]
461–470 of 862 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#462Re: Navier-Stokes – Tristan Buckmaster [pdf]
#463Earlier quoted context omitted.
If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
It’s also very shortsighted to stiff a customer like that. Why would I trust them with my data and ideas?
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#464Earlier quoted context omitted.
Opt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
Edit: the parent comment now seems to better reflect the below. That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on. So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning…
The way for people or companies or universities to control their data and information is to keep it on their own computers.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#465Re: Navier-Stokes – Tristan Buckmaster [pdf]
#466From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#467From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler…
> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#468OpenAI's statement: We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped i…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#469From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
OpenAI should release the agent log, including CoT.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#470Earlier quoted context omitted.
By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.
Isn’t that one of the most salient and straightforward argument against LLMs?