Earlier quoted context omitted.
This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.
Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
Navier-Stokes – Tristan Buckmaster [pdf]
751–760 of 872 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#752Earlier quoted context omitted.
Opt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
This would be a fairly insane breach of trust and common sense if true; the chain-of-thought / reasoning trace is, from an information perspective, close to a superset of the prompt and model response.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#753I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…
This Godfather-like threat in particular pissed me off as well: > I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#754Earlier quoted context omitted.
If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily ju…
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)? No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excell…
A totally reasonable pipeline may be unauditable for this purpose.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#755Earlier quoted context omitted.
OAI started working on this only after they found out it was close to being solved. They threw a team of researchers who spent sleepless nights + a ton of compute. This is not exactly healthy academic competition - it's like if you spend a year hunting for oil fields and finally find a very promising area to be explored, only to find that Exxon tapped their entire exploration unit to go all in and and find it overnig…
that is not related to the accusation that they literally stole Buckmaster's work, which seems baseless I don't agree with the oil claim analogy. this is knowledge, freely given to the world. not something hoarded by a corporation
and you can argue this is "fair use" or whatever, not the point now, the point is that it definitely makes those accusations no longer "baseless".
in addition, it is not given freely to the world, it is the knowledge of the Internet/WWW being sold back to you as a subscription service. it's not free. and it's not even "given", because they can (technically) turn off the tap at any moment and you don't have it any more.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#756Earlier quoted context omitted.
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation. OpenAI et al also stole everything from ever…
Personal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.
Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#757Earlier quoted context omitted.
Well, Buckmaster says both his and Alpöge's use of Codex was non-institutional, and OpenAI claims the right to train their models on inputs and outputs of non-enterprise users in their service policies [0]. So I'm not sure they were even promised that. [0] https://openai.com/policies/how-your-data-is-used-to-improve...
It isn't relevant whether they were promised that. Indeed I think the assumption must be that they were not promised that, since otherwise the author asking if they were would not make much sense. If OpenAI did use the conversations from Buckmaster and Alpoge, then not disclosing it, explicitly, is plagiarism. If they planned to use that plagiarism to pressure the authors to publish, that is even more unethical. What…
(And if memory serves, there is also the opt out from training on consumer subscriptions). Its not plagiarism if you make your data available for the purpose of training their LLMs. It is you giving away your IP for some tokens.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#758Earlier quoted context omitted.
It sounds like OpenAI is trying to appease the author when they don’t have to by allowing him to rewrite their proof. They probably don’t believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
Given that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
They don't need any help publishing.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#759Re: Navier-Stokes – Tristan Buckmaster [pdf]
#760Only here to say, regardless of the drama, shouldn't we all be excited if the Navier-Stokes gap is closed? Time will almost certainly reveal a lot more about the drama and the related ethics, but let's get excited about the actual breakthrough as well!