Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

451–460 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#451

Earlier quoted context omitted.

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained

That doesn't make it fine. We should not excuse this behaviour just because its rampant already, especially when it comes to such a serious prize

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#452

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

Opt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...

Edit: the parent comment now seems to better reflect the below.

That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on.

So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#453

Earlier quoted context omitted.

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable I have terrible news about how literally every leading AI model was trained

I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point.

Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#454
post #138

Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems. "Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."

It's part of their TOS that they can train on users' private chats.

Not if you pay to turn that off. We don't know if Tristan did.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#455
Crazy how optons are OpenAI has superhuman model in frontier mathematics, and OpenAI stealing research data. Like no one will really care what EULA checkbox Tristan ticked, and I think most people will eagerly believe OpenAI is shady org with little scruples, and that big tech data is not actually so siloed that marketer can say we can do XYZ with private data to help with valuations (especially considering timeline). Employees have been creeping on their exes for much less.

IMO the parsimonious answer seems to be OpenAI has a pretty good model (because it did finish) and stole someones work... and threatened them over it. TBH all OpenAI need to do is solve another millennial problem and none of it would matter - people expect them to behave heinously regardless - but if they have generalized superhuman math model... well I guess they're allowed io.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#456

Earlier quoted context omitted.

if they do not deny training on them, they can't deny plagiarism.

By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.

as good academic conduct you may cite the source of the work you are quoting or paraphrasing.

as bad academic conduct you may steal someone else's unpublished work, work on it yourself for a bit, and then publish it as your own work. and then threaten the original author!

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#457
post #85

Earlier quoted context omitted.

Or, given how dedicated he appears to be to the company, a promotion and a raise.

Dark take, but I really hope not. If anything it would be a good opportunity to buy some goodwill by washing themselves from all the alleged shadiness so far.

when it comes to openai and shadiness I'm pretty sure it's a "fish rots from the head" situation.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#458
post #448
post #441

Earlier quoted context omitted.

Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.

Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all it’s kind of their core business model

Indefensible, not impossible. As you say it is quite technically feasible.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#459

Earlier quoted context omitted.

It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course. Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were g…

They are not training a whole model in a matter of days

Models are very obviously continuously updated.

Model editing to remove PII that slipped through, all sorts of things of that sort.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#460

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first.

The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.

Post reply on HN