Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

461–470 of 872 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#461
post #388

Earlier quoted context omitted.

Anthropic too, and evidently others. Amazon were recently confirmed to be doing the same thing: https://www.404media.co/we-tracked-a-shipment-of-rare-books-...

And that’s bad, right?

It's legal. I wouldn't do that myself, but I guess that's why I don't train models for a frontier AI lab.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#463

Earlier quoted context omitted.

If he didn't opt out I'm not sure I'd agree that it was fair game. I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.

It’s also very shortsighted to stiff a customer like that. Why would I trust them with my data and ideas?

I think that’s what we’re all talking about here, you shouldn’t.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#464

Earlier quoted context omitted.

Opt out doesn’t guarantee they can’t train on “your” data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...

Edit: the parent comment now seems to better reflect the below. That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on. So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning…

We come back to the rule: "The cloud is just someone else's computer".

The way for people or companies or universities to control their data and information is to keep it on their own computers.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#466

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

Even if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#467

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler…

Terence Tao said the same[1]

> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.

[1] https://mathstodon.xyz/@tao/117237322160500501

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#468
post #349

OpenAI's statement: We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped i…

It’s plainly false that they cannot rule out whether their de-identified data was used in training their model. Just that they haven’t ruled it out.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#469

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

Eh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints.

OpenAI should release the agent log, including CoT.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#470

Earlier quoted context omitted.

By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.

Isn’t that one of the most salient and straightforward argument against LLMs?

Just overfit ad infinitum:)
Post reply on HN