Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

741–750 of 861 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#741
post #83

Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…

> - Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models.

One minor wrinkle in Alpöge's narrative, he refuses to deny that he hasn't used any non-public Anthropic models for his "independent" research.

https://x.com/giffmana/status/2097560503581069585

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#742

Earlier quoted context omitted.

> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. You…

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#743

Earlier quoted context omitted.

> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. You…

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them.

But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#744

Earlier quoted context omitted.

Frankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.

I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…

It feels convenient to not spend time on engineering around tooling that could be used to answer a question like “did you violate copyright by training on X?”

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#745
There is clearly issues about IP, privacy and accreditation and the motives of powerful companies, but part of me can't help but be excited that whatever the method, the result is the genuine progression of human knowledge - we all win. It's not unusual for mathematical problems to last centuries and we might have a technology that can solve these problem, all these problems (??) in our lifetime. Then there's the repercussions on science and technology... what an astonishing time to be alive.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#747

Earlier quoted context omitted.

A needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?

What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends. In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. Thi…

> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized.

We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I observe that it is absurd to object to a single action being a transgression on the basis of an argument which implies that all actions are inherently transgressions. Notice that nowhere do I take a position on whether or not the argument about all actions being transgressions is true or false.

> you appear to believe that ...

I do not, no. I have not taken a position of my own here. I've merely objected that the one I responded to does not make for a sensible line of argument in context. It seems that you (and many others) have read my objection to position A as support for position B and attempted to infer what I think from that.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#748
post #17

> It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. T…

> "If you don’t want me to be nice, then I don’t have to be nice." Exhibit nr 8453324 to not trust Sam Altman and OpenAI.

He most definitely had a hand in murder of the 26 year old whistleblower: https://www.youtube.com/shorts/lL7WA6i3zVw

Mother's interview: https://www.youtube.com/watch?v=Kev_-HyuI9Y

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#749
post #650

Earlier quoted context omitted.

I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.

Especially because the data that gets fed into training is first anonymized, so they’d need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristan’s permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not…

If the method is indeed found in the training set it's not particularly important to de-anonymize it. You have the proof you need.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#750

Earlier quoted context omitted.

> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. You…

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.

You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).

Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).

In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.

So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.

The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.

And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.

So I'd say a whistleblower is pretty out of luck even becoming one.

Post reply on HN