Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

771–780 of 868 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#771

Earlier quoted context omitted.

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

I work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and m…

> We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.

I feel this is exactly what will happen as they cause all sites to go closed source to protect their intellectual property and the AI companies offer only biased information. They are replace the business on internet model by bankrupting everyone with their own tools. This is predatory pricing under most antitrust laws (imho, not a lawyer) and it is very easy to do when you dont need to pay for the raw material.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#772

Earlier quoted context omitted.

> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. You…

> The chances of getting caught are 0 I'd say non-zero, as seen in the current state of affairs.

sort of agree, but also sort of think an accusation with lots of people arguing is not exactly the same as being caught.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#773

Earlier quoted context omitted.

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

I work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and m…

> If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.

I mean how could you expect them to not given they've trained the existing models on effectively the sum total of all human knowledge available on the internet without regard to copyright/ownership of that material.

It's a little trite but this absolutely runs into the "Frog and the Scorpion", it is simply in their nature.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#774

Earlier quoted context omitted.

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them. But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.

And require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#775

Earlier quoted context omitted.

Did we read the same tweet? It felt very forthright and level-headed to me. Not at all what I expected.

> Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic So using someone’s models makes someone who works for the competitor not “independent”? When their coauthor is? What does that even mean? I almost stopped reading this extra long post entirely at that point. This is not a good look in my book.

Skipped an important word, didn't you?

> internal Anthropic models had been used in their proof ...

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#776

From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…

The wording seems to confirm it was part of training data at some stage. If not, they would be able to prove it quite easily I assume.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#777

Earlier quoted context omitted.

Many do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here: Science papers of a phd level must contain: 1. one or more novel insights 2. a long list of citations to contextualize them and 3. some work to prove that the insights are in fact meaningful --- In this context, consider a prompt based diffusion model which, when aske…

I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context. As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus. So is…

I am not engaging further. You asked for meaningful refutation of your observation, I provided one.

Academia has stricter rules than regular society.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#778

Earlier quoted context omitted.

> If OpenAI is indeed using customer data to train their models to win a $1m prize Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. You…

Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reas…

> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.

The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#779

Earlier quoted context omitted.

This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.

Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves. You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training). Then…

[deleted]

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#780

People here do not seem to be considering the second-order effects of these series of events. No academic institution or enterprise will trust OpenAI, Anthropic or any other non-local AI model with their core IP. There will be severe restrictions on what employees at these companies/institutions can share with AI services even from their personal accounts. (Or I am just overthinking it)

If so that's an extremely good thing
Post reply on HN