Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

581–590 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#581
post #536

Earlier quoted context omitted.

These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.

> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor: 1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details Ergo, the agent could likely decide it would l…

> The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor

It requires humans to verify what agents have done.

Weird

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#582

Earlier quoted context omitted.

I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced so…

I don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.' Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually sea…

> Can you explain the difficulty in engineering a search apparatus over a corpus of text data?

My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#583

Earlier quoted context omitted.

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.

> These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools.

Anyone who just reads headlines, if the lie gets around to more headlines than the truth does.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#584

Earlier quoted context omitted.

For context and balance, Bubeck has tweeted a curiously non-specific denial: > A series of false and inflammatory allegations against me are currently circulating on social channels. To clarify, I came into the discussion following academic norms, and I'm disappointed that it has come to this. Anyone who knows me knows that academic standards are of the highest importance to me. Will have more to say tomorrow.

This is an insane thing to read. Bubeck had a reputation even before he started with OpenAI. Of course it was him that was involved in this drama. This is such a sad mess, and it really didn't have to be this way.

What was his reputation before OpenAI?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#585

Earlier quoted context omitted.

This is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.

A needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?

What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends.

In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. This is not the case, no, because generative models do not "copy" or "create", they do both at different times.

I did not agree or disagree with the original poster, I was explaining to you why I thought you disagreed with them. If you understand what I said above, then why do you disagree with them?

EDIT: I just saw your other post on "general inspiration" and I believe I read the situation exactly; you appear to believe that inputs used to train generative models get "lost in the parameter soup", but it is not always the case.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#587

Earlier quoted context omitted.

my reading was that openai did not deny plagiarising the approach from their prompts. openai then tried to effectively bribe buckmaster with a shared citation, whilst dropping his co-author who works for anthropic. after buckmaster refused, openai tried to threaten him.

This definitely feels like the correct reading unless there is information we were not provided with. If the problem were unimportant, there would be no debate that this is not OK...

I was too slow to edit this post but I'm not sure why I said this. This is too assertive about a situation I don't know much about.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#588

Earlier quoted context omitted.

These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.

I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.

They are already claimed to be liars

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#589

Earlier quoted context omitted.

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good…

Sounds like the plot for Good Will Hunting 2.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#590

Earlier quoted context omitted.

It's specifically the last two bullet poitns - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but…

Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.

Here's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
Post reply on HN