Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

851–860 of 873 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#851
post #724

Earlier quoted context omitted.

> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work > If true, that's generous and beyond the level of generosity one should expect "We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so gen…

The two proofs are structurally very different, and don’t even prove the same conjecture. It’s becoming very clear that OpenAI did not steal anything here.

User name checks out

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#852

Earlier quoted context omitted.

What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends. In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. Thi…

> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I o…

I simply do not see how you can interpret "if they do not deny training on them, they can't deny plagiarism" as "all outputs are necessarily plagiarizations of the training data".

There is a difference between claiming an action is a transgression, and claiming it could be one.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#853

Earlier quoted context omitted.

The chats the professor had are not generic knowledge. And yes of course you still need to cite Newton and Leibniz depending on what result you want to mention. What’s allowed to be not cited are not status quo, the term you’re looking for is “folklore” results aka results that have been around so long that 1) nobody knows who came up with them or 2) everyone knows who came up with them. The second point: if you say…

Even if AI used the result, AI pushed it to the finish line while Levent and Tristan did not. But I understand the approach was different, the information leak was only that it was "doable".

The approach was the same.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#855

Earlier quoted context omitted.

One thing that wasn't obvious to me or adults around me when I was younger: most laws define whatever acts a law punishes as individuals commiting to it, not as situations manifesting anyhow. It's not a murder just because someome died hit by a bullet you fired, but you have to have personally decided to kill that person leading to their death[1][2]. OpenAI's LLMs are not humans, and neither is the company. So by thi…

> OpenAI's LLMs are not humans Neither are guns. Which is why we punish the person shooting the gun and not the gun. Industrial equipment, which is how i would classify LLMs, hurting people is nothing new. The relevant questions are: - did someone intend it to happen? - was someone negligent in taking reasonable steps to prevent something foreseeable? The justice system doesn't punish people for legitimate accidents.…

I mean, they are going to be at least culpable for gross negligence, but who made the decision might become an interesting point of discussions(mainly in deciding precisely how they should be punched in their face and less in whether they should be at all).

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#856

Earlier quoted context omitted.

I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…

It definitely is solvable though. Data versioning is a thing and it can work quite transparently to the mutations done on the data. To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal. Not tracking data ch…

All of your points are valid, and believe me I was trying to make them. The problem is one of culture. Most of the people doing this kind of work didn't like version control, and their work was really just running notebooks (like iPython or Google Colab) until a number was good enough and they'd submit the file for inclusion into training runs.

You can call it a policy failure, but these people were in very high talent demand and so top down dictates would risk "X people leaving lab Y for lab Z" headlines and morale hits.

I am not saying this is good. I am telling you that on the ground it is so much messier than it should be.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#857

Earlier quoted context omitted.

I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…

It feels convenient to not spend time on engineering around tooling that could be used to answer a question like “did you violate copyright by training on X?”

Don't attribute to malice what is better explained by coordination headwinds in extremely large companies.

The engineering around tooling wasn't remotely the issue. It's getting all the (thousands?) data researchers mostly iterating on fine tuning datasets that would get bristly if they couldn't work outside version control in a python notebook iteratively tweaking their dataset that processed and reprocessed a few datasets until a threshold was reached.

The only _guaranteed_ chains of custody are down at the compute job and file read level. Which in a massively distributed computing job is... [redacted] nodes reading [redacted] fanouts of "datasets" that is just an abstraction over [redacted] individual files.

There's no malice here. Just way way way more complex than you'd first think.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#858

Earlier quoted context omitted.

I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…

This sounds like data laundering in effect.

The appearance of heroic efforts to get authoritative lists of what datasets went into which major model versions prevented actual data laundering (up to intent and mistakes). But don't attribute malice to that which is far far easier to explain with coordination headwinds: https://komoroske.com/slime-mold/

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#859

Earlier quoted context omitted.

> OAI tried to share, but didn't want to share with an Ant employee Why does OpenAI get to dictate who Buckmaster can claim co-authorship with? > I find their reaction childish OpenAI may have, with full plausible deniability, taken Buckmaster’s work and passed it off—in substantial part—as their own. (Fitting into a fact pattern of them having tried to do the same with Apple.) There is a material takeaway for anyone…

OAI didn't try to claim Buckmaster's work as their own. OAI tried to let Buckmaster present OAI's work. it is nearly the opposite of your accusation

Pretty strange if they didnt steal Buckmaster's work, right
Post reply on HN