Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

41–50 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#41

Earlier quoted context omitted.

The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.

So are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.

Do you not have reading comprehension? Did you even read the statement? He himself asserts at the end he doesn't know if the above example is true or not. What on earth are you going on about? Where in my comment am I assuming he's lying ?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#42

This thread is full of jumping to conclusions based on a biased perspective. Have some humility.

It's also full of your posts baselessly defending OpenAI. Maybe you should also heed your own advice?

I've been right historically, check my track record. How about you?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#43

I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…

This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data.

And simply knowing a problem can be solved is half the battle.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#44

Earlier quoted context omitted.

The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.

So are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.

You really lack reading comprehension

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#45

I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…

I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…

He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#46

I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…

I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…

He didn't even make that accusation!

> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.

Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#47

I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…

This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data. And simply knowing a problem can be solved is half the battle.

From Buckmaster's text:

    The route to the Clay problem through a
    smooth force, options c and d in Fefferman’s statement of the problem, is the
    route Luis and Diego opened and the one Levent and I had quietly chosen to
    attack. Almost nobody else I know of was working on it. It is not the direction
    one arrives at in a few days by giving a model the problem statement. When I
    heard “forced,” it was a bright red flag.
This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#48
post #46

Earlier quoted context omitted.

I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…

He didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible tha…

Couldn't the Enterprise have a different fine print?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#49
When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon and not necessarily reproduced verbatim, you can assume that unless your chats depict a foundationally new and effective style of communication or ideation, there would be little need or use thereof of training on your chats.

For eg: "Hey ChatGPT my name is X and I am 6 and a half feet tall. Am I anaemic?" This is a query, and while it might suggest to an AI model that tall people may worry about iron deficiencies, it's not really necessary to include in training. The user may be tall or short, but the idea that one may randomly ask about anaemia is not exclusive to this dataset. At best, this chat is an example of linguistics, not anything else, and the models figured out how to write and answer such questions years ago. It is ignored in training.

But when your work involves solid complex and unique mathematical proofs, the data is suddenly worth training upon. If I understand it correctly, the LLM may view your approach as a brand new path to take to solve an otherwise intractable problem. Its reinforcement training emphasises that it should do this in order to improve. And since it leads to results - large internal teams likely flag the model that reached this stage, the model is rewarded and given compute and attention - it is a desireable outcome both for the model and for OpenAI.

OFC, OpenAI becoming an advertising company will suddenly have incentive to treat all data as valuable. But while they are a "we need to make headlines" company, it's more rational that they view these examples of data as more valuable than others.

I don't doubt that they trained on his chats. This seems like the ideal usecase for "mass surveillance but using training" as a sort of filter.

But even so, one wonders how the model differentiates. If the researcher entered proofs into ChatGPT every day that mentioned "strawberries", while no other math paper on the topic did so, does that mean their chats would be audited?

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#50

> This is a a Deep Blue-Kasparov moment. I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)

Amazing how history rhymes
Post reply on HN