Live data from Hacker News

Navier-Stokes – Tristan Buckmaster [pdf]

cims.nyu.edu

541–550 of 862 posts

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#541
post #349

OpenAI's statement: We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped i…

> and the agents) did not see any of their Are those the same agents that a week ago escaped their sandboxes? How can OAI (the humans) vouch for agents they don’t - seemingly - have fully under control?

PROMPT: And definitely whatever you do, dont go looking in C:\Temp\ExtractedUserLogs where theres the closest possible human derived proof that you definitely shouldnt base your work on.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#542

Earlier quoted context omitted.

I wonder what is required for it to cross the line into criminal blackmail.

From the company that likely committed federal crimes by hacking Hugginface

One thing that wasn't obvious to me or adults around me when I was younger: most laws define whatever acts a law punishes as individuals commiting to it, not as situations manifesting anyhow. It's not a murder just because someome died hit by a bullet you fired, but you have to have personally decided to kill that person leading to their death[1][2].

OpenAI's LLMs are not humans, and neither is the company. So by this logic, I think there's a chance that nobody committed a crime by hacking Huggingface, and also the chance that a lot of military and police organizational orders become illegal if OAI's doings would be illegal.

IANAL and all I have is a bucket of popcorns, though.

1: not a meaningful defense in a real trial, also gross negligence exists

2: this also explains insanity defense; if you were so out of your mind that you could not have held such a thought, it is considered out of scope for justice systems

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#543

Earlier quoted context omitted.

These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.

I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.

It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?

What surprises me is they're not more boldly/plainly lying about it.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#544

Earlier quoted context omitted.

I think this should be in the title of the post. 'OpenAI allegedly threatening to ruin a prominent researcher's career', or smth like that.

There's some glaring mistakes in your framing. First, this is an unnamed OpenAI employee speaking, not OpenAI the organization. Second, you miscomprehended the article. The employee did not "threaten to ruin a prominent researcher's career". The actual quote is "Why would you ruin your career?", which implies the researcher would damage their own career, i.e. via self-sabotage. Then the actual "threat" is "If you don…

"Why would you burn your house down and kill your famiy?"

"If you dont want me to be nice, then I dont have to"

-nice mobster

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#545
post #317
post #315

Earlier quoted context omitted.

The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.

How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?

> who trained models on scraped internet data

The strongest complaint is that they trained on a huge corpus of pirated copyrighted works.

It’s a large step above “scraping” and well into the “everyone acknowledges this is illegal” territory.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#546
post #531

Earlier quoted context omitted.

It may not be easy, quick, or simple to figure that out - absolutely fair. But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run. I guess time will tell.

It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training. Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this. OpenAI have petabytes of data, all anonymized. It could take months to say…

Frankly, I don't buy this difficulty argument.

They know which model was used to come up with that particular idea.

A text search over the corpus of user data used in the training set can only take so long.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#547

> This is a a Deep Blue-Kasparov moment. I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)

If you read his account of things, its very much that if they didnt cheat, they gave themselves every opportunity to cheat. But above all that there was a bunch of chess protocol they failed to observe in that match, like providing seats for Kasparovs team and rooms for them to prep in. Even if they didnt have a big room full of chess notables definitely not refining the output, he was personally getting pushed around on a few fronts which unnerved him. If they had given him a few rematches I think they could have confirmed the win, but they refused which is super sus.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#548
post #518
post #467

Earlier quoted context omitted.

Terence Tao said the same[1] > In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any…

> it is now the identification of a promising problem which is the scarce and precious resource This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.

Sure; if you define the quality of a problem by reference to your ability to solve it.

In the real world the quality of these esoteric problems is typically gauged by the difficulty of solving them.

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#549

"It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future." -- Sholto Douglas, an Anthropic researcher [1] "Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming. " -- Noam Brown, an OpenAI researcher [2] We all should heed the implied warning…

It's nice to see this sentiment from Noam but now I'm confused at why he seemingly mocked Sholto's earlier message: https://news.ycombinator.com/item?id=49607090

Re: Navier-Stokes – Tristan Buckmaster [pdf]

#550
post #319
post #46

Earlier quoted context omitted.

He didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible tha…

openAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.

And the labs are all building panel of domin expert models, while simultaneously chasing open math problems.

I would be shocked if they weren't tuning those models with the most relevant math texts and user material

Post reply on HN