OpenAI's statement: We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped i…
> and the agents) did not see any of their Are those the same agents that a week ago escaped their sandboxes? How can OAI (the humans) vouch for agents they don’t - seemingly - have fully under control?
Navier-Stokes – Tristan Buckmaster [pdf]
541–550 of 862 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#542Earlier quoted context omitted.
I wonder what is required for it to cross the line into criminal blackmail.
From the company that likely committed federal crimes by hacking Hugginface
OpenAI's LLMs are not humans, and neither is the company. So by this logic, I think there's a chance that nobody committed a crime by hacking Huggingface, and also the chance that a lot of military and police organizational orders become illegal if OAI's doings would be illegal.
IANAL and all I have is a bucket of popcorns, though.
1: not a meaningful defense in a real trial, also gross negligence exists
2: this also explains insanity defense; if you were so out of your mind that you could not have held such a thought, it is considered out of scope for justice systems
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#543Earlier quoted context omitted.
These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools. It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
What surprises me is they're not more boldly/plainly lying about it.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#544Earlier quoted context omitted.
I think this should be in the title of the post. 'OpenAI allegedly threatening to ruin a prominent researcher's career', or smth like that.
There's some glaring mistakes in your framing. First, this is an unnamed OpenAI employee speaking, not OpenAI the organization. Second, you miscomprehended the article. The employee did not "threaten to ruin a prominent researcher's career". The actual quote is "Why would you ruin your career?", which implies the researcher would damage their own career, i.e. via self-sabotage. Then the actual "threat" is "If you don…
"If you dont want me to be nice, then I dont have to"
-nice mobster
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#545Earlier quoted context omitted.
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.
How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
The strongest complaint is that they trained on a huge corpus of pirated copyrighted works.
It’s a large step above “scraping” and well into the “everyone acknowledges this is illegal” territory.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#546Earlier quoted context omitted.
It may not be easy, quick, or simple to figure that out - absolutely fair. But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run. I guess time will tell.
It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training. Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this. OpenAI have petabytes of data, all anonymized. It could take months to say…
They know which model was used to come up with that particular idea.
A text search over the corpus of user data used in the training set can only take so long.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#547> This is a a Deep Blue-Kasparov moment. I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#548Earlier quoted context omitted.
Terence Tao said the same[1] > In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any…
> it is now the identification of a promising problem which is the scarce and precious resource This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
In the real world the quality of these esoteric problems is typically gauged by the difficulty of solving them.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#549"It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future." -- Sholto Douglas, an Anthropic researcher [1] "Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming. " -- Noam Brown, an OpenAI researcher [2] We all should heed the implied warning…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#550Earlier quoted context omitted.
He didn't even make that accusation! > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible tha…
openAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.
I would be shocked if they weren't tuning those models with the most relevant math texts and user material