It's also difficult to imagine any competent AI researcher would overlook the possibility of training data leaking into the test, especially given that they know the users have been using their model to work on the same problem, and that they jumped on the problem after hearing rumors of the breakthrough.
Navier-Stokes – Tristan Buckmaster [pdf]
841–850 of 873 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#842Earlier quoted context omitted.
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case. The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor: 1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details Ergo, the agent could likely decide it would l…
> The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor It requires humans to verify what agents have done. Weird
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#843Earlier quoted context omitted.
> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work > If true, that's generous and beyond the level of generosity one should expect "We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so gen…
The two proofs are structurally very different, and don’t even prove the same conjecture. It’s becoming very clear that OpenAI did not steal anything here.
Ain’t no way you read one 100 page paper let alone two with enough understanding to make such a claim.
You ought to be ashamed of yourself.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#844Earlier quoted context omitted.
I'm stunned that people are taking this accusation as a fact. OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business. There are things that Buckmaster alleged and things that he specu…
> The entire training data thing is speculation. I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that. I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#845Re: Navier-Stokes – Tristan Buckmaster [pdf]
#846Re: Navier-Stokes – Tristan Buckmaster [pdf]
#847Re: Navier-Stokes – Tristan Buckmaster [pdf]
#848One of the best takedowns of samas/OAIs spin on this I've seen is this post on X: https://xcancel.com/BetterCallMedhi/status/20974723362578637...
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#849Earlier quoted context omitted.
Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS: Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing. Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format. There's a reas…
> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.