Earlier quoted context omitted.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…
It feels convenient to not spend time on engineering around tooling that could be used to answer a question like “did you violate copyright by training on X?”
Navier-Stokes – Tristan Buckmaster [pdf]
761–770 of 871 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#762Does this have any relation to the singularities in black holes?
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#763Earlier quoted context omitted.
What a horrible response from Bubeck. How did he think that would make him and OpenAI look good to tweet that?
I especially loved the part where he claims they spent $60m on compute to push on N-S because of a Twitter rumor, and they totally didn't steal the idea from mathematicians using their tools.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#764This is why I left math even after solving a 20 year old conjecture in grad school. Literally who cares who solved the problem just publish the results. Academia was always politics first results second and I AM GLAD that LLMs are becoming superhuman at math. I like better theorems, not better politics.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#765Earlier quoted context omitted.
Frankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement). And the difficulty is harder than just the extreme scale of text searching. but also explodes…
To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal.
Not tracking data changesets like code changesets is certainly a choice, not really a constraint anymore. A similar choice I feel is implied by "extreme scale of text searching".
> no they will not all add the telemetry you wish they did
...is just a failure of the corporate policy surrounding data handling. Is git-for-data already considered telemetry?
Of course the truth is provenance of data is something best institutionally forgotten as quickly as possible. The only thing that matters is it's there, that the data has no history, and that's why it can be used in whatever way deemed necessary.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#766From OpenAI: > While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool. The fact that this is…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#767Earlier quoted context omitted.
Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof. https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
I work in catastrophe risk modeling and it's a multi billion dollar industry. We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do. There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and m…
Why should your business be any different?
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#768So the entire "new internal model" is 'parallel construction' in criminology speak to justify spying on Buckmaster's work in a legal way.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#769Is it possible that they claim they used a "new internal model" to provide an excuse for "training" on new user data, knowing full well this new user data will include Buckmaster's chat? So the entire "new internal model" is 'parallel construction' in criminology speak to justify spying on Buckmaster's work in a legal way.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#770Earlier quoted context omitted.
If this is a "wake up call" - then your legal team needs immediate education. First - there is this - https://openai.com/policies/how-your-data-is-used-to-improve... (linked from the Navier Stokes writeup) I don't know how much more clearly they can write: > When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models. One of the key selling tactics that co…
ChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated : > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more The "Learn more" link takes you to the link you've shared.