I don't know how the math community handles this but normally I would think if X mathematician comes up with an idea and Y mathematician uses it to solve some problem, Y would get credit. But does that change if Y heavily relied on LLMs? I suppose we're going to find out.
Navier-Stokes – Tristan Buckmaster [pdf]
571–580 of 862 posts
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#572Earlier quoted context omitted.
I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler…
Terence Tao said the same[1] > In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any…
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#573Earlier quoted context omitted.
> Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic So using someone’s models makes someone who works for the competitor not “independent”? When their coauthor is? What does that even mean? I almost stopped reading this extra long post entirely at that point. This is not a good look in my book.
The first part of the sentence is important: > Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. If an Anthropic employee is doing independent research, but with models that aren't available to the public (because they're internal models), then . . . idk. It's not clear to me why that sho…
If true, that's generous and beyond the level of generosity one should expect. Extending that courtesy (beyond academic norms) to a competitor is expecting too much. It take a result OpenAI spent millions of dollars on, and put "Anthropic Researcher" right on the cover.
This is, of course, taking OpenAI's side of the story at face value. But it is a consistent, coherent, and ethically justifiable series of events, if indeed it happened that way.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#574Earlier quoted context omitted.
Frankly, I don't buy this difficulty argument. They know which model was used to come up with that particular idea. A text search over the corpus of user data used in the training set can only take so long.
I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced so…
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#575Re: Navier-Stokes – Tristan Buckmaster [pdf]
#576Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…
> OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but only if they remove Levent as an author, as he works for Anthropic well that sounds like an asshole move.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#577Drama/accusation summary: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler." - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help le…
please could you change 'drama' and 'accusation' to something more formal like 'allegation'. the paper makes a very serious allegation of dishonesty and possible academic misconduct. the governance and integrity of openai is of importance to the welfare of society. this is not a matter of drama.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#578I'm not one to comment often but this really pisses me off. OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!). Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to…
The whole business is based on reselling user data scraped from the whole internet
It’s plagiarism at scale
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#579When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon and not necessarily reproduced verbatim, you can assume that unless your chats depict a foundationally new and effective style of communication or ideation, there would be little need or use thereof of training on your chats. For eg: "Hey ChatGPT my name is X an…
Assuming the results included some external validation such as user's preference, compilation, lean, etc., I'm not sure whether this would lead to model collapse.
Re: Navier-Stokes – Tristan Buckmaster [pdf]
#580Earlier quoted context omitted.
How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
> who trained models on scraped internet data The strongest complaint is that they trained on a huge corpus of pirated copyrighted works. It’s a large step above “scraping” and well into the “everyone acknowledges this is illegal” territory.