Earlier quoted context omitted.
Yeah well, its easy to do if you steal someone elses work and then try to threaten them into staying quiet about it Edit: OpenAI have now admitted they were training on prompts at the time they made their breakthrough: https://mastodon.social/@tristanbuckmaster/11723647135247030...
All the ai labs are open about training on prompts. The question is if buckmaster had disabled that with the toggle openAI provides.
On the Navier–Stokes Millennium Prize Problem
661–670 of 1001 posts
Re: On the Navier–Stokes Millennium Prize Problem
#662Earlier quoted context omitted.
Pre-IPO marketing?
I'm so tired of this "It's just marketing!!" commentary. An AI model just proved one of the top 3 unsolved problems in mathematics, they have a Lean certificate showing it's valid. How much more evidence do you need that these models are actually highly capable?
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
Re: On the Navier–Stokes Millennium Prize Problem
#663OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet another advantage of using open models right here.
fwiw there is a big "TRAIN ON MY DATA" toggle you can turn off (that they almost certainly did) and Anthropic MTS are posting that they almost certainly did not "steal" their methods
Re: On the Navier–Stokes Millennium Prize Problem
#664Earlier quoted context omitted.
I don't know what anyone's been saying on Twitter and I don't care. If it's really true that there's a model out there that's that capable two weeks after the start of training, then that's objectively a much bigger deal than a priority dispute, even if the latter involves juicy allegations of espionage and skulduggery.
It isn't a priority dispute, the more concerning allegation is that OpenAI may be training their models on prompts that mathematicians were using to solve this problem, and then surprise surprise OpenAI were able to replicate that work in their latest model What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question If OpenAI is training mode…
Re: On the Navier–Stokes Millennium Prize Problem
#665Earlier quoted context omitted.
All of those statements sound true, based on what I've heard. - "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input - it's all true that a team worked on this, a bunch of compute was burned, and the problem was…
>> I'm not sure how any of this provides evidence that OpenAI took any of their work. Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
Re: On the Navier–Stokes Millennium Prize Problem
#666Earlier quoted context omitted.
NS is a question for natural science. Q: can we model these bodies of discrete particles with a continuous approximation? A: if you do, you can get aphysical singularities. "If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."
This is a wrong interpretation. Physicists have a shit-ton of models that produce "aphysical singularities", they just work around those to get meaningful answers anyway. This is a whole trope and stereotype. Some of the most successfull and accurate predictions in all of physics come out after you discard a bunch of singularities. See e.g. https://en.wikipedia.org/wiki/Renormalization Nobody who actually works in fl…
Re: On the Navier–Stokes Millennium Prize Problem
#667Earlier quoted context omitted.
This is going to be dramatic in so many different ways. - First off, to reiterate, WOW. - Second of all, when does this end? Are we at the dawn of the singularity now? - People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them? - Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can butt…
2) We are witnessing the intelligence explosion from the first row, wherever this takes us 3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms. But apart from A…
Minor productivity boost in mathematics as people are no longer nerdsniped by the problem
Re: On the Navier–Stokes Millennium Prize Problem
#668"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Isn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
Re: On the Navier–Stokes Millennium Prize Problem
#669Earlier quoted context omitted.
That said, perhaps it will take over the world government and decree that all corrupt officials shall be imprisoned and all weapons of mass destruction shall be destroyed.
You think a model with an effective memory of 200-500k words, that can be unplugged, is going to "run the world" You people gotta put down the sci-fi
Re: On the Navier–Stokes Millennium Prize Problem
#670"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
Everyone knows that they train on the discounted rate plans data. All the labs are upfront about this too. If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
At this point, how can we even trust that they aren't accidentally training on those tokens too?