Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

411–420 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#411

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

If they have solved hundreds of open problems in math, why are they publishing results for the ones other mathematicians happen to be working on at the same time? Why not the others?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#412
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.

Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.

Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!

Re: More questions about whether researchers can trust OpenAI with unpublished math

#413

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

The Cult tells us the AI is almight andpowerful; unfortunately, the cult cant actually describe the indescribable.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#414

Earlier quoted context omitted.

OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, th…

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

Apart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs.

You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property.

You can learn a lot from a single side of a conversation.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#416

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

> Anyone remotely familiar with CC over the years

years?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#417
Astra is strange. I asked it to design a treehouse and it just stopped every couple of minutes telling me what it still had left to do. After dozens of continue prompts it finally gave me a structure that would work but it was 10x more wood than I needed. I think the key mistake I made was asking it to “approve” the design for building. As soon as I asked that of it, it started getting “scared” and “apprehensive” and wouldn’t complete what I asked of it.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#418
post #394
post #360

Earlier quoted context omitted.

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

Could "superintelligence" arrive as basically applying this overhang to all other domains?

It already did.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#419
post #293

Earlier quoted context omitted.

If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.

> were the first victims Spinning it negatively like that doesn't do anybody good. Were mathematicians the "victims" of calculators? of Matlab? Were writers the ""vIcTiMs"" of word processors?? (apparently yes, according to old TV shows about computers during the 1980s, that you can see on YouTube) > "tHiS iS nOt ThE sAmE" — Everyone every time. No, just look it up. Look into old magazines and TV shows or newspaper a…

It’s not the same. AI potentially completely replaces intellectual work without creating any* new jobs (*almost any - there will be some extra jobs for building data centers but that’s negligible).

Re: More questions about whether researchers can trust OpenAI with unpublished math

#420
Forget researchers you as a business are putting in your business optimizations, your processes in order to train it so that Ai can then give that information to your competitors once incorporated into its training set. You are literally training your competitors.
Post reply on HN