Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

631–640 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#631

Earlier quoted context omitted.

Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritu…

> I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritual truth of this, so please give it zero weight)

Why would you include a statement that you want us to give zero weight to, unless you don’t actually want us to give it zero weight?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#632

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable.

I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#633

Most people here are missing the forest for the trees. We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights. We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish lin…

This reminded me of anecdotes of people discussing with friends about buying a random specific item, and then suddenly seeing it advertised everywhere before even googling about it.

Next step, discussing your Navier Stokes solutions with friends might require leaving your phone in another room.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#634
post #398

Earlier quoted context omitted.

I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What…

> They intentionally threw $15 million in compute at the problem what? really?

When I worked at Google, we spent $100M in power on protein folding and drug discovery (this was long before AlphaFold). Never underestimate the willingness of smart rich people to invest in speculative science.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#635
post #486

Earlier quoted context omitted.

No. 1. Buckmaster contacted OpenAI first. Not the other way. 2. Giving the $1M bounty to a human mathematician for the effort and giving him credit would be excellent PR. They had already burned much more than $1M for the generation. Adding him as author also costs nothing. Purely pragmatical. 3. “As long as he removed Alpöge” part itself is against academic honesty by all means. 4. Buckmaster rejected fame and $1M o…

First of all I put "generosity" in quotes because I don't believe a corporation as big as OpenAI is even capable of acting out of generosity. It's always one of the three: A) PR B) commoditizing complements C) stupidity. In this case it's more like C) though, as in hindsight the best move OpenAI could do is insisting that they just used an insurmountable number of tokens to exhaust all the published directions. They…

> the best move OpenAI could do is insisting that they just used an insurmountable number of tokens to exhaust all the published directions

So… lie more? They knew the approach and started there.

At least they were honest about that.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#636
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

The problem here is that OAI (and others) pretend or claim that this is uncharted legal territory, where in fact it is very simple. We have a machine that is fed data, and produces new data as a result. If that new data depends (in any way) on the fed data, then from a legal viewpoint it is derived from that data. Whether they anthropomorphize the operation performed by the machine does not matter. They can anthropom…

> in fact it is very simple

Even if this opinion were backed up by a court ruling, it would definitely not be “simple”. It will be a very ugly case if it is ever litigated. A lot of money will be spent and no guarantee at all the plaintiff wins.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#637
All players in this space are doing the same thing with all data, no surprise here. They are stealing IP across the board with support to allow it: https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...

IMO, it is extremely naive to trust these black box remote service API calls, especially at an institution that can provide $$$ for local compute.

This whole fiasco reminds one of this story: https://www.theregister.com/offbeat/2010/05/14/facebook-foun...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#638
post #473
post #469

Earlier quoted context omitted.

Tools don’t turn around and scoop you. What OpenAI did here was use the same tool that the researcher did which might have coupled their work together.

You’re right, they don’t. It was scooped by the humans at OpenAI who published the paper. The tool they used to do it isn’t that relevant.

> isn’t that relevant

“Might not be” that relevant. You’re dismissing the whole controversy without addressing why it’s controversial.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#639

Earlier quoted context omitted.

Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritu…

> we have zero reason to believe our models did anything fishy.

Obviously. They cannot do anything "fishy". They are just computer programs.

Now, how about their operators?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#640

Earlier quoted context omitted.

Nonsense; the claim was that they wouldn't do anything that would mean they'd be > risking massive lawsuits and a total loss of trust Evidence of the massive lawsuits and lack of trust seems pretty relevant.

By "loss of trust", I meant that this is something they would risk losing a lot of users over, which isn't the case with the other lawsuits. There is a massive distinction between fighting third parties in a legal grey area and committing blatant fraud against your own users. Even if you have no regards for ethics, intentionally shipping a noop "do not train" toggle offers negligible upside for a massive downside.

They're being sued for several issues that resulted in the deaths of users; that's not fraud, but it is against their users.

The parallel still holds, and the information is still on the page I linked.

Post reply on HN