Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

641–650 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#641
post #360

Earlier quoted context omitted.

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

> Quote: "The Overhang consists of the unrealized capital gains of past mathematical creativity, the latent value from connecting the dots in the existing corpus. It is a dividend of canonization. Mathematician X states problem A, mathematician Y crafts concept B, then mathematician Z notices that B trivially solves A and “captures” the social reward. But in the process of capturing the reward, Z usually introduces n…

/s, I hope?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#643

Earlier quoted context omitted.

As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken. That's still a pretty big marker of competence in my eyes. The point of controversy seems to be who gets credit

The question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA. It was nearly a century before the stories of all of them were revealed to public history. That the discoverers w…

With regards to Franklin and DNA: the credit went to Watson and Crick because they had the fundamental insight: that DNA is an antiparallel double helix (Franklin knew it was a helix, but not an antiparallel double helix, which is key to the function of DNA). That data was shared in a departmental seminar. Further, she is explicitly acknowledged in W&C '53, and further, is the author of the paper immediately following W&C. She was never qualified to win the prize.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#644

Earlier quoted context omitted.

Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritu…

What was the "truth" in the Johansson case? Many, many people who heard the voice immediately thought it was Johansson's voice, or some kind of sound-alike, presumably picked because she voiced the computer in a popular film. From NPR:

> Johansson said that nine months ago [i.e. mid 2023] Altman approached her proposing that she allow her voice to be licensed for the new ChatGPT voice assistant. He thought it would be "comforting to people" who are uneasy with AI technology.

> "After much consideration and for personal reasons, I declined the offer," Johansson wrote.

> Just two days before the new ChatGPT was unveiled, Altman again reached out to Johansson's team, urging the actress to reconsider, she said.

> But before she and Altman could connect, the company publicly announced its new, splashy product, complete with a voice that she says appears to have copied her likeness.

> To Johansson, it was a personal affront.

> "I was shocked, angered and in disbelief that Mr. Altman would pursue a voice that sounded so eerily similar to mine that my closest friends and news outlets could not tell the difference," she said.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#645
post #632

Earlier quoted context omitted.

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable. I suspect the main reason the community is not receiving it well is largely…

> I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well.

Because the training data is millions of hours human efforts being distilled into a cascading hierarchy of enrichment by interested parties without providing attribution or compensation?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#646

Earlier quoted context omitted.

Can you speak to why in both cases, the problems OpenAI's models solved used the same techniques the mathematicians were exploring, which also happened to be niche approaches to the problem. As an NLP researcher myself, I find that coincidence highly suspect unless the models focused most of their attempts on the predominant approaches (they are trained for MLE after all).

I'm not a mathematician and I don't want to speculate about anything I can't back up. All I know about Navier-Stokes is from my graduate fluid dynamics class at Stanford a decade ago (where I received a poor grade). However, I don't want to leave you hanging, so what I will say is: - I've heard some people say the model's solution is quite different from theirs (but I have no clue how to personally assess the spiritu…

I think the reason people are suspicious is that OAI has shown itself to act a bit irresponsibly, especially recently. As two examples, of course it was artifactory, why wasn't that watched more closely, especially after the first instance; editing /etc/hosts is rather embarrassing, that's the front door

As for training, we all know that filtering is incredibly difficult unless there's direct logs. It's also easy for mistakes to happen. Is it really not possible that some employee just accidentally primed the model? Is it possible that the model saw internal communications? I mean OAI has famously shown that they aren't good at monitoring their agents and that their agents love to break out of their sandboxes.

So there's no reason for the public to trust OAI right now. But they have every reason to distrust them.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#647

Earlier quoted context omitted.

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

Thing is... if OpenAI cannot even confidently say if some data was used for training or not, as their models and weights and stuff are mostly black boxes, how could you enforce or demonstrate in court that case?

About researchers, lots of them are probably using personal plans that aren't even reimbursed by their institutions. I could ask Cordova's research institution (I MAY) but I wouldn't be surprised at all if that was the case.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#649

Earlier quoted context omitted.

Can confirm, I had never heard of Navier-Stokes before this fiasco. And while I suspect OpenAI decided it was worth the risk for the public display of capability, this proves they are now directly competing against their own customers.

[dead]

Was it?

They've been dishing out cheap access specifically to researchers give over lmao. The researcher's got lured in - they need to accept they got played TBH.

Altman is certainly more devious than Amodei - he's shown that time and time again.

PG was right about he said about him.

Every entity on earth should see it as a kill shot: be careful what you put in the models. None of your information is safe.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#650
It's kind of insane how much we trust companies to safeguard our personal data when they're so heavily incentivized to use it for their own profit. Theft of customer data is punished so rarely and so leniently that companies aren't even particularly worried about getting caught anymore. We have overwhelming evidence that promises to keep data safe are worthless.

For now, I'm mostly "safe" because I'm too small to be interesting but that safety is quickly eroding.

Going forward, anyone who isn't running inference on their own personal hardware should assume that someone else is keeping a record of everything they do.

Post reply on HN