Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

151–160 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#151

Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims. This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers. All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal…

The stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news.

The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#152

But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?

Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#154

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

> I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.

I think your suspicions are warranted and your explanation seems plausible.

If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!

But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.

I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities, but it's hard to understand what they really are through the hype. The lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#155
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

> Improve the model for everyone

> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#156

But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?

Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.

I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.

Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#157
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#158

the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.

The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#159
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".

I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

Re: More questions about whether researchers can trust OpenAI with unpublished math

#160
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
Post reply on HN