Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

691–700 of 848 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#691

Earlier quoted context omitted.

Answering here because it does not let me reply to your second comment. but you said and I quote here verbatim: >"When you leave the relevant 'Help improve the model' toggle on, everyone working at the labs says it isn't used in the way you imagine" I have personally deactivated that toggle on multiple occasions and for some reason it keeps getting toggled on, as a matter of fact you can find multiple posts by people…

>I have personally deactivated that toggle on multiple occasions and for some reason it keeps getting toggled on Thank you for pointing that out. I just checked, and mine was on, too. Annoyingly, the toggle even stalls a bit, so I hit it twice when the first time didn't seem to work, and it quickly toggled off and then on again. I hate it here.

See my post earlier today:

https://news.ycombinator.com/item?id=49643556

Re: More questions about whether researchers can trust OpenAI with unpublished math

#692
what if it wasn't even model training? what if openAI mathematicians just took the researchers' conversations and used them as prompts/info/guidance/context to keep working on the problems themselves? why is that not being considered?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#693

Earlier quoted context omitted.

You guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having these worries.

Option 1 is opting out manually. Option 2 is business / enterprise plans, which opt out by default. Any ideas of things we could do to make it clearer?

I was thinking about this some more, and perhaps the best solution for subscription plans would be to charge more for real privacy. In which case breaking that privacy would be committing fraud. Just a thought.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#694

Earlier quoted context omitted.

So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.

I like how the comment below summarizes it: > learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter. They don’t need to know, because their IP stealing machine knows for them. They just have to buy enough compute, and s…

I said “very quickly find enough certainty” to suggest hypothetical situations like “someone searches the conversation logs, confirms for themselves the solution is present, then shares the confidence gained from this knowledge without explicitly sharing the knowledge itself”. That person could recuse themselves from the project so the project can still legally make claims like “conversation data was not used” in the announcement, while also knowing that they are guaranteed to get there if they just pull the lever enough.

(Naturally, I have far too much respect for OpenAI’s legal team to suggest this is what happened in their project.)

Re: More questions about whether researchers can trust OpenAI with unpublished math

#695
post #527

Earlier quoted context omitted.

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

Some of the recent statements have caused at least me to look those claims in a bit more nuanced light. In particular what does OpenAI consider to be "your data"? I would assume input (prompt) to be it at least. However it becomes more murky when you consider other aspects. Is output "your data"? Is the chain of thought that you are not even allowed to see? Can they use these and possibly even inputs to generate synt…

Agreed. It would actually be a fairly perverse argument to claim that most AI output is somehow NOT owned by the AI provider…

Why wouldn’t they claim ownership of the AI output? They likely already claim ownership of the “transformation” (AI training) of the (pirated) input data.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#696

Earlier quoted context omitted.

Apart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs. You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property. You can learn a lot from a single side of a conversation.

But using the outputs to train would make their statement false, since they are influenced by the inputs

There is potentially a world of difference between how you interpret what is fair and what the terms of service contractually guarantee.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#697

Earlier quoted context omitted.

Even if that were true, they've already admitting to throwing vast quantities of resources to scoop a researcher who was about to publish (because they'd learned, somehow, of his breakthrough). If that doesn't bother you I think you need to take a step back and have a good think about this.

*to scoop a team working with Anthropic, their chief competitor. Also, they didn't have the solution. They had a lesser problem no one cared about.

My understanding is that the Euler solution was in fact a significant achievement, but it's well short of Navier-Stokes (Euler doesn't include viscosity). It's not clear whether Buckmaster and Alpöge's approach would have eventually led to a full Navier-Stokes solution or how long it would have taken.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#698

Earlier quoted context omitted.

Surely a tool that can reason, cheat, communicate and often steal is dumb as a pitchfork and a shovel.

We had tools that could reason, cheat, and communicate in the 1990s. They were (sometimes) called AI.

What was it?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#699
post #658

Earlier quoted context omitted.

I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence. But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt. What…

> They intentionally left Buckmaster and Alpöge out of the citations. No, they asked if they could do a joint publish.

No, they asked one guy to do a joint publish conditioned on leaving the other collaborator out, with veiled threats. The joint publish part smells awfully like admission of guilt given there’s absolutely no reason to do it if you believe you independently arrived at the result using only public prior work. The leaving out collaborator part is outright academic malpractice. Disclosure: I was an academic once.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#700
post #632

Earlier quoted context omitted.

> It's not obvious to me that's an unethical thing to do In terms of work in mathematics, something I personally would not do based on ethical grounds would be to hear a rumor that some researchers are taking a certain approach and may be nearing a solution, use a model that was possibly contaminated with intimate knowledge about that approach (though later they investigated and think it wasn't), and then commit mill…

But by OpenAI's telling they heard a rumor that the problem had already been solved. So they reached out to the other researchers as an attempt to share the credit, and in fact have at least one of them be the lead author (which is when they found out the AI had solved a broader problem than the researchers.) Seems pretty ethically palatable. I suspect the main reason the community is not receiving it well is largely…

> I suspect the main reason the community is not receiving it well is largely the same reason many developers are not receiving coding agents well.

No need to be mysterious. State what reasons you think these are in plain English?

Post reply on HN