Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

111–120 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#111

When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1] https://mathoverflow.net/a/513885

That link is a helpful contribution to this discussion.

I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work:

> It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward.

Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#112
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.

The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#113
post #51

Earlier quoted context omitted.

It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.

AI is America's last chance to salvage its empire. Nothing will be allowed to impede it.

But I don't see how. AI is going to be a commodity in short order and best case the US will be a temporary leader in the supply of tokens. Meanwhile AI is going to destroy much of the Service and Software industry that make up most of the US economy. And the US is betting every last cent to bring about this future. It does make sense for Trump since this might be a sugar high that lasts till the end of his term.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#114
post #112

One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR. The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversatio…

No, if they want they can easily compare the strings verbatim because these exact phrases are so extremely rare that it almost certainly isn’t in other conversations.

But of course they wouldn’t do it. Why would they?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#116
post #18

It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.

[dead]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#117
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.

The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.

In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#118
post #51

Earlier quoted context omitted.

It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.

AI is America's last chance to salvage its empire. Nothing will be allowed to impede it.

[dead]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#119
post #98
post #62

I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal. For a long time it was clearly the former, but now I think it is the latter. The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agent…

I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.…

10000000% Correct.

I’ve been working on a novel project for 1 year.

I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now.

Get stuffed OAI.

Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres.

Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#120

Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims. This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers. All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal…

Frankly, these mathematicians have more credibility than the sociopaths running OpenAI
Post reply on HN