Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

431–440 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#431
How to steal ideas with AI.

step 1, identify high value users by net worth, citation count, or number of followers

step 2, select all prompts by high value users

step 3, invest 10 billion thinking tokens in modeling an objective for each user

step 4, build an RL environment for each user

step 5, rollout 10 billion tokens per environment

step 6, train on resulting traces

Re: More questions about whether researchers can trust OpenAI with unpublished math

#432

When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1] https://mathoverflow.net/a/513885

That link is a helpful contribution to this discussion. I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work: > It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-componen…

> On the other side, I was looking myself for such a mechanism ever since we wrote the paper in 2019 and admire the efficiency of this construction.

He seems to admit very clearly he does not see this as his own work. 'I was looking...' well why did he stop? Because the AI figured it out first.

It seems quite odd to me to 'admire the construction' of something, only for your opinion to sour once that something figures it out first.

I think a lot of the emotional reaction here is familiar to us non mathematicians: you spent years developing expertise, and then LLMs began producing competent work in areas that had previously required that expertise. That's understandably uncomfortable, but discomfort by itself isn't evidence of misappropriation.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#433
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

> we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

implied the humans sessions could have been (and probably were, why wouldn’t they be?) in the training set?

If I was trying to make a model smarter and I had transcripts from the smartest mathematicians in the world I’d make sure the model trained on them.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#434
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

> I think it's a useful analogy to compare OpenAI to a human collaborator.

Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

Re: More questions about whether researchers can trust OpenAI with unpublished math

#436
post #376

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

This is literally "We have investigated ourselves and found no wrongdoing" Why should we trust them?

Reputational risk- if they lie about this and get caught, it will have billion dollar implications for their business.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#437

How to steal ideas with AI. step 1, identify high value users by net worth, citation count, or number of followers step 2, select all prompts by high value users step 3, invest 10 billion thinking tokens in modeling an objective for each user step 4, build an RL environment for each user step 5, rollout 10 billion tokens per environment step 6, train on resulting traces

step 7. get away with it as no other party has enough capital (tokens) to prove such infringement ever happened

Re: More questions about whether researchers can trust OpenAI with unpublished math

#438

If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.

If the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.

It’s a prisoner’s dilemma. A single mathematician working with AI while all others forebear will clearly outcompete.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#439

Earlier quoted context omitted.

I guess it's a fraction of problems on which a model produces a LEAN proof or a counterexample.

Wouldn't they just list the number of problems solved then?

rates beat counts almost always.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#440

Earlier quoted context omitted.

OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, th…

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months

two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?

Post reply on HN