Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

591–600 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#591
post #526

Earlier quoted context omitted.

OK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much. 1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick. 2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be n…

> and that humans are NOT making nice progress on They've pretty much said their own work was heavily agent driven. Levent is in a particularly bad place here because while he probably had a lot of background in the Jacobian Conjecture problem, he made the solution to that one sound like someone asked the question and he just fed it to Fable during the world cup. Whether that nonchalantness was to just seem hip or wa…

I was referring to the overall pattern of apparently sniffing around for recent mathematical progress then setting the AI on it to see if the problem is now easy enough to solve (if you have the money).

Terrance Tao has lamented this practice as being unhelpful for mathematics, and likely to lead to humans working in private to avoid this.

Tao has also noted that many of these AI math proofs don't really help mathematics (nor does it seem they are intended to), since for many of them the proof was never the point, it was the math expected to be needed to be developed along the way, which the AI solutions don't provide.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#592

Earlier quoted context omitted.

I put it like that because of Brubeck's own words on the matter. You're acting like we've gotten email receipts here. I'm not really interested in going over a he-said she-said about strangers.

Brubeck has admitted what he said, but claims he immediately retracted it as a "poor choice of words". Given Buckmaster's telling, this seems beyond "poor choice of words"... It was a veiled threat, that he then doubled down on with his "If you don’t want me to be nice, then I don’t have to be nice." follow-up. ** I said that if OpenAI released its result in the way proposed I would go public with what happened. The…

Fair enough then. I had only seen some earlier comments.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#593
post #434

Earlier quoted context omitted.

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

Surely a tool that can reason, cheat, communicate and often steal is dumb as a pitchfork and a shovel.

Nah, that just makes it a shitty tool.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#594
post #567

Earlier quoted context omitted.

I heard they also tried to strong-arm them into removing the name of their collaborator who happened to work at a different company (Anthropic)... I haven't looked into it myself, but if true, that seems incredibly scummy.

The flip side is that the Anthropic researcher is clearly pushing the case against Open AI and one might have reason to suspect their motivation and version of events for the same reasons. What a mess.

> the Anthropic researcher is clearly pushing the case against Open AI

Things like the nytimes interview are with Buckmaster, who works at NYU, not Alpöge. I saw a couple of tweets from him over the last week. Any chance of clarifying what makes you think he's "clearly pushing the case"?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#595
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

> Provenance is hard to track

Right, which is going to open a lot of doors to a lot of questions.

I don't think there's any legal ramifications on this, just ethical ones about when and how you publish research, but it's yet another point in favor of "if provenance is hard to track, should we be using this for things where it needs to be".

Obviously copyright/trademark is a huge discussion on this, and I could absolutely see this devolving into that as well with how certain findings wind up monetized.

We have a response in this topic from someone claiming to be from OpenAI and linking an article where they, roughly, say "we are sure nothing from the 2 month period made its way into the solution". If that is true, that should mean it is provable, but leads to some more open ended questions like "well what data did it use then?". Is this still okay if someone close to the author did plug data into open AI and it extrapolated it?

Obviously that's probably an unreasonable expectation for these models to track and prove, but it also used to be an unreasonable expectation to scrape every single piece of digital and physical info for consolidated data.

If I opine to a friend on a park bench about a story I'm writing, do they get to pull it from the flock feed, shove it in the model, and then provide it to disney?

Legally, right now, probably. But there's going to need to be a serious look at laws and standards. Or a major shift in what is and isn't discussed in public if literally every breath and move you make can become monetized.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#596

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

We were also curious and we looked further into this. We've determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. This goes beyond what we said earlier, when we were less sure. If prompts were submitted earlier than that and training was not opted out, there may be a chance they made their way into our training pipeline i…

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#597

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

Generative models copy training data verbatim and also generalize, the two are not mutually exclusive.

You can have a look at the literature on exact copying in image models if it interests you, but just online we often see online examples of agents outputting code that already exists, even if its not the common case.

I very much doubt OpenAI points the model towards a specific conversation, but these trillion parameter models can very much "remember" their training data. For instance I can ask GPT to summarize my papers from their title alone, without looking them up, and it works decently.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#598

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?

Just knowing that there had been progress is enough to have an idea that throwing more compute at it might work (OpenAI had previously tried all the Millennium Prize problems with somewhat limited compute and failed).

It's comparable to Magnus Carlson saying that if he wanted to cheat, all he would need would be for someone to tell him to spend more time thinking about a specific move (just a wink would be enough) as an indication that a computer had found something interesting.

It's as-if after OpenAI first failing on Navier-Stokes (which OpenAI had just tweeted about 2 days earlier!), someone winked at them and said "you might want to try a little harder ...".

Re: More questions about whether researchers can trust OpenAI with unpublished math

#599

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

We were also curious and we looked further into this. We've determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training. This goes beyond what we said earlier, when we were less sure. If prompts were submitted earlier than that and training was not opted out, there may be a chance they made their way into our training pipeline i…

This may be true but nobody trusts your employer. The shadiest drips downward too, with the mob-like way they treated Dr. Buckmaster.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#600

Earlier quoted context omitted.

These examples aren't really similar. None of those situations involve harming and lying to their own customers.

Nonsense; the claim was that they wouldn't do anything that would mean they'd be > risking massive lawsuits and a total loss of trust Evidence of the massive lawsuits and lack of trust seems pretty relevant.

By "loss of trust", I meant that this is something they would risk losing a lot of users over, which isn't the case with the other lawsuits. There is a massive distinction between fighting third parties in a legal grey area and committing blatant fraud against your own users. Even if you have no regards for ethics, intentionally shipping a noop "do not train" toggle offers negligible upside for a massive downside.
Post reply on HN