Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

531–540 of 850 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#531

Earlier quoted context omitted.

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

The open internet is now a cesspit, with very little new good data. Expect everyone to train on user data always. They just got clever about whitening it.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#533
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#534
post #434

Earlier quoted context omitted.

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

i think this line of argument is outdated

Care to explain why?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#535
post #491

Earlier quoted context omitted.

I'd wager a fair chunk of my money that money breaks OpenAI before OpenAI breaks money.

OpenAI != AI. If you were in 1999 you'd be saying pets.com = internet.

I think this leads to an interesting question. What happens when the money runs out?

Right now, a lot of money is going to train new models. And we need to train new models because they get gated by their training data. And models are only as useful as their training data.

So let's say the money stops.

Do we stop training models? Do we train them slowly? Do we accept the then current models as the limit?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#538

Earlier quoted context omitted.

Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems

You forgot possibility 3: OpenAI solved the problem without using any private training data from the two researchers. Everyone in this thread seems to have made up their mind about OpenAI's guilt though.

If the new model is that good, and is chewing through open problems at an unprecedented rate, the smart move would have been to let the humans have their W on this one and present solutions to those other problems.

Especially if there really is a long list of them.

"Here are a few hundred proofs" is far more convincing than "We really Navier Stokes and coincidentally someone else did too but we don't know the details or anything, who us, definitely not."

It's a PR fiasco, and a cynic might wonder if it's entirely about the IPO.

I'm consistently entertained by how these companies, with the most advanced models on the planet, consistently do the most idiotic things.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#539

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

OK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much. 1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick. 2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be n…

1. I would agree if the rumours were that some mathematician(s) had solved them, but the rumors alleged it was Anthropic. I don't really see what the big deal was. They had a new model that was going along great and wanted to test its mettle.

2. Yes Brubeck's comments were weird at face value. That said, Open AI's proof isn't a duplication of anything. Not only is Tristan's work a sub problem but the methods are different. And what OpenAI didn't want was Levant on the paper OpenAI authored not whatever they were working on (Euler). It's petty sure but it's fair enough. Tristan and Levant didn't have anything to do with the Navier Stokes solution, so it's really their call if they didn't want to collaborate on their own paper with the Anthropic employee.

>OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute,

$20M in approximated API prices doesn't mean they spent $20M worth of compute. The real number would obviously be substantially less.

>and the assistance of a whole team of people at OpenAI

You can't eat your cake and have it. What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one.

>I'd say advantage humans this time....to work on problems that have not been solved yet, and that humans are NOT making nice progress on.

Interesting way to frame progress that didn't move along till an LLM generated proof.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#540

Earlier quoted context omitted.

> but right now there's no credible evidence, only claims. since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried. you're being naive

OpenAI has said that their models were definitely not trained on any of Buckmaster's sessions after July 3rd (from https://archive.ph/75WcF ); likely they found that's when he switched the "allow training" setting off.

It's also possible he shared drafts of the work with someone else, who asked chatgpt to explain it to them with training on. Tao seemed to know lots of details of the work before anything was published, though also worked on the problem in the past with big results so maybe just guessed.
Post reply on HN