Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

521–530 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#521

This smells of extreme 'cope'. Am I really supposed to believe that all of these problems could have been solved, were right about to be solved, etc. But it just happens they are all getting solved now when AI is getting really good at Math...

They’re not necessarily claiming the solutions were imminent. The big questions here from a mathematician’s perspective are to what extent these LLM systems are discovering conceptually new ideas compared to merely combining and pursuing known frameworks.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#522
post #34

I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?

That doesn't stop them from training on your data apparently. I have that disabled but still has to disable "Don't train on my data" in the privacy center too. https://privacy.openai.com/policies?modal=take-control

Even that sort of thing they could just as easily go 6 months from now

"oopsie guys, turns out our vibe coded "don't train on my data" toggle was just flipping the ui asset not changing any underlying boolean flag associated with your account. sorry but all that stuff is in the training set now and we don't know how to get it out either and no we won't be doing a 6 month rollback."

Re: More questions about whether researchers can trust OpenAI with unpublished math

#523

Earlier quoted context omitted.

First, OpenAI is not claiming that the model wasn't trained on those sessions. What they've said is “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.” and “We did not use their prompts or proofs to prompt our models or direct our agents.” and “While unlikely, we cannot…

As with all press releases I assume it was written/re viewed/redacted by their lawyers, so: > no specific user data was accessed in order to solve this problem Data was accessed in order to (and then accidentally used in training) Also, is llm’s answer to the prompt actually “user data”? > We did not use their prompts or proofs … So they used llm’s answers to those prompts. > … to prompt our models or directew our ag…

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#524
post #434
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

Surely a tool that can reason, cheat, communicate and often steal is dumb as a pitchfork and a shovel.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#525

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?

When reading human comments, we should be generous; when we read corporate texts, we may assume paltering.

(TIL: paltering: exact and technically correct statement usage to create misleading impression)

Re: More questions about whether researchers can trust OpenAI with unpublished math

#526

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

OK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much. 1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick. 2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be n…

> and that humans are NOT making nice progress on

They've pretty much said their own work was heavily agent driven. Levent is in a particularly bad place here because while he probably had a lot of background in the Jacobian Conjecture problem, he made the solution to that one sound like someone asked the question and he just fed it to Fable during the world cup. Whether that nonchalantness was to just seem hip or was to promote Anthropic, which he has stock in, or was just the truth I don't know though. But it makes this one seem similar, when they might have had really had nearly a year of very valuable feedback to the models.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#527

Earlier quoted context omitted.

When you say "confidential in-house version", what are you referring to? Local models? Bedrock deployment with "guardrails"? A different thing?

Enterprise Agreements can have binding terms for this. When I launch the ChatGPT desktop app, and open the options pane it says "Corpname data is not used for OpenAI training". I would expect academic institutions to require equivalent contractual terms.

Some of the recent statements have caused at least me to look those claims in a bit more nuanced light. In particular what does OpenAI consider to be "your data"? I would assume input (prompt) to be it at least. However it becomes more murky when you consider other aspects. Is output "your data"? Is the chain of thought that you are not even allowed to see? Can they use these and possibly even inputs to generate synthetic data that is then used?

All of these would seem to be "your data", but when they are carefully only including certain aspects (like prompts) in their statements it starts to sound they want to hide something.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#529
post #419

Earlier quoted context omitted.

It’s not the same. AI potentially completely replaces intellectual work without creating any* new jobs (*almost any - there will be some extra jobs for building data centers but that’s negligible).

> without creating any* new jobs So fucking make it so that people don't -need- "jobs" It's about fucking time already. Don't fucking try to hold back electricity just so people still have to manually light street lamps to earn food and shelter: https://en.wikipedia.org/wiki/Lamplighter

Ok I’ll make it so, you’ve convinced me.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#530
post #434

Earlier quoted context omitted.

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

i think this line of argument is outdated

[deleted]
Post reply on HN