Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

471–480 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#471

Earlier quoted context omitted.

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

Apart from the well-known dubious position of OpenAI wrt truth, the prompts/inputs do mot include the outputs. You can train on a sequence of outputs. In the end, OpenAI outputs are OpenAI's property. You can learn a lot from a single side of a conversation.

But isn’t Tristan’s breakthrough happens in August? OpenAI can’t really train with text that doesn’t exist

Re: More questions about whether researchers can trust OpenAI with unpublished math

#472
post #48
post #27

[flagged]

Let's turn this question around. If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?

Sort of like Apple takes validated market products and snipes the last steps to an actual good UX (at least in theory)?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#473
post #469
post #434

Earlier quoted context omitted.

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

Tools don’t turn around and scoop you. What OpenAI did here was use the same tool that the researcher did which might have coupled their work together.

You’re right, they don’t. It was scooped by the humans at OpenAI who published the paper. The tool they used to do it isn’t that relevant.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#474
The contamination framing is a proxy for a deeper problem: we have no tools to track the provenance of ideas in model weights.

OpenAI saying they "cannot rule out" training on user data isn't a hedge. It's an accurate description of the epistemic situation for anyone in their position. Current interpretability methods can't answer questions like "did this proof technique originate from training on Session X?" The ideas in a model's weights don't have clear lineage -- they're smeared across millions of examples in ways we can't localize. This is different from citation in human research, where influence is presumed to flow through legible chains (reading, citing, corresponding). In a trained model, the nearest equivalent to "you read their work" is undetectable.

Lean makes this worse, not better. It verifies that the proof is correct, but provides zero information about its intellectual genealogy. So OpenAI now has a proof that is formally verified and provably mysterious about its origins. The "we cannot rule it out" statement is the honest answer, but it's also an answer that can never become more certain in either direction with current tools.

The researchers are pointing at something structurally new: the normal academic attribution apparatus depends on influence being legible. If AI intermediaries can soak up ideas from private conversations, synthesize them, and produce outputs that are formally correct but intellectually unattributable, we don't have norms for that situation yet. This specific case may or may not involve misconduct. But the structural problem it reveals exists independently of OpenAI's behavior.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#475

Earlier quoted context omitted.

“ On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.” - https://openai.com/index/navier-stokes-solution/ They do not explicitly admit to knowing about NS specifically, but are e…

So then they DIDN'T "learn the secret to cracking the problem". They simply knew that part of the problem was solved. Knowing a problem can be solved and knowing the solution are not the same thing.

So because they didn’t admit to it they didn’t do it?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#476
post #419

Earlier quoted context omitted.

> were the first victims Spinning it negatively like that doesn't do anybody good. Were mathematicians the "victims" of calculators? of Matlab? Were writers the ""vIcTiMs"" of word processors?? (apparently yes, according to old TV shows about computers during the 1980s, that you can see on YouTube) > "tHiS iS nOt ThE sAmE" — Everyone every time. No, just look it up. Look into old magazines and TV shows or newspaper a…

It’s not the same. AI potentially completely replaces intellectual work without creating any* new jobs (*almost any - there will be some extra jobs for building data centers but that’s negligible).

> without creating any* new jobs

So fucking make it so that people don't -need- "jobs"

It's about fucking time already.

Don't fucking try to hold back electricity just so people still have to manually light street lamps to earn food and shelter: https://en.wikipedia.org/wiki/Lamplighter

Re: More questions about whether researchers can trust OpenAI with unpublished math

#477
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#478

Earlier quoted context omitted.

You can read their Wikipedia page [1]. [1] https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_is...

These examples aren't really similar. None of those situations involve harming and lying to their own customers.

Nonsense; the claim was that they wouldn't do anything that would mean they'd be

> risking massive lawsuits and a total loss of trust

Evidence of the massive lawsuits and lack of trust seems pretty relevant.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#479
> The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

> The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

https://archive.is/lWzkk

Re: More questions about whether researchers can trust OpenAI with unpublished math

#480
post #434
post #412

I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical. Now, OpenAI is claiming that the model…

> I think it's a useful analogy to compare OpenAI to a human collaborator. Frankly I don’t buy this. It’s not a human or a collaborator. It’s a tool. This is like saying it’s not Microsoft’s fault if they extract a bunch of data from people’s Excel sheets because they willingly put it into the program. Anthropomorphizing software is ignorant and foolhardy

i think this line of argument is outdated
Post reply on HN