Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

371–380 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#371

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, th…

OpenAI have come out and said:

>The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”

>The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#372

Earlier quoted context omitted.

>> Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] Maybe I'm failing to read that graph properly but the y axis says "pass rate" and it only goes up to 0.5. That would mean every single problem is at most half-solved. I don't know what that means though. What is "0.5 pass rate" in the context of "open math problems" (as in the graph title)?

I guess it's a fraction of problems on which a model produces a LEAN proof or a counterexample.

Wouldn't they just list the number of problems solved then?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#373
post #329
post #314

Earlier quoted context omitted.

> What we don't know is just how recent this model was, and therefore what it may have been trained on. OpenAI's statement says that they began training their new model on August 28.

omitting when training concluded edit: ffsm8 makes a great point below, it doesn't matter. I'm not great with dates, sorry.

[deleted]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#374
post #309

Earlier quoted context omitted.

Yes. See https://www.anthropic.com/research/small-samples-poison?from... . 250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size. I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the mod…

Thats not what they are asking. This paper is discussing documents in the training dataset poisoning the LLM for malicious behavior. This person are asking if anyone has deliberately put something in a private chat (presumably with retrain on my data turned off), to see if they can get it to leak across sessions from distinct users. I am positive this happens but I have not seen the proof. I also want to know the ans…

* with train on my data turned ON, yes. Though OFF would of course be even more notable!

Thank you – the non-adversarial reproduction paper ( https://arxiv.org/abs/2411.10242 ) nails it – from chat, to training corpus, to subsequent model. Though in my hasty read, it is not entirely clear whether the snippets it finds are nonces, i.e. present exactly once in the internet.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#375

https://x.com/markchen90/status/2097400166554993041?s=20 that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

Not sure what you are seeing in that tweet that gives you the impression that the toggle does nothing.

They say they train on your “deidentified data”

Passing your output through a second model and telling it to remove identifying data would count as “deidentified”

So they could scrape all the IP in your company as long as they take the names out first…

Re: More questions about whether researchers can trust OpenAI with unpublished math

#376

Earlier quoted context omitted.

OpenAI said they sicced this agent army on Navier-Stokes on Sept 1st, while only a couple of days earlier OpenAI's Noam Brown happened to reply to a tweet saying that they had already tried to solve all the Millennium Prize problems and failed... So, it seems either the previous attempt didn't have the training to succeed, or was just not given the compute to do so. Once OpenAI heard that Navier-Stokes was solved, th…

OpenAI have come out and said: >The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” >The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way…

This is literally "We have investigated ourselves and found no wrongdoing"

Why should we trust them?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#377

Earlier quoted context omitted.

This is absolutely not "common sense opsec". If I type information about some proof I'm exploring into a Google Doc, I do not worry even a tiny bit that the Docs team might forward it to a team of advanced mathematicians in case they have an advanced technique they want to show off by scooping me. That would be a crazy thing to do, nobody would even consider it, and if it happened Sundar would fire everyone involved.…

That’s … Googles entire reason for making these “you don’t pay with money” tools. Did you not understand that?

Google docs exists as a competitor to microsoft office. Google gives it away free to consumers for the same reason AI labs sell subscriptions for 10% of the price of the api. They hope businesses will switch to what employees know how to use.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#378

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

The leakage wouldn't be from training, but from other uses of Personal Data.

As far as I understand it, users can opt out from the training aspect, but they cannot stop their conversations (“User Content”) being used “[t]o improve and develop our Services and conduct research, for example to develop new features”.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#379
post #360

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

On your second point: there is a more plausible explanation which David Bessis calls the "overhang". The short version is that there is a large amount of relatively low hanging fruits in mathematics, because no human has broad enough knowledge and enough time to try them all. AI is not constraint by that, and therefore can systematically pluck all those low hanging fruits. Quote: "The Overhang consists of the unreali…

It's like AlphaGo but playing against all living mathematicians. (Overhang being low hanging fruit is what allows this comparison, of course the general moot point is the skepticism that LLMs are also innovative etc.)

Re: More questions about whether researchers can trust OpenAI with unpublished math

#380

Earlier quoted context omitted.

There is a very active open letter of over 1000 signatures from mathematicians in protest of this event. This event is targeting undergraduates. It previously suggested that math researchers already have no place in mathematics, and presents a limited and heavily distorted view of what mathematics research is.

I'm an AI skeptic, but I don't see how this squares with what the organisers of the event actually say. "It previously suggested that math researchers already have no place in mathematics"? I don't see this.

Also I want to mention that the letter is not about AI skepticism, in any direct way at least.
Post reply on HN