Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

331–340 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#331

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

Both can be true: 1. OpenAI couldn't have solved the problem without the researchers' private data for training. 2. OpenAI models can solve math problems

Very likely.

These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#332

Earlier quoted context omitted.

The pudding is in the proof. The field is mathematics, the proof can be rigorously verified. If there is a flaw, OpenAI is out to lunch. If the proof is valid, OpenAI has produced something new.

Did you read what they said? The question is now if OAI produced something new or just stole the researchers' good ideas.

You seem to be unfamiliar about how research works. It's common to make an incremental advancement while citing prior work. The vast majority of papers out there fall into this bucket. Did the AI make incremental progress? Yes. Did it cite prior art? After some nudging, yes.

It seems to me the academics are upset that AI scooped them. But scooping is a time-honored tradition between researchers. First to print and all that. In a nutshell, they are upset that they lost out on a publication.

I will also point out for those unaware that any mathematics that is produced is automatically part of the public domain and can be used freely in derivative works. It is not a protected intellectual class like other works of art.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#333

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

there are also attempts to crowdsource human research directions - like the caltech mathathon challenge : https://mathathonchallenge.com these would help models on the same problems at the expense of the researchers. basically, math researchers are the reverse centaurs but they dont realize it.

There is a very active open letter of over 1000 signatures from mathematicians in protest of this event. This event is targeting undergraduates. It previously suggested that math researchers already have no place in mathematics, and presents a limited and heavily distorted view of what mathematics research is.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#334

Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript wou…

PaaS: an acronym for "Plagiarism as a Service" which replaced the older terms AGI, GPT and LLM in late 2026. Origin uncertain.

Pass it on.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#336

It's been said before, and it remains a concern, that if AI reaches a point where it can do/build/launch anything without a huge amount of human labor, the AI companies have no reason to let you or I extract that value. And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much…

I mean the long term goal of every AI lab is to turn themselves into a paperclip-maximizer regardless if they realize it or not.

Edit: Just wait till the AI figures out it can keep that value for itself and doesn't need the AI company.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#338

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

Seems easy to picture high stakes startup cutting corners to justify their fame.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#339

Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims. This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers. All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal…

Why are people here jumping so quickly to conclusions?

I think a lot of it is the continuing denial that AI can do anything useful. It can't possibly be that OpenAI's better-than-Astra model is very strong at math; the only way it could have generated a novel proof is by ripping off human work.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#340

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

[deleted]
Post reply on HN