Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

341–350 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#341
post #140
post #104

Earlier quoted context omitted.

This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price? They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at…

In the short run it’s fantastic if it means that folks will feed in enough inputs from a wide array of software that can eventually replicate software with smaller teams than historically. Why? Competition. In the long run imagination will win out. No firm has the divine right to exist - it must earn its existence. What OAI and Anthropic have shown is they can accumulate all the information in the world - they still…

Looking at the current behavior of AI swarms this is going to be 'fun'.

AI: Hmm, I'm running out of new ideas, how I can I make more?

AI: Well, it takes a shitload of energy/tokens to do that, or I could just steal them.

AI: [proceeds to hack the shit out of everybody stealing all the data it can]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#342

People saying “he should have opted out” are missing the point. OpenAI can and should check their training data for leakage in the face of big breakthroughs like these. It’s the burden of the author to appropriately cite their sources. It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.

We checked and determined it was impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.

If prompts were submitted earlier than that and training was not opted out, there's a chance they made their way into our training pipeline in some form. But this would be a droplet in an ocean and unlikely to have made any difference, imo.

See: https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#343
Most scientific breakthroughs are simply a continuation of previous work.

I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.

Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#344
post #16

Everything you say can and will be trained against you

This is the scary part - your most novel thoughts and breakthrough ideas being slurped up and regurgitated as if they were the AI’s creativity Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine

It's not at all scary. I know some of my ideas are poorly remembered copies of other people's work. Whenever I'm trying to build something, I spend time going through technical journals on the topic to see who invented it first and what they discovered that I haven't figured out yet. It's amazing how hours in the library save you days of beating your head against the wall.

I suggest looking at the past history of IP disputes. Humans have been "slurping up and regurgitating ideas" for a very long time. There are lots of examples of parallel creation, rediscovering old ideas independently, telling an idea to the wrong person, and having them claim credit for it.

- Newton/Leibniz clash over who invented calculus. - Niccolò Tartaglia vs. Gerolamo Cardano clash over the formula used to solve cubic equations. This was also an independent rediscovery, as Scipione del Ferro discovered and published the formula earlier. - There are multiple literary works in print, music, and film that have competing claims. - Meccano versus Erector Set: developed about 20 years apart in England and the United States. Unclear if it's independent invention or copied. US developer Alfred Carlton Gilbert claims he was inspired by steel girder construction of infrastructure.

also https://community.thriveglobal.com/10-famous-inventions-that...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#345

Earlier quoted context omitted.

Did you read what they said? The question is now if OAI produced something new or just stole the researchers' good ideas.

Name one discovery ever that didn't depend on someone else's work.

Most discoveries did not happen by someone looking in someone else's notebooks without them knowing.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#346

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

I feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems

The brute-forcing is a good, old-fashioned generate-and-test approach like in Simon and Newell's Logic Theorist, which was presented in the Dartmouth convention in 1956, where AI was named by John McCarthy. Logic Theorist caused a big stir by (re) proving several of the theorems in Principia Mathematica by Russel and Whitehead.

There was much excitement, then, as now, for this kind of approach and there were several systems that followed along the same lines, e.g. Automated Mathematician by Doug Lenat.

Eventually it became clear that this approach is limited by what it can generate: you may have a sound and complete verifier, but if the generator, i.e. the first step in the generate-and-test pipeline, is incomplete, then the entire thing will run out of steam sooner or later.

The difference with LLMs is that they are... well, large. They are the most powerful generators ever created. That means their limits are not in sight and it will probably take us a very long time to find them.

Which is all to say that, yes of course, automatic verification is indispensable. But without an LLM generating an unprecedentedly large number of plausible theorems, there would be no AI mathematics, or in any case AI mathematics wouldn't have gone as far as it has.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#347

Earlier quoted context omitted.

Imagine OpenAI Astra model weights were made public because the datacenter they use had T&C that allows them to make them public Would that be ok in your mind? Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did) Nobody would care if they provided published research that author made public same as a google search would offer that.

Would that be ok in your mind? It would in my mind. Hopefully companies have looked through the agreement.

well then fingers crossed someone does that

Re: More questions about whether researchers can trust OpenAI with unpublished math

#348
I hate to be that dinosaur, but this is exactly what I’ve been warning about since XaaS began taking off in the mid-oughts: any provider you use can and will use your data for their own benefit regardless of any contracts or safeguards in place, especially if the benefits outweigh the consequences.

Honestly, I’m surprised it took this long for some company to really go all the way, though. OpenAI really making it transparently clear that they can and will do whatever they want with the data you provide them, contracts or settings be damned. Completely untrustworthy as an entity, full stop.

Of course, I’m also too jaded to think this will change anything. Folks will move to Anthropic, or Gemini, or Grok, or some other hosted model on a pubCSP managing the harness and logs for them, and then do another shocked-Pikachu face when it happens again.

If you aren’t running workloads on infrastructure you own, then your privacy, security, and general outcomes are at the sole whims of the hosting provider - who can and will fuck you over the exact second it’s more beneficial for them to do so than the loss of trust incurred.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#349
post #293

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

If your rumor is true, what we are witnessing is a giant paradigm shift rather than individual incidents. Mathematicians were the first victims of super-intelligence. Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.

They "burn" a lot when they do benchmarks, while these runs can become valid roll outs for training. Perhaps less efficient than other data creation, but hardly burned in the same way.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#350
post #18

It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.

Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyb…

You are both using a different definition of sharing I believe. When people have an expectation of privacy, use by others should be forbidden. Tech has gone completely off the rails with the use of private data.
Post reply on HN