Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

131–140 of 844 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#131

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

> they could be significantly piggybacking on human progress,

This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#132
post #5

This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point". They don't even claim to have had a proof, only to have been working on it.

I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capt…

> capturing large amounts of data and connecting the dots.

This is what research is; collecting data and connecting the dots.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#133
post #37

All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.? The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all. That doesn't mean they can't be useful, or that their products a…

They have throughout this period of AI products shown to reproduce works that they were trained on. They are getting sued all over the place for the theft of content right now and it seems courts and governments want to wave copyright protection (and ignore criminal acts because the "ai did it") to see where this leads.

Its why I stopped writing open source software, my code was stolen and put behind a paywall and the license under which it was published has not been adhered to. Doing work in the public domain at all now is just stupid, these companies are allowed to steal it and call it their own.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#134
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#135
post #33

It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.

I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence.

But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.

What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations.

Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse.

[^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#136
post #34

I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?

That doesn't stop them from training on your data apparently. I have that disabled but still has to disable "Don't train on my data" in the privacy center too. https://privacy.openai.com/policies?modal=take-control

I think that flow is an easy way to disable everything, so there isn’t a risk of forgetting to flip one thing back off after accidentally setting it on. I set my ChatGPT environment to allow model improvement for example but had to check my codex settings to make sure ‘Include environments’ for model improvement is off.

I think if I had both on and turned off the ChatGPT setting, ‘Include environments’ has a chance of still being flipped on.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#137

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

> they could be significantly piggybacking on human progress, This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.

That’s one perspective.

I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.

No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.

I’m very pro AI long term btw but I’m not blinded.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#138

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

I’d argue the invitation of researchers was incredibly strategic.

Sam Altman knows what he’s doing. He will happily screw these folks to one-up his competition.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#139
post #37

All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.? The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all. That doesn't mean they can't be useful, or that their products a…

They have throughout this period of AI products shown to reproduce works that they were trained on. They are getting sued all over the place for the theft of content right now and it seems courts and governments want to wave copyright protection (and ignore criminal acts because the "ai did it") to see where this leads. Its why I stopped writing open source software, my code was stolen and put behind a paywall and th…

Yes open source code was the first - it’s what has got Anthropic and OAI its revenues from selling outputs associated with producing code.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#140
post #104

This is the second wake up call. Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”. Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge…

This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price? They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at…

In the short run it’s fantastic if it means that folks will feed in enough inputs from a wide array of software that can eventually replicate software with smaller teams than historically.

Why? Competition. In the long run imagination will win out.

No firm has the divine right to exist - it must earn its existence.

What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though.

Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages.

Interesting times ahead.

Post reply on HN