Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

61–70 of 849 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#61
post #31
post #4

TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.

I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformation…

And humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#62
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.

For a long time it was clearly the former, but now I think it is the latter.

The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#63
post #27

[flagged]

We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal of egoless service. There's nothing wrong with that, either. But it's misanthropic to pass such harsh judgment on people for behaving in such a normal and natural manner.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#65

If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.

It looks to me more like they made a math engine that can sift through a huge number of combinations, most them absurd, to prove a statement. Just like a chess engine, but for math.

At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#66
post #27

[flagged]

HN loves drive-by downvotes. It's a real shame.

Downvotes might work as an abuse sponge, absorbing the impulse to make personal attacks. Other than that possible advantage, the downvote functionality seems contradictory to the concept of a discussion forum, I agree.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#67
post #51
post #37

All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.? The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all. That doesn't mean they can't be useful, or that their products a…

It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.

what the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees?

can you realize what this means?

focus on this part:

"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"

don't threat this as a minor dispute!

also why not nitter link? not even in comments?

https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#68

Only after reading this post did I learn that my preferred AI trains on my inputs (prompts). How was I not aware of this before?

Don't make this our fault. I would even ask how is this not off by default or why aren't we asked upfront about it if they really care. It's disguising data collection as good faith. I don't even understand how this is legal under GDPR/EU given how much of PII they receive through chats.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#69

If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.

If the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#70
post #5

Earlier quoted context omitted.

I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capt…

But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?

The difference is that Einstein didn't literally have someone prompting him towards his result.
Post reply on HN