TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformation…
More questions about whether researchers can trust OpenAI with unpublished math
61–70 of 848 posts
Re: More questions about whether researchers can trust OpenAI with unpublished math
#62For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#63[flagged]
Re: More questions about whether researchers can trust OpenAI with unpublished math
#64Re: More questions about whether researchers can trust OpenAI with unpublished math
#65If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#66Re: More questions about whether researchers can trust OpenAI with unpublished math
#67All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.? The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all. That doesn't mean they can't be useful, or that their products a…
It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.
can you realize what this means?
focus on this part:
"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"
don't threat this as a minor dispute!
also why not nitter link? not even in comments?
https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...
Re: More questions about whether researchers can trust OpenAI with unpublished math
#68Only after reading this post did I learn that my preferred AI trains on my inputs (prompts). How was I not aware of this before?
Re: More questions about whether researchers can trust OpenAI with unpublished math
#69If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
Re: More questions about whether researchers can trust OpenAI with unpublished math
#70Earlier quoted context omitted.
I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well. To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capt…
But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?