Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

211–220 of 848 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#211

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

> Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.

Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#212
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

I refer you to this: https://news.ycombinator.com/item?id=49643556 Quoting: > "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

>"reset this more than once ... to my surprise I found it re-enabled"

At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.

What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#213
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

I refer you to this: https://news.ycombinator.com/item?id=49643556 Quoting: > "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

In all fairness this is someone saying something. Misremembering happens. Unless we have something with a bit more evidence, the simpler explanation suffices.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#214

[flagged]

What are you talking about?

I think they are referring to the Apple Watch 12 & 4 ultra feature announced yesterday. But I will note they are not on by default and has several configuration options.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#215

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

There's also the fact that the setting has been getting turned on by some users: https://news.ycombinator.com/item?id=49643556

Re: More questions about whether researchers can trust OpenAI with unpublished math

#216
post #160

Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.

What does "with no substantive human input" mean? All of the training data is human input, isn't it?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#217
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

"Improve the model for everyone" can be implemented in so many ambiguous ways. https://news.ycombinator.com/item?id=49643513

>>"Improve the model for everyone"

e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers

I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).

----

My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."

Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."

Re: More questions about whether researchers can trust OpenAI with unpublished math

#218

So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training? How do people become that trusting? The phrasing itself is guilt tripping

Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that: 1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest. And also: 2) someone in the c-suite or marketing has decided to mitigate that thr…

[dead]

Re: More questions about whether researchers can trust OpenAI with unpublished math

#219
post #160

Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.

Public domain doesn’t mean anyone can assert copyright. It specifically means no one can.

No, it means you can use that work in the creation of new works, which can indeed be copyrighted.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#220

Earlier quoted context omitted.

Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that: 1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest. And also: 2) someone in the c-suite or marketing has decided to mitigate that thr…

> To me, dark patterns (like manipulative wording) imply that: Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.

That's my take as well. At some point, they will claim that you cannot use their service without contributing back. If you quibble with them using your info in exchange for using their service, you don't get to use the service. Hence, I don't use their service. I do not trust these companies at all.

At this point, I'm left wondering what is wrong with me that I don't just go with the flow, otherwise, what's wrong with everyone that does.

Post reply on HN