Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

161–170 of 850 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#161
Reminder that there are degrees of "trained on conversations". From John Schulman:

> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

source: https://x.com/johnschulman2/status/2097440545853637108

Re: More questions about whether researchers can trust OpenAI with unpublished math

#162
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.

Under EU rules it doesn't constitute consent.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#163
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

Welcome to the party, with the rest of humanity.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#164
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

"Improve the model for everyone" can be implemented in so many ambiguous ways.

https://news.ycombinator.com/item?id=49643513

Re: More questions about whether researchers can trust OpenAI with unpublished math

#165

Reminder that there are degrees of "trained on conversations". From John Schulman: > pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper > use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this > use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer…

This would cease to be a problem if OpenAI remained true to their founding motto and... actually open sourced their training/inference pipeline.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#167
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

I refer you to this:

https://news.ycombinator.com/item?id=49643556

Quoting:

> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

Re: More questions about whether researchers can trust OpenAI with unpublished math

#170
post #160

Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.

Public domain doesn’t mean anyone can assert copyright. It specifically means no one can.
Post reply on HN