Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

271–280 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#272
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

I believe it can be "off by default" depending on terms negotiated between the enterprise customer and ChatGPT.

We have ChatGPT at work and it explicitly says that "workspace data isn't used to train models"

Re: More questions about whether researchers can trust OpenAI with unpublished math

#275

Earlier quoted context omitted.

If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.

Sam Altman has provided people with more generous usage than Claude or Gemini. There is no doubt ChatGPT is the most generous LLM provider!

Just because a guy is giving you free meth, doesn't make him generous.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#276
It would be really useful if the researchers disclose their notes and/or chats (or the key pieces thereof) so people can determine how close their work was to whatever the models produced.

I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#277

Earlier quoted context omitted.

Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that: 1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest. And also: 2) someone in the c-suite or marketing has decided to mitigate that thr…

> To me, dark patterns (like manipulative wording) imply that: Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.

Sorry, no. Explaining why someone would have taken them at their word is definitely not stupider than blaming people who could have been lied to for trusting a company that lied to them.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#278
post #247

The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes). People need to understand how all these AI company policies around training data work before working with them,…

> If it is found that OpenAI and other labs are not respecting the training opt out HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!

What do we need to audit? The researcher in question here did not opt out of training until a few months ago.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#279

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

Just use Bedrock...

Re: More questions about whether researchers can trust OpenAI with unpublished math

#280
post #10
post #9

Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts. Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are…

Curious, what other options are prospective pure math grad students considering?

At this school? Big four internships. Consider this a complete squandering of their potential (at least imo) These are mathematicians at top institutions, which is partly why they're being prodded for ideas, and getting the best and brightest to not take these consulting firms' offers was already a challenge.
Post reply on HN