Live data from Hacker News

More questions about whether researchers can trust OpenAI with unpublished math

mathstodon.xyz

281–290 of 845 posts

Re: More questions about whether researchers can trust OpenAI with unpublished math

#281
post #27

[flagged]

We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal o…

Identity isn't the question. Eating is the question. If you can't come up with things you don't eat. If you come up with 90% of things and some overarching parasitic process comes in, puts in the 10%, and now they get 100% and you get 0%, you don't eat.

The problem is that AI is capital, and having to rent AI to keep up when it can just steal your mostly done work is something somehow even lower than wage-labor. They can use your own risked investment (the cash you paid to work) to get out in front of you and take credit.

I have yet to trust LLMs with anything important that can be capitalized on. I only use it to work on projects that if they stole and expanded on them, I'd actually be happy to see.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#282
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

I refer you to this: https://news.ycombinator.com/item?id=49643556 Quoting: > "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

Has anyone else seen this happen? I checked and my setting is still off.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#283

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

my understanding is that a sufficiently large model will memorize the training data once enough representations are built up. Opus 4 scale seems to have been sufficient. cf NYT vs OAI.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#285

I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.

I always thought that due to the big batch size in SGD/Adam/Muon the model will not memorize a single conversation when trained on, but idk how true that is. The idea of AI companies pin-pointing users that do novel scientific research and then tracking their activity is the direction this points to. I hope that's not the case; that would be bad.

They're almost certainly pin-pointing high-quality conversations and giving them a special weighting. Seems stupid to not do that.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#286

Both things can be true: 1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation. 2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and…

I feel that we don’t praise Lean enough. AFAIU it’s what enables LLMs to brute force those problems

Re: More questions about whether researchers can trust OpenAI with unpublished math

#287
post #155

I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting: > Improve the model for everyone > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more. It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem". I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

I suspect the fundamental problem here is it's hard (if not impossible) to determine if someone who tried the winning approach deserves the credit for the discovery, because there's always the chance that they could've done something differently, or stopped before finishing, and thus never actually made the discovery. They might've even tried the approach just based on a whim, without really thinking it would work, and might've given up without a final insight. And fundings run out, people end up in hospitals, etc. What do you credit them with when the work isn't finished? For trying an approach that sounded promising? You can do that I guess, but is that what they want?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#288

https://x.com/markchen90/status/2097400166554993041?s=20 that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.

Not sure what you are seeing in that tweet that gives you the impression that the toggle does nothing.

Re: More questions about whether researchers can trust OpenAI with unpublished math

#289

Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings. My naive instincts would be that it seems unlikely that a single chat transcript wou…

Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?

Re: More questions about whether researchers can trust OpenAI with unpublished math

#290

I've been wondering whether AI really is improving rapidly at open problems or we're being fooled. - OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay - Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2] - But researchers will typically work on open problems. A researcher who is using Co…

It tells me that AI companies are just another mechanism to extract and extort value from the masses for the rich. Just another rich man’s trick Perhaps the last one before they destroy that world and try to hide away as people forget and history is rewritten again. I don’t think they’ll succeed this time.

AI providers are pretty much the end boss of rent seeking, that’s for sure
Post reply on HN