Live data from Hacker News

Tell HN: OpenAI keeps re-enabling the 'allow training' setting

news.ycombinator.com

161–170 of 180 posts

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#161
Dude, they probably just generate a piece of synthetic data from all of our inputs in some ambiguous way. Legally protects them, but your input is absolutely their input. You can bet your ass on that because their original input was all the shit humans ever wrote, why would your shit be different to them? Thieves are thieves. They absolutely train on your data whether your checked that thing or not.

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#162
post #119

Earlier quoted context omitted.

There is nothing even close to a proof. A lot of accusations, a lot of people ready with pitchforks and torches (sadly, also here on HN), but not a lot of facts. Did the researches opt out from data sharing on subsidised subs? Did anyone prove that their methods enabled OpenAI models to produce the solution? For a discussion about science, there is almost no scientifical method applied to proving anyone stole anythin…

On one side, yes we don't have hard evidence that intentional plagiarism is exactly what happened. On the other side, the lack of evidence is pretty damning. Only OpenAI can try to prove that they came by these results legitimately, and the case they're making is quite weak. They could make public metadata about what their model was trained on and whether it did train on the conversations in question; they have not.…

That it was a dick move, I think there is no doubt about that. OpenAI wanted to scoop Anthropic, and the two guys working on the problem got caught in the crossfire.

Both OpenAI and the researchers know if the sessions in questions were subject to data sharing. Why neither the scientists nor OpenAI is clear about that is weird - it would seem at least one party has the incentive to report that. But even if their sessions were in training data sets its hard to tell whether it influenced the outcome. Those models are big, but are they big enough to preserve subtle, niche techniques enough to draw from them while solving a related problem? Probably nobody knows.

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#164
post #48

Lol, "Outrageous that the company that chose to ignore copyright holder claims, chose to ignore my checkbox of intent despite the implied pinky promise".

Ignoring the checkbox is an utterly offensive move, but repeatedly manipulating it contrary to stated consumer intent is a whole other level. I didn't know we were supposed to take 'frontier' literally in every sense of the word. I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!

Pirating copyrighted works is absolutely illegal in most jurisdictions, but pirating every single copyrighted work in the world is somehow exempt from law.

We can't apply plebeian laws or ethics to our benevolent overlords, they are above our worldly worries.

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#165
post #48

Earlier quoted context omitted.

Ignoring the checkbox is an utterly offensive move, but repeatedly manipulating it contrary to stated consumer intent is a whole other level. I didn't know we were supposed to take 'frontier' literally in every sense of the word. I have now witnessed this myself after not believing this at first. Of course, screenshots etc. will hardly prove anything. This needs a proper third-party audit!

Oh they aren't a frontier on this. Undoing user configuration is well tested in Windows land.

"Remember my choice", right?:)

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#166

Earlier quoted context omitted.

Is there a description of what "Improve the model for everyone" actually means or is it just a straight up *Dark* pattern?

It's pretty clear once you click on it: > Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more

So not in the basement behind a sign saying "Beware of the Leopard" but could also just be "Allow your content to be used to train our models: ".

Could be worse I guess.

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#169
post #160
post #35

Earlier quoted context omitted.

Out of curiosity, how one is supposed to "document it properly"?

You can obtain a cryptographic proof by recording the tls exchange, including the keys You need to use a tls intercepting proxy for that. I couldn't find any ready-made tool unfortunately, there's tlsnotary.org but it seems far from simple.

So basically I record what the browser sends to the server when I click the toggle and the server response?

I wonder how I can attach a timestamp that cannot be faked.

Re: Tell HN: OpenAI keeps re-enabling the 'allow training' setting

#170
post #147

Earlier quoted context omitted.

Reverting... to on? I've seen some stuff in Apple-land where I'm pretty sure it went back to the default option after an update but I've never seen anything get granted more permissions.

No, like specific apps that have been granted full disk access etc lose their permissions. It’s hell for IT, constantly breaks AV. I would not characterize that as reverting to a default. More often it’s a breaking change for how the security model works that ends up requiring reauthorization

Ah, sure, it's annoying but it's never giving new permissions you didn't give it before. And yes "revert to default" is probably not what's happening, but reverting to forcing a specific permission grant when the existing one was superceeded or replaced.
Post reply on HN