Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

231–240 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#231

I've noticed a pattern of companies writing their customers open letters asking them to do their contract negotiations for them. First it was ESPN vs. YouTube (not watching MNF this week was the best 3 hours I've ever saved, sorry advertisers). Now it's OpenAI vs. The New York Times. Little do they know that I care very little for either party and enjoy seeing both of them squirm. You went to business school, not me.…

Two orgs helmed by supporters of authoritarianism (to put it nicely): let them fight.

Yup, well said. Companies want me to have an emotional investment in them. I don't.

Re: Fighting the New York Times' invasion of user privacy

#232

Earlier quoted context omitted.

> Also, you need to understand, that for huge corps like OpenAI, the lying on your ToS will do orders of magnitude more damage to your brand than what you would gain Is this true? I can’t recall anything like this (look at Ashley Madison which is alive and well)

It's not national news when a company is found to be doing what they say they are doing.

> It's not national news when a company is found to be doing what they say they are doing.

You said there would be ‘orders of magnitude’ of brand damage. What is the proof?

Re: Fighting the New York Times' invasion of user privacy

#233

I've noticed a pattern of companies writing their customers open letters asking them to do their contract negotiations for them. First it was ESPN vs. YouTube (not watching MNF this week was the best 3 hours I've ever saved, sorry advertisers). Now it's OpenAI vs. The New York Times. Little do they know that I care very little for either party and enjoy seeing both of them squirm. You went to business school, not me.…

> Finally, both parties should find a neutral third party. That's next to impossible. And if that party fails to be neutral you've just generated a new lawsuit entangled with this one. The current procedure is each side gets their own expert. The two expert can duke it out and the crucible of the courtroom decides who was more credible.

That's fair. I understand why OpenAI wouldn't want to give anyone transcripts (as a user, I frankly wouldn't even want OpenAI to keep my transcripts), and I understand why the NYT doesn't want to give OpenAI all their articles.

Maybe the NYT needs to bloom-filter-ify their articles in 10 word chunks (or something, I don't know enough about linguistics to tell you what's unique enough for copyright infringement or to prove "copying"), have OpenAI search transcripts, and turn over the matches. That limits the scope of the search dramatically, but is still invasive.

Re: Fighting the New York Times' invasion of user privacy

#234
post #228

Earlier quoted context omitted.

> What is the purpose of OpenAI storing millions of private conversations Your previous ChatGPT conversations show up right in the ChatGPT interface. They have to store the private conversations to enable users to bring them up in the interface. This isn't a secretive, hidden data collection. It's a clear and obvious feature right in the product. They're fighting for the ability to not retain secret records of past c…

They could have been stored at the client, and encrypted before optionally synced back to OpenAI servers in a way that the stored chats can only be read back by the user. Signal illustrates how this is possible. OpenAI made a choice in how the feature was and is implemented.

> Our long-term roadmap includes advanced security features designed to keep your data private, including client-side encryption for your messages with ChatGPT. We believe these features will help keep your private conversations private and inaccessible to anyone else, even OpenAI.

Re: Fighting the New York Times' invasion of user privacy

#235
post #2

> Trust, security, and privacy guide every product and decision we make. -- openai

- any corporation remember a corporation generally is an object owned by some people. Do you trust "unspecified future group of people" with your privacy? You can't. Best we can do is understand the information architecture and act accordingly.

> - any corporation

I don’t recall seeing many food, furniture, plant, or generally anything not related to tech talking about trust, security, and privacy as guiding principles.

Re: Fighting the New York Times' invasion of user privacy

#236

Earlier quoted context omitted.

NB. There is no order to "collect". The order is to preserve what is already being collected and stored in the ordinary course of business https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6... https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6...

Why does OpenAI collect and retain for 30 days^1 chats that the user wants to be deleted It was doing this prior to being sued by the NYT and many others OpenAI was collecting chats even when the user asked for deletion, i.e., the user did not want them saved That's why a lawsuit could require OpenAi to issue a hold order, retain these chats for longer and produce them to another party in discovery If OpenAI was not…

Maybe an append only data store where actual hard deletes only happen as an async batch job? Still 30 days seems really long for this.

Re: Fighting the New York Times' invasion of user privacy

#237

If OpenAI hadn't used data from the NYT without permission in the first place this wouldn't have happened. That is the root cause of all this. I'm glad the NYT is fighting them. They've infringed the rights of almost every news outlet but someone has to bring this case.

They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...

Re: Fighting the New York Times' invasion of user privacy

#238

Earlier quoted context omitted.

NB. There is no order to "collect". The order is to preserve what is already being collected and stored in the ordinary course of business https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6... https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6...

Why does OpenAI collect and retain for 30 days^1 chats that the user wants to be deleted It was doing this prior to being sued by the NYT and many others OpenAI was collecting chats even when the user asked for deletion, i.e., the user did not want them saved That's why a lawsuit could require OpenAi to issue a hold order, retain these chats for longer and produce them to another party in discovery If OpenAI was not…

I'm not commenting on the core point of your comment, only the "why retain for 30 days" question.

Im an age of automated backups and failovers, deleting can be really hard. Part of the answer could simply be that syncing a delete across all the redundancies (while ensuring those redundancies are reliable when a disaster happens and they need to recover or maintain uptime) may take days to weeks. Also the 30 days could be the limit, as oppose to the average or median time it takes.

Re: Fighting the New York Times' invasion of user privacy

#239

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

Even if OpenAI is reproducing pieces of NYT articles, they still have a difficult argument because in no way is is a practical means of accessing paywalled NYT content, especially compared to alternatives. The entire value proposition of the NYT is news coverage, and probably 99.9% of their page views are from stories posted so recently that they aren't even in the training set of LLMs yet. If I want to reproduce a NYT story from LLM it's a prompt engineering mess, and I can only get old ones. On the other hand I can read any NYT story from today by archiving it: https://archive.is/5iVIE. So why is the NYT suing OpenAI and not the Internet Archive?

Re: Fighting the New York Times' invasion of user privacy

#240

Earlier quoted context omitted.

You don't hate the media nearly enough. "Credible" my ass. They hired "experts" who used prompt engineering and thousands of repetitions to find highly unusual and specific methods of eliciting text from training data that matched their articles. OpenAI has taken measures to limit such methods and prevent arbitrary wholesale reproduction of copyrighted content since that time. That would have been the end of the situ…

I'm not a fan of NYT either, but this feels like you're stretching for your conclusion: > They hired "experts" who used prompt engineering and thousands of repetitions to find highly unusual and specific methods of eliciting text from training data that matched their articles....would have been the end of the situation if NYT was engaging in good faith. I mean, if I was performing a bunch of investigative work and my…

> my publication was considered the source of truth

Their publication is not considered the source of truth, at least not by anyone with a brain.

Post reply on HN