Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

221–230 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#221

Almost every comment (five) so far is against this: 'An incredibly cynical attempt at spin', 'How dare the New York Times demand access to our vault of everything-we-keep to figure out if we're a bunch of lying asses', etc. In direct contrast: I fully agree with OpenAI here. We can have a more nuanced opinion than 'piracy to train AI is bad therefore refusing to share chats is bad', which sounds absurd but is genuine…

OpenAI is the one who chose to store the information. Nobody twisted their arm to do so.

If you store data it can come up in discovery during lawsuits and criminal cases. Period.

E.g., storing illegal materials on Google Drive, Google WILL turn that over to the authorities if there’s a warrant or lawsuit that demands it in discovery.

E.g., my CEO writes an email telling the CFO that he doesn’t want to issue a safety recall because it’ll cost too much money. If I sue the company for injuring me through a product they know to be defective, that civil suit subpoena can ask for all emails discussing the matter and there’s no magical wall of privacy where the company can just say “no that’s private information.”

At the same time, I don’t get to trawl through the company’s emails and use some email the CEO flirting with their secretary as admissible evidence.

There are many ways the court is able to ensure privacy for the individuals. Sexual assault victims don’t have their evidence blasted across the the airwaves just because the court needs to examine that physical evidence.

The only way to avoid this is to not collect the data in the first place, which is where end to end encryption with user-controlled keys or simply not collecting information comes into play.

Re: Fighting the New York Times' invasion of user privacy

#222

WTF with all these comments. Regardless on OpenAI reputation and practices, I don't want NYT or anyone else to see my conversations, I completely agree to OpenAI here.

>I don't want NYT or anyone else to see my conversations Except for OpenAI, apparently.

Obviously for OpenAI, because they providing the service. Same for my doctor, my pharmacist and my lawyer. What was your point?

Re: Fighting the New York Times' invasion of user privacy

#223

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

> NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" Credible to whom? In their supposed "investigation", they sent a whole page of text and complex pre-prompting and still failed to get the exact content back word for word. Something users would never do anyways. And that's probably the best the…

Agreed, they could carefully coerce the model to more or less output some of their articles, but the premise that users were routinely doing this to bypass the paywall is silly.

Re: Fighting the New York Times' invasion of user privacy

#224

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

NB. There is no order to "collect". The order is to preserve what is already being collected and stored in the ordinary course of business https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6... https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6...

Why does OpenAI collect and retain for 30 days^1 chats that the user wants to be deleted

It was doing this prior to being sued by the NYT and many others

OpenAI was collecting chats even when the user asked for deletion, i.e., the user did not want them saved

That's why a lawsuit could require OpenAi to issue a hold order, retain these chats for longer and produce them to another party in discovery

If OpenAI was not collecting these chats in the ordinary course of its business before being sued by the NYT and many others, then there would be no "deleted chats" for OpenAI to be compelled by court order to retain and produce to the plaintiffs

1. Or whatever period OpenAI decides on. It could change at any time for any reason. However OpenAI cannot change their retention policy to some shortened period after being sued. Google tried this a few years ago. It began destroying chats between employees after Google was on notice it was going to be sued by the US government and state AGs

Re: Fighting the New York Times' invasion of user privacy

#226
post #196

Earlier quoted context omitted.

Of course this principle applies to Gmail too, if you’re willing to accept the absurdity. I could copy-paste copyrighted NYT snippets into emails and send them to everyone I know. Under the same logic, the NYT would be entitled to have access to everyone's Gmail account in order to verify who's sending what and get compensated if anyone is infringing their copyright. That’s not justice. That’s legal extortion. I get…

> That’s not justice. That’s legal extortion. If you made it your business to publish a newsletter containing copied NYT articles, then wouldn't they have the right to go after you and discover your sent emails?

Exactly, they wouldn't even need all of the emails in gmail for that example, just the ones from a specific account.

The real equivalent here would be if gmail itself was injecting NYT articles into your emails. I'm assuming in that scenario most people would see it as straightforward that gmail was infringing NYT content.

Re: Fighting the New York Times' invasion of user privacy

#227

Earlier quoted context omitted.

If OpenAI truly didn't keep conversation records for any length of time, they would not be subject to this kind of order. Lots of stateless services get these and are able to defeat them because they never store the user's data. The fact that they store them at all means that they are in scope for a preservation order. It also means that they are in scope for all manner of usage by OpenAI themselves even if a user re…

It seems as if the court has forced OpenAI into collecting logs that they weren't otherwise collecting, or that they were deleting at user request. So in this case not keeping logs as ordered by the court would be contempt of court.

Respectfully, it doesn’t matter the way it “seems,” it matters what is. They were collecting these logs, and as soon as they got the preservation order, they disabled deletion functionality and notified their customers of that.

There is a separate higher-tier private API customers can pay for that never had logging enabled, and the court did not force the company to add it.

Re: Fighting the New York Times' invasion of user privacy

#228

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

> What is the purpose of OpenAI storing millions of private conversations Your previous ChatGPT conversations show up right in the ChatGPT interface. They have to store the private conversations to enable users to bring them up in the interface. This isn't a secretive, hidden data collection. It's a clear and obvious feature right in the product. They're fighting for the ability to not retain secret records of past c…

They could have been stored at the client, and encrypted before optionally synced back to OpenAI servers in a way that the stored chats can only be read back by the user. Signal illustrates how this is possible.

OpenAI made a choice in how the feature was and is implemented.

Re: Fighting the New York Times' invasion of user privacy

#229

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

> This case is unusual because the New York Times can't point to any harm It helps to read the complaint. If that was the case, the case would have been subject to a Rule 12(b)(6) (failure to state a claim for which relief can be granted) challenge and closed. Complaint: https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... See pages 60ff.

My observation is that section does not articulate any harm. It _claims_ harm, but doesn't actually explain what the harm is. Reduced profits? Lower readership? All they say is "OpenAI violated our copyrights, and we deserve money."

> 167. As a direct and proximate result of Defendants’ infringing conduct alleged herein, The Times has sustained and will continue to sustain substantial, immediate, and irreparable injury for which there is no adequate remedy at law. Unless Defendants’ infringing conduct is enjoined by this Court, Defendants have demonstrated an intent to continue to infringe the copyrighted works. The Times therefore is entitled to permanent injunctive relief restraining and enjoining Defendants’ ongoing infringing conduct. > 168. The Times is further entitled to recover statutory damages, actual damages, restitution of profits, attorneys’ fees, and other remedies provided by law.

They're simply claiming harm, nothing more. I want to see injuries, scars, and blood if there's harm. As far as I can tell, the NYT was on the ropes long before AI came along. If they could actually articulate any harm, they wouldn't need to read through everyone's chats.

Re: Fighting the New York Times' invasion of user privacy

#230

Earlier quoted context omitted.

Exactly. And the OpenAI corporates speak acting like they give a shit about our best interests. Give me a break, Sam Altman. How stupid do you think everyone is? They have proven that they are the most untrustworthy company on the planet And this isn't AI fear speaking. This is me raging at Sam Altman for spreading so much fear, uncertainty, and doubt just to get investments. The rest of us have to suffer for the las…

To me, no company has the customers’ best interests in mind. This whole thing is akin to when Apple was refusing to unlock phones for the FBI. Of course, Apple profits by having people think that they take privacy seriously, and they demonstrate it by protecting users’ privacy. Same thing here; OpenAI needs chats to have some expectation of privacy, especially because a large use case of AI is personal advice on thin…

> To me, no company has the customers’ best interests in mind.

Lavabit opted to stop operating rather than give the FBI access to client emails.

https://archive.ph/20200915083857/https://www.nytimes.com/20...

Post reply on HN