Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

331–340 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#331
It ridiculous for OpenAI to attempt to claim some moral high-ground here. They're a company that has demonstrated zero respect for the copyright or data privacy regulations of other organisations. I think they take users dignity and rights with a grain of salt.

Their statements are all aspirational, "we're working toward de-identifying" etc. They've built one of the most powerful AIs ever seen and now they're claiming it's difficult to delete, de-identify / anonymize. Maybe they should ask their AI to do it :-)

It's impossible to take this company seriously. They're nothing but a carny barker stealing everything of value that they can lay their (creepy) hands on.

Re: Fighting the New York Times' invasion of user privacy

#332
post #186

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

>What is the purpose of OpenAI storing millions of private conversations Have you used ChatGPT? Your conversation history is on the left rail

As requested

http://news.ycombinator.com/item?id=45708542

http://news.ycombinator.com/item?id=45708573

Re: Fighting the New York Times' invasion of user privacy

#333
post #136

Earlier quoted context omitted.

> What is the purpose of OpenAI storing millions of private conversations Its needed for the conversation history feature, a core feature of the ChatGPT product Its like saying "What is the purpose of Google Photos storing millions of private images"

This is true but why retain deleted conversations?

Because the New York Times sued them and made them.

https://openai.com/index/response-to-nyt-data-demands/

Re: Fighting the New York Times' invasion of user privacy

#334
post #228

Earlier quoted context omitted.

> What is the purpose of OpenAI storing millions of private conversations Your previous ChatGPT conversations show up right in the ChatGPT interface. They have to store the private conversations to enable users to bring them up in the interface. This isn't a secretive, hidden data collection. It's a clear and obvious feature right in the product. They're fighting for the ability to not retain secret records of past c…

They could have been stored at the client, and encrypted before optionally synced back to OpenAI servers in a way that the stored chats can only be read back by the user. Signal illustrates how this is possible. OpenAI made a choice in how the feature was and is implemented.

[deleted]

Re: Fighting the New York Times' invasion of user privacy

#335

Earlier quoted context omitted.

ChatGPT (the app) specifically says they keep deleted conversations for up to 30 days. That's probably why.

yeah but the link states "The 20 million user conversations were randomly sampled from Dec. 2022 to Nov. 2024" so this makes no sense. 2024 was much longer than 30 days ago

Because the court ordered them to retain the records longer than they normally would.

Re: Fighting the New York Times' invasion of user privacy

#336

Earlier quoted context omitted.

No it's not. It's literally a court order mandating them to collect this data. - [1] https://arstechnica.com/tech-policy/2025/08/openai-offers-20...

This article says nothing of the sort. The court order is to preserve existing logs they already have, not to disable logging, and hand all the logs over the plaintiffs. OpenAI's objections are mainly that 1/there are too many logs (so they're proposing a sample instead) and that 2/there's identifying data in the logs and so they are being "forced" to anonymize the logs at their expense (even though it's what they wa…

This response is misleading. Almost all computer services keep logs for a short period of time, so the court order to retain existing information is quite a bit more powerful than a layman would think. Because a huge amount of data is retained for a short period of time and then rapidly deleted in most web services I've worked on for the past 30 years.

This is true in services like Datadog, New Relic, and logging services like Splunk. But even privacy-focused services like Mullvad keep logs for 24 hours to monitor for abuse. So this concept that retaining logs is significantly weaker than not ordering the collection is really a bit of misdirection. I'm not sure whether it's intentional, but it's definitely misleading.

Re: Fighting the New York Times' invasion of user privacy

#337

Earlier quoted context omitted.

NB. There is no order to "collect". The order is to preserve what is already being collected and stored in the ordinary course of business https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6... https://ia801404.us.archive.org/31/items/gov.uscourts.nysd.6...

Why does OpenAI collect and retain for 30 days^1 chats that the user wants to be deleted It was doing this prior to being sued by the NYT and many others OpenAI was collecting chats even when the user asked for deletion, i.e., the user did not want them saved That's why a lawsuit could require OpenAi to issue a hold order, retain these chats for longer and produce them to another party in discovery If OpenAI was not…

> Why does OpenAI collect and retain for 30 days^1 chats that the user wants to be deleted

When working on an e-commerce gig we would get "delete my data" requests from customers, which we're legally obliged to comply with. A script would delete everything we could from the DB immediately. Since we had 30 day backups, their data would only be deleted from the backups on day 31. I think this was acceptable to the GDPR consultant.

Going in to the backups to delete their data there in insane.

Re: Fighting the New York Times' invasion of user privacy

#338

Can this legal principle be used on Gmail too?

Gmail is an Electronic Communication Service as defined in 18 U.S.C § 2510, meaning its contents are protected under the Stored Communications Act (18 U.S.C. Chapter 121 §§ 2701–2713).

Communications with an AI system do not involve a human so are not protected by ECPA or the SCA and get less protection. This is controversial and some people have called on ECPA/SCA to be extended to cover AI services. That means a warrant would be necessary to get your OpenAI history, not just a subpoena.

Re: Fighting the New York Times' invasion of user privacy

#339

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

> The user has no right to privacy The correct term for this is prima facie right. You do have a right to privacy (arguably) but it is outweighed by the interest of enforcing the rights of others under copyright law. Similarly, liberty is a prima facie right; you can be arrested for committing a crime.

Seems to me my right to privacy is far more important than their right to copyright enforcement.

Re: Fighting the New York Times' invasion of user privacy

#340

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

It's not credible. Using AI to regurgitate news articles is not a good use of the tool, and it is not credible that any statistically significant portion of their user base is using the tool for that.
Post reply on HN