Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

81–90 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#81
"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations."

As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims

Why has OpenAI collected and stored 20 million conversations (including "deleted chats")

What is the purpose of OpenAI storing millions of private conversations

By contrast the purpose of NYT's request is both clear and limited

The documents requested are not being made public by the plaintiffs. The documents will presumably be redacted to protect any confidential information before being produced to the plaintiffs, the documents can only be used by the plaintiffs for the purpose of the litigation against OpenAI and, unlike OpenAI who has collected and stored these conversations for as long as OpenAI desires, the plaintiffs are prohibited from retaining copies of the documents after the litigation is concluded

The privacy issue here has been created by OpenAI for their own commercial benefit

It is not even clear what this benefit, if any, will be as OpenAI continues to search for a "business model"

Wanton data collection

Re: Fighting the New York Times' invasion of user privacy

#82

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

>That's what NYTimes lawyers are after. They want the chat logs so they can do their own searches to find NYTimes text within the responses.

The trouble with this logic is NYT already made that argument and lost as applied to an original discovery scope of 1.4 billion records. The question now is about a lower scope and about the means of review, and proposed processes for anonymization.

They have a right to some form of discovery, but not to a blank check extrapolation that sidesteps legitimate privacy issues raised both in OpenAIs statement as well as throughout this thread.

Re: Fighting the New York Times' invasion of user privacy

#83

I fully believe that OpenAI is essentially stealing the work of others by training their models on it without permission. However, giving a corporation infamous for promoting authoritarianism full access to millions of private conversations is not the answer. OpenAI is right here. The NYT needs to prove their case another way.

> giving a corporation infamous for promoting authoritarianism

The NYT is certainly open to criticism along many fronts, but I don't have the slightest idea what you mean in claiming it promotes authoritarianism.

Re: Fighting the New York Times' invasion of user privacy

#84
This is BS. It’s like saying “We robbed a jewelry store and sold the jewelry. Now the police are poking around to see if anyone is wearing the jewelry we stole. Blasphemy! But don’t worry we will protect your privacy!”

Of course the Times wants more evidence that the content OpenAI allegedly stole is ending in things OpenAI is selling.

Re: Fighting the New York Times' invasion of user privacy

#85
>They claim they might find examples of you using ChatGPT to try to get around their paywall.

Is this a joke? We all know people do this. There is no "might" in it. They WILL find it.

OpenAI is trying to make it look like this is a breach of user's privacy, when the reality is that it's operating like a pirate website and if it were investigated that would become proven.

Re: Fighting the New York Times' invasion of user privacy

#86

Open AI deservedly getting a beating in this HN comments section but any comments about NYT overreach and what it means in general? And what if they for example find evidence of X other thing such as: 1. Something useful for a story, maybe they follow up in parallel. Know who to interview and what to ask? 2. A crime. 3. An ongoing crime. 4. Something else they can sue someone else for. 5. Top secret information

1. That sounds useful.

2. That sounds useful.

3. That sounds useful.

4. That sounds useful.

5. That sounds useful.

Are these supposed to be examples of things that shouldn't be found out about? This has to be the worst pro-privacy argument I've ever seen on the internet. "Privacy is good because they will find out about our crimes"

Re: Fighting the New York Times' invasion of user privacy

#87
post #84

This is BS. It’s like saying “We robbed a jewelry store and sold the jewelry. Now the police are poking around to see if anyone is wearing the jewelry we stole. Blasphemy! But don’t worry we will protect your privacy!” Of course the Times wants more evidence that the content OpenAI allegedly stole is ending in things OpenAI is selling.

It's more like a torrent tracker telling users that a newspaper wants to know what people are torrenting because they "claim" people are torrenting the newspaper, but investigating this would be an invasion of privacy of the users of the torrent tracker.

This isn't even a hyperbole. It's literally the same thing.

Re: Fighting the New York Times' invasion of user privacy

#88

I fully believe that OpenAI is essentially stealing the work of others by training their models on it without permission. However, giving a corporation infamous for promoting authoritarianism full access to millions of private conversations is not the answer. OpenAI is right here. The NYT needs to prove their case another way.

I'll bet you're right in some cases. I don't think that it is as pervasive as it has been made out to be though, but the argument requires some framing and current rules, regulation, and laws aren't tuned to make legal sense of this. (This is a little tangential, because the complaint seems to be about getting ChatGPT to reproduce content verbatim to a third party.)

There are two things I think about:

First, and generally, an AI ought to be able to ingest content like news articles because it's beneficial for users of AI. I would like to question an AI about current events.

Secondly, however, the legal mechanism by which it does that isn't clear. I think it would be helpful if these outlets would provide the information as long as the AI won't reproduce the content verbatim. If that does not happen, then another framing might liken the AI ingestion as an individual going to the library to read the paper. In that case, we don't require the individual to retroactively pay for the experience or unlearn what he may have learned while at the library.

Re: Fighting the New York Times' invasion of user privacy

#89

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

No it's not. It's literally a court order mandating them to collect this data.

- [1] https://arstechnica.com/tech-policy/2025/08/openai-offers-20...

Re: Fighting the New York Times' invasion of user privacy

#90

Earlier quoted context omitted.

Its clearly propaganda. "Your data belongs to you." I'm sure the ToS says otherwise, as OpenAI likely owns and utilizes this data. Yes, they say they are working on end-to-end encryption (whatever that means when they control one end), but that is just a proposal at this point. Also their framing of the NYT intent makes me strongly distrust anything they say. Sit down with a third party interviewer who asks challengi…

"Your data belongs to you" but we can take any of your data we can find and use it for free for ever, without crediting you, notifying you, or giving you any way of having it removed.

We can even download it illegally to train our models on it!
Post reply on HN