Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

31–40 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#31

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

> The user has no right to privacy. The same as how any internet service can be (and have been) compelled to produce private messages.

This is nonsense. I’ve personally been involved in these things, and fought to protect user privacy at all levels and never lost.

Re: Fighting the New York Times' invasion of user privacy

#32

Almost every comment (five) so far is against this: 'An incredibly cynical attempt at spin', 'How dare the New York Times demand access to our vault of everything-we-keep to figure out if we're a bunch of lying asses', etc. In direct contrast: I fully agree with OpenAI here. We can have a more nuanced opinion than 'piracy to train AI is bad therefore refusing to share chats is bad', which sounds absurd but is genuine…

I suspect that many of those comments are from the Philosopher's Chair (aka bathroom), and are not aspiring to be literal answers but are ways of saying "OpenAI Bad". But to your point there should be privacy preserving ways to comply, like user anonymization, tailored searches and so on. It sounds like the NYT is proposing a random sampling of user data. But couldn't they instead do a random sampling of their most widely read articles, for positive hits, rather than reviewing content on a case by case basis?

Re: Fighting the New York Times' invasion of user privacy

#33

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

>But conversations people thought they were having with OpenAI in private ...had never been private in the first place. not only is the data used for refining the models, OpenAI had also shariah policed plenty of people for generating erotica.

> OpenAI had also shariah policed plenty of people for generating erotica.

That framing is retorically brilliant if you think about it. I will use that more. Chat Sharia Law for Chat Control. Mass Sharia Surveillance from flock etc.

Re: Fighting the New York Times' invasion of user privacy

#34

Open AI deservedly getting a beating in this HN comments section but any comments about NYT overreach and what it means in general? And what if they for example find evidence of X other thing such as: 1. Something useful for a story, maybe they follow up in parallel. Know who to interview and what to ask? 2. A crime. 3. An ongoing crime. 4. Something else they can sue someone else for. 5. Top secret information

> 5. Top secret information

https://en.wikipedia.org/wiki/Pentagon_Papers

Re: Fighting the New York Times' invasion of user privacy

#35

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

>But conversations people thought they were having with OpenAI in private ...had never been private in the first place. not only is the data used for refining the models, OpenAI had also shariah policed plenty of people for generating erotica.

Yeah I don’t get why more people don’t understand this - why would you think your conversation was private when it wasnt actually private. Have you not been paying attention.

Re: Fighting the New York Times' invasion of user privacy

#36

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

The original lawsuit has lots of examples of ChatGPT (3.5? 4?) regurgitating article...snippets. They could get a few paragraphs with ~80-90% perfect replication. But certainly not full articles, with full accuracy.

This wasn't solid enough for a summary judgement, and it seems the labs have largely figured out how to stop the models from doing this. So it looks like NYT wants to comb all user chats rather than pay a team of people tens of thousands a day to try an coax articles out of ChatGPT-5.

Re: Fighting the New York Times' invasion of user privacy

#37

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

[flagged]

Re: Fighting the New York Times' invasion of user privacy

#38

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

> In copyright cases, typically you need to show some kind of harm.

NYT is suing for statutory copyright infringement. That means you only need to demonstrate that the copyright infringement, since the infringement alone is considered harm; the actual harm only matters if you're suing for actual damages.

This case really comes down to the very unsolved question of whether or not AI training and regurgitation is copyright infringement, and if so, if it's fair use. The actual ways the AI is being used is thus very relevant for the case, and totally within the bounds of discovery. Of course, OpenAI has also been engaging this lawsuit with unclean hands in the first place (see some of their earlier discovery dispute fuckery), and they're one of the companies with the strongest "the law doesn't apply to US because we're AI and big tech" swagger.

Re: Fighting the New York Times' invasion of user privacy

#39

I wouldn't want to make it out like I think OpenAI is the good guy here. I don't. But conversations people thought they were having with OpenAI in private are now going to be scoured by the New York Times' lawyers. I'm aware of the third party doctrine and that if you put something online it can never be actually private. But I think this also runs counter to people's expectations when they're using the product. In c…

>But conversations people thought they were having with OpenAI in private ...had never been private in the first place. not only is the data used for refining the models, OpenAI had also shariah policed plenty of people for generating erotica.

This is about private chats, which are not used for training and only stored for 30 days.

Also, you need to understand, that for huge corps like OpenAI, the lying on your ToS will do orders of magnitude more damage to your brand than what you would gain through training on <1% more user chats. So no, they are not lying when they say they don't train on private chats.

Re: Fighting the New York Times' invasion of user privacy

#40

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

> The user has no right to privacy. The same as how any internet service can be (and have been) compelled to produce private messages. This is nonsense. I’ve personally been involved in these things, and fought to protect user privacy at all levels and never lost.

You've successfully fought a subpoena on the basis of a third party's privacy? More than once? I'd love to hear more.
Post reply on HN