Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

191–200 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#191

Open AI deservedly getting a beating in this HN comments section but any comments about NYT overreach and what it means in general? And what if they for example find evidence of X other thing such as: 1. Something useful for a story, maybe they follow up in parallel. Know who to interview and what to ask? 2. A crime. 3. An ongoing crime. 4. Something else they can sue someone else for. 5. Top secret information

1-5: not a concern

It'll be the lawyers who need to go through the data, and given the scale of it, they won't be able to do anything more than trawl for the evidence they need and find specific examples to cite. They don't give a shit if you're asking chatgpt how to put a hit out on your ex, and they're not there to editorialize.

I wont pretend to guess* how they'll perform the discovery, but I highly doubt it will result in humans reading more than a handful of the records in total outside of the ones found via whatever method they automate the discovery process.

If there's top secret information in there, and it was somehow stumbled upon by one of these lawyers or a paralegal somewhere, I find it impossibly unlikely they'd be stupid enough to do anything other than run directly to whomever is the rightful possessor of said information and say "hey we found this in this place it shouldn't be" and then let them deal with it. Which is what we'd want them to do.

*Though if I had to speculate on how they'd do it, I do think the funniest way would be to feed the records back into chatgpt and ask it to point out all the times the records show evidence of infringement

Re: Fighting the New York Times' invasion of user privacy

#192

Earlier quoted context omitted.

>Discovery isn't binary yes/no, it involves competing proposals regarding methods and scope for satisfying information requests. Sometimes requests are egregious or excessive, sometimes they are reasonable and subject to excessively zealous pushback. There is a court order that OpenAI must produce these documents. OpenAI litigated this issue and lost. I'm not sure what point you are trying to make. The court decided…

Meanwhile back in reality, as of today that order is being challenged, and challenging the scope of an interlocutory order in discovery is a normal thing and part of a coherent legal position. So I don't know why you're pretending you don't understand what it means not to "roll over like a beaten dog" in response to overzealous discovery. >I don't think you read TFA. It was in TFA. If you don't like their number whic…

>Meanwhile back in reality, as of today that order is being challenged, and challenging the scope of an interlocutory order in discovery is a normal thing and part of a coherent legal position. So I don't know why you're pretending you don't understand what it means not to "roll over like a beaten dog" in response to overzealous discovery.

That's all OpenAI's argument and not coherent with regard to the facts. That's why they lost. I'd be willing to bet you they lose again. The rest of what you post is just verbatim their side, without any real analysis w/r/t to the facts (again) and I find your responses to this article to be a bit ridiculous in that regard. No less while you criticize others for pointing out OpenAI's incredibly self serving, smarmy, BS about privacy that they otherwise do not actually care about.

Again, this is not an issue of user privacy. OpenAI already represented to the court that they could anonymize the logs and that they did anonymize the logs (something you repeatedly fail to acknowledge while ranting that I didn't read the "article"). The issue is that OpenAI does not want to produce these logs because it will demonstrate that they are wrong. If you're gullible enough to believe otherwise, sure, but it certainly doesn't warrant the ridiculous attitude you bring to communicating with others here.

Re: Fighting the New York Times' invasion of user privacy

#193
post #136

Earlier quoted context omitted.

> What is the purpose of OpenAI storing millions of private conversations Its needed for the conversation history feature, a core feature of the ChatGPT product Its like saying "What is the purpose of Google Photos storing millions of private images"

This is true but why retain deleted conversations?

ChatGPT (the app) specifically says they keep deleted conversations for up to 30 days. That's probably why.

Re: Fighting the New York Times' invasion of user privacy

#195

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

You don't hate the media nearly enough. "Credible" my ass. They hired "experts" who used prompt engineering and thousands of repetitions to find highly unusual and specific methods of eliciting text from training data that matched their articles. OpenAI has taken measures to limit such methods and prevent arbitrary wholesale reproduction of copyrighted content since that time. That would have been the end of the situ…

I'm not a fan of NYT either, but this feels like you're stretching for your conclusion:

> They hired "experts" who used prompt engineering and thousands of repetitions to find highly unusual and specific methods of eliciting text from training data that matched their articles....would have been the end of the situation if NYT was engaging in good faith.

I mean, if I was performing a bunch of investigative work and my publication was considered the source of truth in a great deal of journalistic effort and publication of information, and somebody just stole my newspaper off the back of a delivery truck every day and started rewriting my articles, and then suddenly nobody read my paper anymore because they could just ask chatgpt for free, that's a loss for everyone, right?

Even if I disagree with how they editorialize, the Times still does a hell of a lot of journalism, and chatgpt can never, and will never be able to actually do journalism.

> they want to insert themselves as middlemen - pure rent seeking, second hander, sleazy lawyer behavior

I'd love to hear exactly what you mean by this.

Between what and what are they trying to insert themselves as middlemen, and why is chatgpt the victim in their attempts to do it?

What does 'rent seeking' mean in this context?

What does 'second hander' mean?

I'm guessing that 'sleazy lawyer' is added as an intensifier, but I'm curious if it means something more specific than that as well, I suppose.

> Copyright law....the rest of it

Yeah. IP rights and laws are fucked basically everywhere. I'm not smart enough to think of ways to fix it, though. If you've got some viable ideas, let's go fix it. Until then, the Times kinda need to work with what we've got. Otherwise, OpenAI is going to keep taking their lunch money, along with every other journalist's on the internet, until there's no lunch money to be had from anyone.

Re: Fighting the New York Times' invasion of user privacy

#196

Can this legal principle be used on Gmail too?

Of course this principle applies to Gmail too, if you’re willing to accept the absurdity. I could copy-paste copyrighted NYT snippets into emails and send them to everyone I know. Under the same logic, the NYT would be entitled to have access to everyone's Gmail account in order to verify who's sending what and get compensated if anyone is infringing their copyright. That’s not justice. That’s legal extortion. I get…

> That’s not justice. That’s legal extortion.

If you made it your business to publish a newsletter containing copied NYT articles, then wouldn't they have the right to go after you and discover your sent emails?

Re: Fighting the New York Times' invasion of user privacy

#197

Please correct me if I am wrong, but couldn't OpenAi just encrypt every conversation before saving them? With each query to the model the full conversation is fed into the model again, so I guess there is no technical need to store them unencrypted. Unless, of course, OpenAi wants to analyze the chats. The way I see it, the problem is that OpenAI employees can look at the chats and the fact that some NYT lawyer can l…

Encryption that you have the keys to won't save you from a court order

what about encryption only the users have the keys to? I'm assuming thats what parent meant

Re: Fighting the New York Times' invasion of user privacy

#198

If it's about* proving that people are getting around the paywall with OpenAI, won't it be much easier to prove this with a live reproduction in the court? * I am not too familiar with this matter and hence definitely am not rooting for one party or another. Asking this just out of technical curiosity.

No, because OpenAI can change their server at any time - and does, to patch over the cases where their copyright infringement is obvious.

Re: Fighting the New York Times' invasion of user privacy

#199
post #7

Cynicism aside, this seems like an attempt to prune back a potentially excessive legal discovery demand by appealing to public opinion. The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations. They claim they might find examples of you using ChatGPT to try to get around their paywall.

Yeah, I'm not sure why everyone feels the need to take a side here. Both of these organizations are ghoulish.

The NYT has problems with being a stooge of the military-industrial complex, but I really don't see them doing anything wrong in this case.

Re: Fighting the New York Times' invasion of user privacy

#200
post #186

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

>What is the purpose of OpenAI storing millions of private conversations Have you used ChatGPT? Your conversation history is on the left rail

"Have you used ChatGPT?"

No

Large number of upvotes on the quoted comment however. Maybe some of those voters are ChatGPT users

I do searching from the command line in text mode. The script I use keeps a "log" (a customised SERP) of all query strings and search result URLs. I also have these URLs stored in the logs from the forward proxy. These are compressed using RePair. I can search the compressed logs faster this way than with something like

    ztsd -dc log.zst|grep pattern
or

    rg -z pattern log.zst
Post reply on HN