Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

371–380 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#371

Earlier quoted context omitted.

To me, no company has the customers’ best interests in mind. This whole thing is akin to when Apple was refusing to unlock phones for the FBI. Of course, Apple profits by having people think that they take privacy seriously, and they demonstrate it by protecting users’ privacy. Same thing here; OpenAI needs chats to have some expectation of privacy, especially because a large use case of AI is personal advice on thin…

Both OpenAI and NYT are bad. I don't know about NYT's privacy policy, because that's not really the industry they're in, but they did admit to fabricating a story that led to a now 2-year-long war, so.

> Both OpenAI and NYT are bad.

Both -1 and -1,000,000 are negative numbers.

We need to be careful and mindful of our framing. Saying "X is bad" is a drastic oversimplification and not necessarily useful. Pointing at any one company and saying "bad" doesn't move the needle much in terms of figuring out how to steer us towards better outcomes. For that, we have to identify incentives and understand motivations.

Re: Fighting the New York Times' invasion of user privacy

#372

Earlier quoted context omitted.

I got one sentence in and thought to myself, "This is about discovery, isn't it?" And lo, complaints about plaintiffs started before I even had to scroll. If this company hadn't willy-nilly done everything they could to vacuum up the world's data, wherever it may be, however it may have been protected, then maybe they wouldn't be in this predicament.

How do you feel about Google vacuuming up the world's data when they created a search engine? I feel like everybody just ignores this because Google was ostensibly sending traffic to the resulting site. The actual infringement of scraping should be identical between OpenAI and Google. Why is nobody complaining about Google scraping their sites? Is it only because they're getting paid off to not complain? Everybody ac…

At the time Google created a search engine, they were not showing the data themselves, they were pointing to where those are. When they started to actually print articles themselves, they got sued. Showing where the thing is and showing content of the thing are two different actions.

So, when google did the same thing, there were complains.

> Why is nobody complaining about Google scraping their sites?

And second, search engines were actually pretty gentle with their sites scrapping. They needed the sites to work, so they respected robots.txt and made sure they wont accidentally DDoS sites by too many requests. AI companies just DDoS sites, do not respect robots.txt and if you block them, they will use another from their infinite amount of IPs.

Otherwise said, even back then, Google was kind trying to be ok non evil citizen. They became sociopathic only much later and even now kind of try to hide it. OpenAI and the rest of AI companies are openly sociopathic and proud of damage they cause.

Re: Fighting the New York Times' invasion of user privacy

#373

Earlier quoted context omitted.

A rope isn’t going to tell you to make sure you don’t leave it out on your bed so your loved ones can’t stop you from carrying out the suicide it helped talk you in to.

This is a good observation! The LLM can tell you to kill yourself. The rope can actually actually help you do it.

Ok

Re: Fighting the New York Times' invasion of user privacy

#374

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

>That's what NYTimes lawyers are after. They want the chat logs so they can do their own searches to find NYTimes text within the responses. The trouble with this logic is NYT already made that argument and lost as applied to an original discovery scope of 1.4 billion records. The question now is about a lower scope and about the means of review, and proposed processes for anonymization. They have a right to some for…

Again, as I pointed out to you numerous times in this thread. OpenAI already represented to the court that the data was anonymized and that they can anonymize it, so you are significantly departing from the actual facts in your discussion here. There are no genuine privacy issues left here. The data is anonymous and it is under a protective order so it must be maintained confidentially.

Re: Fighting the New York Times' invasion of user privacy

#375
post #307

Earlier quoted context omitted.

I get that you're mad, and rightly should be for an invasion of your privacy, but the NYT would be foolish to use any of your data for anything other than this lawsuit, and to not delete it afterwards, as per their request. They can't use this data against any individual, even if they explicitly asked, "How do I hack the NYT?" The only potential issue is them finding something juicy in someone's chat, that they could…

> The only potential issue is them finding something juicy in someone's chat, that they could publish as a story; and then claiming they found out about this juicy story through other means, (such as a confidential informant) Which is concerning since this is a news organization that's getting the data. Let's say they do find some juicy detail and use it, then what? Nothing. It's not like you can ever fix a privacy v…

>Let's say they do find some juicy detail and use it, then what? Nothing. It's not like you can ever fix a privacy violation. Nobody involved would get a serious punishment, like prison time, either.

There are no privacy violations. OpenAI already told the court they anonymized it. What they say in court and what they say in the blog is different and so many people here are (unfortunately) falling for it!

Re: Fighting the New York Times' invasion of user privacy

#376

Earlier quoted context omitted.

> The user has no right to privacy The correct term for this is prima facie right. You do have a right to privacy (arguably) but it is outweighed by the interest of enforcing the rights of others under copyright law. Similarly, liberty is a prima facie right; you can be arrested for committing a crime.

Seems to me my right to privacy is far more important than their right to copyright enforcement.

Have you read OpenAI's terms of service? Which part is being violated by producing anonymized logs in response to discovery? OpenAI's ToS state that they will produce your data in response to discovery. What's not clicking for you?

Re: Fighting the New York Times' invasion of user privacy

#377

Earlier quoted context omitted.

I get the feeling, but that's not what this is. NYTimes has produced credible evidence that OpenAI is simply stealing and republishing their content. The question they have to answer is "to what extent has this happened?" That's a question they fundamentally cannot answer without these chat logs. That's what discovery, especially in a copyright case, is about. Think about it this way. Let's say this were a book store…

Even if OpenAI is reproducing pieces of NYT articles, they still have a difficult argument because in no way is is a practical means of accessing paywalled NYT content, especially compared to alternatives. The entire value proposition of the NYT is news coverage, and probably 99.9% of their page views are from stories posted so recently that they aren't even in the training set of LLMs yet. If I want to reproduce a N…

OpenAI is not allowed to reproduce the NYT's articles, that's copyright infringement. It does not really matter if it is a practical thing or not, that would only go to damages, not liability.

Re: Fighting the New York Times' invasion of user privacy

#378

Can this legal principle be used on Gmail too?

Of course this principle applies to Gmail too, if you’re willing to accept the absurdity. I could copy-paste copyrighted NYT snippets into emails and send them to everyone I know. Under the same logic, the NYT would be entitled to have access to everyone's Gmail account in order to verify who's sending what and get compensated if anyone is infringing their copyright. That’s not justice. That’s legal extortion. I get…

If you make a business out of that, then yes, it is copyright infringement and thus you can be sued. Are we supposed to be outraged over someone making a business out of newspaper articles they did not wrote being potentially sued?

Your example is not nearly an example of copyright troll or overreach.

Re: Fighting the New York Times' invasion of user privacy

#379

If OpenAI hadn't used data from the NYT without permission in the first place this wouldn't have happened. That is the root cause of all this. I'm glad the NYT is fighting them. They've infringed the rights of almost every news outlet but someone has to bring this case.

They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...

Damn, you'd think OpenAI would have made this argument! Maybe there's something you're missing if this didn't save the day for them.

Re: Fighting the New York Times' invasion of user privacy

#380

Earlier quoted context omitted.

This article says nothing of the sort. The court order is to preserve existing logs they already have, not to disable logging, and hand all the logs over the plaintiffs. OpenAI's objections are mainly that 1/there are too many logs (so they're proposing a sample instead) and that 2/there's identifying data in the logs and so they are being "forced" to anonymize the logs at their expense (even though it's what they wa…

This response is misleading. Almost all computer services keep logs for a short period of time, so the court order to retain existing information is quite a bit more powerful than a layman would think. Because a huge amount of data is retained for a short period of time and then rapidly deleted in most web services I've worked on for the past 30 years. This is true in services like Datadog, New Relic, and logging ser…

There is an important distinction that relates to a court’s ability to order a defendant to perform work to facilitate discovery. A court can order preservation of records, but they generally cannot order a defendant to create new ones. I was responding to your use of the word “collect,” which implies significantly more effort than merely not destroying logs (i.e. logging new information that they weren’t already).

It’s not misdirection or misleading; it lies in an understanding of the law. There’s plenty of case law out there on the subject if you’re interested.

Post reply on HN