Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

381–390 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#382

Earlier quoted context omitted.

Even if OpenAI is reproducing pieces of NYT articles, they still have a difficult argument because in no way is is a practical means of accessing paywalled NYT content, especially compared to alternatives. The entire value proposition of the NYT is news coverage, and probably 99.9% of their page views are from stories posted so recently that they aren't even in the training set of LLMs yet. If I want to reproduce a N…

OpenAI is not allowed to reproduce the NYT's articles, that's copyright infringement. It does not really matter if it is a practical thing or not, that would only go to damages, not liability.

What do you think it is you are liable for?

Re: Fighting the New York Times' invasion of user privacy

#383

Earlier quoted context omitted.

They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...

Damn, you'd think OpenAI would have made this argument! Maybe there's something you're missing if this didn't save the day for them.

No I wouldn't since this is discovery. Maybe there's something you're missing here.

Re: Fighting the New York Times' invasion of user privacy

#384

Earlier quoted context omitted.

We're not talking about collaborative tooling, just a record of what you've asked an AI assistant. If it doesn't sync right away, it's not the end of the world. I find that's true with most things. And the clients don't need to be running at the same time if you have a third device that's always on and receiving the changes from either (like a backup system). Eventually everything arrives. It's not as robust as what…

Chatgpt.com is essentially a CRUD app. What you're saying here amounts to saying that it could conceivably have been designed to work dramatically differently from all other CRUD apps. And obviously that's true, but why would it be? It's a website! You submit text, that you'll view or edit later, so the server stores it. How is that controversial to a HN audience? Also: > the clients don't need to be running at the s…

> An always-on device that stores data in order to sync it to clients is a server.

Yes. But it's my server. I burden myself to operate it so that persistence does not come at the cost of control.

I think we might be tilting at different windmills here.

Re: Fighting the New York Times' invasion of user privacy

#385
This is rich coming from the company that scraped the entire internet and tons of pirated books and scientific papers to train their models.

Maybe if you didn't scrape every single site on the internet they wouldn't have a basis for their case that you've stolen all of their articles through training your models on them. If anyone is to blame for this its openAI, not the NYT.

Play stupid games win stupid prizes.

Re: Fighting the New York Times' invasion of user privacy

#386
If you do anything in America that results in a stored record it's possible it will be released in discovery and a lawyer will read it. This happens all the time, and has happened for hundreds years.

It's not like the NYT will be published this shit in the news. Their lawyers and experts will have access to make a legal case, under a protective order. I'm not going to lose my law license because I'm doing doc review and you asked it something naughty and I think it's funny.

Courts and lawyers deal with this stuff all the time. What's very very weird to me is how upset OpenAI is about it.

They look like they are hiding something.

Re: Fighting the New York Times' invasion of user privacy

#387

Earlier quoted context omitted.

> As a direct and proximate result of Defendants’ infringing conduct alleged herein, The Times has sustained and will continue to sustain substantial, immediate, and irreparable injury for which there is no adequate remedy at law. Unless Defendants’ infringing conduct is enjoined by this Court, Defendants have demonstrated an intent to continue to infringe the copyrighted works. The Times therefore is entitled to per…

Very much appreciate the clarification and nuance here. I understand that legally they don't have to provide any of this detail, but I'm also somewhat astonished that there doesn't appear to be any evidence that they've been harmed in any way other than them claiming that they are.

It’s because 1/the damages aren’t clearly articulable and would be speculative at the time of filing, and 2/they don’t have to claim the specific nature of the injury at this point in the case.

Re: Fighting the New York Times' invasion of user privacy

#388
post #292

"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…

> The documents requested are not being made public by the plaintiffs In fact, as far as I understand it, they could not be made public by the plaintiffs even if they wanted to do so, or even if one of their employees decided to leak them. That's because the plaintiffs themselves never actually see the documents. They will only be seen by the plaintiff's lawyers and any experts hired by those lawyers to analyze them.

You are correct. I've operated under many protective orders that require me to redact portions of reports clients paid for because they were not authorized to see those specific parts due to the order.

Re: Fighting the New York Times' invasion of user privacy

#389

Earlier quoted context omitted.

I'm not commenting on the core point of your comment, only the "why retain for 30 days" question. Im an age of automated backups and failovers, deleting can be really hard. Part of the answer could simply be that syncing a delete across all the redundancies (while ensuring those redundancies are reliable when a disaster happens and they need to recover or maintain uptime) may take days to weeks. Also the 30 days coul…

The most likely explanation is whatever storage solution they’re using has a built in “recycle bin” functionality and deleted data stays the for 30 days before it’s actually deleted. I see this a lot in very large databases. The recycle bin functionality is built in to the data store product.

I'm doubtful that a data store product used at their scale can't be configured to not keep data for 30 days; for large clients that could be TB of deleted data or more. This would be neither cheap or easy to manage.

Re: Fighting the New York Times' invasion of user privacy

#390
post #331

It ridiculous for OpenAI to attempt to claim some moral high-ground here. They're a company that has demonstrated zero respect for the copyright or data privacy regulations of other organisations. I think they take users dignity and rights with a grain of salt. Their statements are all aspirational, "we're working toward de-identifying" etc. They've built one of the most powerful AIs ever seen and now they're claimin…

The “aspirational” language is what really stood out to me as well. “We’re building our privacy and security protections to match the responsibility” and “we are accelerating our security and privacy roadmap” and “our long term roadmap includes advanced security features designed to keep your data private, including client-side encryption” (what does this have to do with what OpenAI stores server-side?) and “we will build.” If OpenAI cared that much, then the privacy and security protections should be baked in rather than “tacked on.” Their statement makes me feel even less optimistic in their abilities to protect information.
Post reply on HN