Fighting the New York Times' invasion of user privacy
381–390 of 441 posts
Re: Fighting the New York Times' invasion of user privacy
#382Earlier quoted context omitted.
Even if OpenAI is reproducing pieces of NYT articles, they still have a difficult argument because in no way is is a practical means of accessing paywalled NYT content, especially compared to alternatives. The entire value proposition of the NYT is news coverage, and probably 99.9% of their page views are from stories posted so recently that they aren't even in the training set of LLMs yet. If I want to reproduce a N…
OpenAI is not allowed to reproduce the NYT's articles, that's copyright infringement. It does not really matter if it is a practical thing or not, that would only go to damages, not liability.
Re: Fighting the New York Times' invasion of user privacy
#383Earlier quoted context omitted.
They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...
Damn, you'd think OpenAI would have made this argument! Maybe there's something you're missing if this didn't save the day for them.
Re: Fighting the New York Times' invasion of user privacy
#384Earlier quoted context omitted.
We're not talking about collaborative tooling, just a record of what you've asked an AI assistant. If it doesn't sync right away, it's not the end of the world. I find that's true with most things. And the clients don't need to be running at the same time if you have a third device that's always on and receiving the changes from either (like a backup system). Eventually everything arrives. It's not as robust as what…
Chatgpt.com is essentially a CRUD app. What you're saying here amounts to saying that it could conceivably have been designed to work dramatically differently from all other CRUD apps. And obviously that's true, but why would it be? It's a website! You submit text, that you'll view or edit later, so the server stores it. How is that controversial to a HN audience? Also: > the clients don't need to be running at the s…
Yes. But it's my server. I burden myself to operate it so that persistence does not come at the cost of control.
I think we might be tilting at different windmills here.
Re: Fighting the New York Times' invasion of user privacy
#385Maybe if you didn't scrape every single site on the internet they wouldn't have a basis for their case that you've stolen all of their articles through training your models on them. If anyone is to blame for this its openAI, not the NYT.
Play stupid games win stupid prizes.
Re: Fighting the New York Times' invasion of user privacy
#386It's not like the NYT will be published this shit in the news. Their lawyers and experts will have access to make a legal case, under a protective order. I'm not going to lose my law license because I'm doing doc review and you asked it something naughty and I think it's funny.
Courts and lawyers deal with this stuff all the time. What's very very weird to me is how upset OpenAI is about it.
They look like they are hiding something.
Re: Fighting the New York Times' invasion of user privacy
#387Earlier quoted context omitted.
> As a direct and proximate result of Defendants’ infringing conduct alleged herein, The Times has sustained and will continue to sustain substantial, immediate, and irreparable injury for which there is no adequate remedy at law. Unless Defendants’ infringing conduct is enjoined by this Court, Defendants have demonstrated an intent to continue to infringe the copyrighted works. The Times therefore is entitled to per…
Very much appreciate the clarification and nuance here. I understand that legally they don't have to provide any of this detail, but I'm also somewhat astonished that there doesn't appear to be any evidence that they've been harmed in any way other than them claiming that they are.
Re: Fighting the New York Times' invasion of user privacy
#388"The New York Times is demanding that we turn over 20 million of your private ChatGPT conversations." As might any plaintiff. NYT might be the first of many others and the lawsuits may not be limited to copyright claims Why has OpenAI collected and stored 20 million conversations (including "deleted chats") What is the purpose of OpenAI storing millions of private conversations By contrast the purpose of NYT's reques…
> The documents requested are not being made public by the plaintiffs In fact, as far as I understand it, they could not be made public by the plaintiffs even if they wanted to do so, or even if one of their employees decided to leak them. That's because the plaintiffs themselves never actually see the documents. They will only be seen by the plaintiff's lawyers and any experts hired by those lawyers to analyze them.
Re: Fighting the New York Times' invasion of user privacy
#389Earlier quoted context omitted.
I'm not commenting on the core point of your comment, only the "why retain for 30 days" question. Im an age of automated backups and failovers, deleting can be really hard. Part of the answer could simply be that syncing a delete across all the redundancies (while ensuring those redundancies are reliable when a disaster happens and they need to recover or maintain uptime) may take days to weeks. Also the 30 days coul…
The most likely explanation is whatever storage solution they’re using has a built in “recycle bin” functionality and deleted data stays the for 30 days before it’s actually deleted. I see this a lot in very large databases. The recycle bin functionality is built in to the data store product.
Re: Fighting the New York Times' invasion of user privacy
#390It ridiculous for OpenAI to attempt to claim some moral high-ground here. They're a company that has demonstrated zero respect for the copyright or data privacy regulations of other organisations. I think they take users dignity and rights with a grain of salt. Their statements are all aspirational, "we're working toward de-identifying" etc. They've built one of the most powerful AIs ever seen and now they're claimin…