Live data from Hacker News

Fighting the New York Times' invasion of user privacy

openai.com

311–320 of 441 posts

Re: Fighting the New York Times' invasion of user privacy

#312
post #186

Earlier quoted context omitted.

>What is the purpose of OpenAI storing millions of private conversations Have you used ChatGPT? Your conversation history is on the left rail

Using RePair for compression I can also search inside compressed tarballs full of logs To do this, I first insert a blank line at the top of each log file before adding to the tarball IME, RePair is faster than compressing with zstd and the size reduction is almost the same The only "catch" is that RePair requires more memory during compression

Pardon, but do you have a link for this RePair compressor?

Unfortunately, different searches for this RePair you mentioned have only revealed links to resources for repairing broken air compressors, damaged compressed files, spinal injuries, etc.

Re: Fighting the New York Times' invasion of user privacy

#313

Earlier quoted context omitted.

Agreed, they could carefully coerce the model to more or less output some of their articles, but the premise that users were routinely doing this to bypass the paywall is silly.

Especially when you can just copy paste the url into Internet Archive and read it. And yet they aren't suing Internet Archive.

Let's be real, they are suing OpenAI because they have way more money than the Internet Archive and they would be happy with a cut

Re: Fighting the New York Times' invasion of user privacy

#314

Earlier quoted context omitted.

I really should take the "invest in companies you hate" advice seriously.

I don't hate them. It is just plain to see they have discovered no scalable business model outside of getting larger and larger amounts of capital from investors to utilize intellectual property from others (either directly in the model aka NYT, or indirectly via web searches) without any rights. It is better for all of us the sooner this fails.

to utilize intellectual property from others (either directly in the model aka NYT, or indirectly via web searches) without any rights

... and put the liability for retrieving said property and hence the culpability for copyright infringement on the enduser:

Since the output would only be generated as a result of user inputs known as prompts, it was not the defendants, but the respective user who would be liable for it, OpenAI had argued.

https://www.reuters.com/world/german-court-sides-with-plaint...

Re: Fighting the New York Times' invasion of user privacy

#315

If OpenAI hadn't used data from the NYT without permission in the first place this wouldn't have happened. That is the root cause of all this. I'm glad the NYT is fighting them. They've infringed the rights of almost every news outlet but someone has to bring this case.

They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...

Two idiot judges.

Re: Fighting the New York Times' invasion of user privacy

#317
> Fighting the New York Times' invasion of user privacy

OpenAI is lying about why they are doing this. They want the public to attack the New York Times because OpenAI probably broke the law in so many ways...

If they cared about privacy they would no training their models on that same private data. But here we are.

We need very strong regulations to rule in all these tech companies and make them work for their users instead of working against them and lying about it.

Re: Fighting the New York Times' invasion of user privacy

#319

Earlier quoted context omitted.

The most likely explanation is whatever storage solution they’re using has a built in “recycle bin” functionality and deleted data stays the for 30 days before it’s actually deleted. I see this a lot in very large databases. The recycle bin functionality is built in to the data store product.

That sounds very plausible.

The problem when dealing with any company that has proven itself untrustworthy is that by default the innocent "plausible" option is probably no longer the "likely" one.

And I say this knowing that intentionally deleting data is harder than it looks.

Re: Fighting the New York Times' invasion of user privacy

#320

If OpenAI hadn't used data from the NYT without permission in the first place this wouldn't have happened. That is the root cause of all this. I'm glad the NYT is fighting them. They've infringed the rights of almost every news outlet but someone has to bring this case.

They infringed nothing. Two judges have already ruled that training on copyrighted data is fair use https://www.whitecase.com/insight-alert/two-california-distr...

[flagged]
Post reply on HN